David A. Maltz

dblp:65/1459 · DBLP profile ↗
← Back
45ranked-venue papers
7as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 40 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
31 papers
Network management and operations · 27% Software-defined and programmable networks · 21% Network measurement and analytics · 14%
Computer architecture, parallel and distributed computing, and storage systems
13 papers
Cloud and datacenter computing · 79% Storage systems · 10% Distributed systems · 9%

Topics — the 30 heaviest of 89, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software-defined and programmable networks › programmable data plane
programmable switch
1.012026
Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch · NSDI 2026
Cloud and datacenter computing › cloud networking
cloud network services
1.012026
Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch · NSDI 2026
Cloud and datacenter computing › computation offloading
network function offloading
1.012026
Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch · NSDI 2026
Storage systems › networked storage › storage networking
RDMA storage
0.712023
Empowering Azure Storage with RDMA · NSDI 2023
Network management and operations › fault management
fault diagnosis
0.542020
Packet-Level Telemetry in Large Datacenter Networks · SIGCOMM 2015
Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing · SIGCOMM 2020
Towards highly reliable enterprise network services via inference of multi-level dependencies · SIGCOMM 2007
Cloud and datacenter computing
datacenter operations
0.412020
Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing · SIGCOMM 2020
Network management and operations
network configuration
0.452013
Unraveling the Complexity of Network Management · NSDI 2009
Mining policies from enterprise network configuration · Internet Measurement Conference 2009
Towards systematic design of enterprise networks · CoNEXT 2008
Cloud and datacenter computing › virtualization
network virtualization
0.312018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload
0.312018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Cloud and datacenter computing
virtualization
0.312018
Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018
Distributed systems › replication › replica management
replica selection
0.322013
Dynamic Request Splitting for Interactive Cloud Applications · IEEE J. Sel. Areas Commun. 2013
Dealer: application-aware request splitting for interactive cloud applications · CoNEXT 2012
Routing and switching › routing
anycast routing
0.212015
FastRoute: A Scalable Load-Aware Anycast Routing Architecture for Modern CDNs · NSDI 2015
Content delivery and video streaming
content delivery network
0.212015
FastRoute: A Scalable Load-Aware Anycast Routing Architecture for Modern CDNs · NSDI 2015
Network measurement and analytics
latency measurement
0.212015
Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis · SIGCOMM 2015
Internet of things and sensor networks › wireless sensor network
network diagnosis
0.212015
Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis · SIGCOMM 2015
Cloud and datacenter computing
cloud storage
0.212023
Empowering Azure Storage with RDMA · NSDI 2023
Software-defined and programmable networks › network update
congestion-free update
0.212013
zUpdate: updating data center networks with zero loss · SIGCOMM 2013
Datacenter networks
load balancing
0.212013
Per-packet load-balanced, low-latency routing for clos-based data center networks · CoNEXT 2013
Software-defined and programmable networks
network update
0.212013
zUpdate: updating data center networks with zero loss · SIGCOMM 2013
Cloud and datacenter computing › cloud networking
cloud load balancing
0.212013
Ananta: cloud scale load balancing · SIGCOMM 2013
Cloud and datacenter computing
datacenter infrastructure
0.212013
Challenges in cloud scale data centers · SIGMETRICS 2013
Cloud and datacenter computing › datacenter operations
datacenter workload characterization
0.212013
Challenges in cloud scale data centers · SIGMETRICS 2013
Cloud and datacenter computing
geo-distributed applications
0.212013
Dynamic Request Splitting for Interactive Cloud Applications · IEEE J. Sel. Areas Commun. 2013
Network optimization and economics › resource allocation
bandwidth optimization
0.112012
Surviving failures in bandwidth-constrained datacenters · SIGCOMM 2012
Network management and operations › network robustness
fault tolerance
0.112012
Surviving failures in bandwidth-constrained datacenters · SIGCOMM 2012
Distributed systems › distributed system architecture
geo-distributed systems
0.112012
Dealer: application-aware request splitting for interactive cloud applications · CoNEXT 2012
Privacy and data protection
anonymization
0.122009
Structure preserving anonymization of router configuration data · IEEE J. Sel. Areas Commun. 2009
Structure preserving anonymization of router configuration data · Internet Measurement Conference 2004
Internet architecture and protocols
domain name system
0.112011
Public DNS system and Global Traffic Management · INFOCOM 2011
Network measurement and analytics
internet measurement
0.112011
Public DNS system and Global Traffic Management · INFOCOM 2011
Routing and switching
MPLS
0.112011
Latency inflation with MPLS-based traffic engineering · Internet Measurement Conference 2011

Methods — techniques the papers use, named apart from their topics

machine learning · 0.9optimization · 0.6large-scale measurement · 0.5traffic trace analysis · 0.2packet-level telemetry · 0.2latency analysis · 0.2algorithms · 0.2replica selection · 0.2performance monitoring · 0.2network modeling · 0.2digit-reversal bouncing · 0.2component graph modeling · 0.1automation · 0.1packet-level trace analysis · 0.1SNMP statistics analysis · 0.1fingerprinting analysis · 0.1experiments · 0.1analysis · 0.1
YearPublicationVenuePosition
2026 Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch
Shaofeng Wu, Zhixiong Niu, Riff Jiang, Lawrence Lee, Junhua Zhai, Ze Gan, Vasundhara Volam, Prabhat Aravind, Prince Sunny, Prince George, Evan Langlais, Soumya Tiwari, Venkat Satish Katta, Weixi Chen, Rishiraj Hazarika, Sachin Jain, Deven Jagasia, Michal Zygmunt, Avijit Gupta, Neeraj Motwani, Pranjal Shrivastava, Anil Reddy Pannala, Kristina Moore, James Grantham, Anupam Pandey, Guohan Lu, Gerald DeGrace, Rishabh Tewari, Erica Lan, Deepak Bansal, David A. Maltz, Yongqiang Xiong, Hong Xu 0001
NSDI35
2023 Empowering Azure Storage with RDMA
Wei Bai 0001, Shanim Sainul Abdeen, Ankit Agrawal 0013, Krishan Kumar Attre, Paramvir Bahl, Ameya Bhagat, Gowri Bhaskara, Tanya Brokhman, Ahmad Cheema, Rebecca Chow, Jeff Cohen, Mahmoud Elhaddad, Vivek Ette, Igal Figlin, Daniel Firestone, Mathew George, Ilya German, Lakhmeet Ghai, Eric Green, Albert G. Greenberg, Randy Haagens, Matthew Hendel, Ridwan Howlader, Neetha John, Julia Johnstone, Tom Jolly, Greg Kramer, David Kruse, Erica Lan, Avi Levy, Marina Lipshteyn, Guohan Lu, Yuemin Lu, Xiakun Lu, Vadim Makhervaks, Ulad Malashanka, David A. Maltz, Ilias Marinos, Rohan Mehta, Sharda Murthi, Anup Namdhari, Aaron Ogus, Jitendra Padhye, Madhav Pandya, Douglas Phillips, Adrian Power, Suraj Puri, Shachar Raindel, Jordan Rhee, Anthony Russo, Maneesh Sah, Ali Sheriff, Chris Sparacino, Ashutosh Srivastava, Weixiang Sun, Nick Swanson, Fuhou Tian, Lukasz Tomczyk, Vamsi Vadlamuri, Alec Wolman, Joyce Yom, Yanzhao Zhang, Brian Zill
NSDI43
2020 Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing
abstract
Incident routing is critical for maintaining service level objectives in the cloud: the time-to-diagnosis can increase by 10x due to mis-routings. Properly routing incidents is challenging because of the complexity of today's data center (DC) applications and their dependencies. For instance, an application running on a VM might rely on a functioning host-server, remote-storage service, and virtual and physical network components. It is hard for any one team, rule-based system, or even machine learning solution to fully learn the complexity and solve the incident routing problem. We propose a different approach using per-team Scouts. Each teams' Scout acts as its gate-keeper --- it routes relevant incidents to the team and routes-away unrelated ones. We solve the problem through a collection of these Scouts. Our PhyNet Scout alone --- currently deployed in production --- reduces the time-to-mitigation of 65% of mis-routed incidents in our dataset.
Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri, Vincent Liu 0001, Ricardo Bianchini, Ramaswamy Aditya, Xiaohang Wang 0008, Henry Lee, David A. Maltz, Minlan Yu, Behnaz Arzani
SIGCOMM10
2018 Azure Accelerated Networking: SmartNICs in the Public Cloud
Daniel Firestone, Andrew Putnam, Sambrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian M. Caulfield, Eric S. Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitendra Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw 0001, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, Albert G. Greenberg
NSDI31
2017 Distributed load management algorithms in anycast-based CDNs
Abhishek Sinha, Pradeepkumar Mani, Jie Liu 0001, Ashley Flavel, David A. Maltz
Comput. Networks5
2015 FastRoute: A Scalable Load-Aware Anycast Routing Architecture for Modern CDNs
Ashley Flavel, Pradeepkumar Mani, David A. Maltz, Nick Holt, Jie Liu 0001, Oleg Surmachev
NSDI3
2015 Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis
abstract
Can we get network latency between any two servers at any time in large-scale data center networks? The collected latency data can then be used to address a series of challenges: telling if an application perceived latency issue is caused by the network or not, defining and tracking network service level agreement (SLA), and automatic network troubleshooting. We have developed the Pingmesh system for large-scale data center network latency measurement and analysis to answer the above question affirmatively. Pingmesh has been running in Microsoft data centers for more than four years, and it collects tens of terabytes of latency data per day. Pingmesh is widely used by not only network software developers and engineers, but also application and service developers and operators.
Chuanxiong Guo, Yingnong Dang, Ray Huang, David A. Maltz, Vin Wang, Varugis Kurien
SIGCOMM6
2015 Packet-Level Telemetry in Large Datacenter Networks
abstract
Debugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task.
Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001
SIGCOMM7
2014 Theia: Simple and Cheap Networking for Ultra-Dense Data Centers
abstract
Recent trends to pack data centers with more CPUs per rack have led to a scenario in which each individual rack may contain hundreds, or even thousands, of compute nodes using system-on-chip (SoC) architectures. At this increased scale, traditional rack-level star topologies with a top-of-rack (ToR) switch as the hub and servers as the leaves are no longer feasible in terms of monetary cost, physical space, and oversubscription. We propose Theia, an architecture to connect hundreds of SoC nodes within a rack, using inexpensive, low-latency, hardware elements to group the rack's servers into subsets which we term SubRacks. We then replace the traditional per-rack ToR with a low-latency, passive, circuit-style patch panel that interconnects these SubRacks. We explore alternatives for the rack-level topology implemented by this patch panel, and we consider approaches for interconnecting racks within a data center. Finally, we investigate options for routing over these new topologies. Our proposal of Theia is unique in that it offers the flexibility of a packet-switched networking over a fixed circuit topology.
Meg Walraed-Sullivan, Jitendra Padhye, David A. Maltz
HotNets3
2014 Re-evaluating the responsiveness of DNS-based network control
abstract
For more than a decade, special purpose DNS [1] servers have been used to direct users to different IP addresses based on characteristics such as proximity, load balancing and failover. In this paper, we examine two of the mechanisms these special purpose DNS servers may use to control user traffic:-short time-to-lives (TTL) and time-splicing multiple responses. We find that although 2% of users are still using a stale DNS response after 15 minutes (despite a 20 second TTL), 90% of user requests will have switched to a new answer within 60 seconds. Further, when trying to direct users to two different locations in the ratio 90:10, we find time-splicing CNAME and A responses (corresponding to the two locations, respectively) is incompatible with a commonly deployed recursive resolver, and results in a ratio policy closer to 10:90. Overall we found that although DNS TTLs are not always respected, we are able manage load at a coarse granularity - provided time-spliced responses are of the same response type.
Ashley Flavel, Pradeepkumar Mani, David A. Maltz
LANMAN3
2013 Per-packet load-balanced, low-latency routing for clos-based data center networks
abstract
Clos-based networks including Fat-tree and VL2 are being built in data centers, but existing per-flow based routing causes low network utilization and long latency tail. In this paper, by studying the structural properties of Fat-tree and VL2, we propose a per-packet round-robin based routing algorithm called Digit-Reversal Bouncing (DRB). DRB achieves perfect packet interleaving. Our analysis and simulations show that, compared with random-based load-balancing algorithms, DRB results in smaller and bounded queues even when traffic load approaches 100%, and it uses smaller re-sequencing buffer for absorbing out-of-order packet arrivals. Our implementation demonstrates that our design can be readily implemented with commodity switches. Experiments on our testbed, a Fat-tree with 54 servers, confirm our analysis and simulations, and further show that our design handles network failures in 1-2 seconds and has the desirable graceful performance degradation property.
Jiaxin Cao, Pengkun Yang, Chuanxiong Guo, Guohan Lu, Yixin Zheng, Yongqiang Xiong, David A. Maltz
CoNEXT10
2013 zUpdate: updating data center networks with zero loss
abstract
Datacenter networks (DCNs) are constantly evolving due to various updates such as switch upgrades and VM migrations. Each update must be carefully planned and executed in order to avoid disrupting many of the mission-critical, interactive applications hosted in DCNs. The key challenge arises from the inherent difficulty in synchronizing the changes to many devices, which may result in unforeseen transient link load spikes or even congestions. We present one primitive, zUpdate, to perform congestion-free network updates under asynchronous switch and traffic matrix changes. We formulate the update problem using a network model and apply our model to a variety of representative update scenarios in DCNs. We develop novel techniques to handle several practical challenges in realizing zUpdate as well as implement the zUpdate prototype on OpenFlow switches and deploy it on a testbed that resembles real DCN topology. Our results, from both real-world experiments and large-scale trace-driven simulations, show that zUpdate can effectively perform congestion-free updates in production DCNs.
Hongqiang Harry Liu, Ming Zhang 0005, Roger Wattenhofer, David A. Maltz
SIGCOMM6
2013 Ananta: cloud scale load balancing
abstract
Layer-4 load balancing is fundamental to creating scale-out web services. We designed and implemented Ananta, a scale-out layer-4 load balancer that runs on commodity hardware and meets the performance, reliability and operational requirements of multi-tenant cloud computing environments. Ananta combines existing techniques in routing and distributed systems in a unique way and splits the components of a load balancer into a consensus-based reliable control plane and a decentralized scale-out data plane. A key component of Ananta is an agent in every host that can take over the packet modification function from the load balancer, thereby enabling the load balancer to naturally scale with the size of the data center. Due to its distributed architecture, Ananta provides direct server return (DSR) and network address translation (NAT) capabilities across layer-2 boundaries. Multiple instances of Ananta have been deployed in the Windows Azure public cloud with combined bandwidth capacity exceeding 1Tbps. It is serving traffic needs of a diverse set of tenants, including the blob, table and relational storage services. With its scale-out data plane we can easily achieve more than 100Gbps throughput for a single public IP address. In this paper, we describe the requirements of a cloud-scale load balancer, the design of Ananta and lessons learnt from its implementation and operation in the Windows Azure public cloud.
Parveen Patel, Deepak Bansal, Ashwin Murthy, Albert G. Greenberg, David A. Maltz, Randy Kern, Marios Zikos, Changhoon Kim, Naveen Karri
SIGCOMM6
2013 Challenges in cloud scale data centers
abstract
Data centers are fascinating places, where the massive scale required to deliver on-line services like web search and cloud hosting turns minor issues into major challenges that must be addressed in the design of the physical infrastructure and the software platform. In this talk, I'll briefly overview the kinds of applications that run in mega-data centers and the workloads they place on the infrastructure. I'll then describe a number of challenges seen in Microsoft's data centers, with the goals of posing questions more than describing solutions and explaining how economic factors, technology issues, and software design interact when creating low-latency, low-cost, high availability services.
David A. Maltz
SIGMETRICS1
2013 Dynamic Request Splitting for Interactive Cloud Applications
abstract
Deploying interactive applications in the cloud is a challenge due to the high variability in performance of cloud services. In this paper, we present Dealer — a system that helps geo-distributed, interactive and multi-tier applications meet their stringent requirements on response time despite such variability. Our approach is motivated by the fact that, at any time, only a small number of application components of large multi-tier applications experience poor performance. Dealer continually monitors the performance of individual components and communication latencies between them to build a global view of the application. In serving any given request, Dealer seeks to minimize user response times by picking the best combination of replicas (potentially located across different data centers). While Dealer requires modifications to application code, we show the changes required are modest. Our evaluations on two multi-tier applications using real cloud deployments indicate the 90%ile of response times could be reduced by more than a factor of 6 under natural cloud dynamics. Our results indicate the cost of inter-data-center traffic with Dealer is minor, and that Dealer can in fact be used to reduce the overall operational costs of applications by up to 15% by leveraging the difference in billing plans of cloud instances.
Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai
IEEE J. Sel. Areas Commun.3
2012 Dealer: application-aware request splitting for interactive cloud applications
abstract
Deploying interactive applications in the cloud is a challenge due to the high variability in performance of cloud services. In this paper, we present Dealer-- a system that helps geo-distributed, interactive and multi-tier applications meet their stringent requirements on response time despite such variability. Our approach is motivated by the fact that, at any time, only a small number of application components of large multi-tier applications experience poor performance. Dealer abstracts application structure as a component graph, with nodes being application components and edges capturing inter-component communication patterns. Dealer continually monitors the performance of individual component replicas and communication latencies between replica pairs. In serving any given user request, Dealer seeks to minimize user response times by picking the best combination of replicas (potentially located across different data-centers). While Dealer does require modifications to application code, we show through integration with two multi-tier applications that the changes required are modest. Our evaluations on two multi-tier applications using real cloud deployments indicate the 90%ile of application response times could be reduced by a factor of 3 under natural cloud dynamics compared to conventional data-center redirection techniques which are agnostic of application structure.
Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai
CoNEXT3
2012 Surviving failures in bandwidth-constrained datacenters
abstract
Datacenter networks have been designed to tolerate failures of network equipment and provide sufficient bandwidth. In practice, however, failures and maintenance of networking and power equipment often make tens to thousands of servers unavailable, and network congestion can increase service latency. Unfortunately, there exists an inherent tradeoff between achieving high fault tolerance and reducing bandwidth usage in network core; spreading servers across fault domains improves fault tolerance, but requires additional bandwidth, while deploying servers together reduces bandwidth usage, but also decreases fault tolerance. We present a detailed analysis of a large-scale Web application and its communication patterns. Based on that, we propose and evaluate a novel optimization framework that achieves both high fault tolerance and significantly reduces bandwidth usage in the network core by exploiting the skewness in the observed communication patterns.
Peter Bodík, Ishai Menache, Mosharaf Chowdhury, Pradeepkumar Mani, David A. Maltz, Ion Stoica
SIGCOMM5
2012 NetPilot: automating datacenter network failure mitigation
abstract
Driven by the soaring demands for always-on and fast-response online services, modern datacenter networks have recently undergone tremendous growth. These networks often rely on commodity hardware to reach immense scale while keeping capital expenses under check. The downside is that commodity devices are prone to failures, raising a formidable challenge for network operators to promptly handle these failures with minimal disruptions to the hosted services.
Daniel Turner, Chao-Chih Chen, David A. Maltz, Xiaowei Yang 0001, Ming Zhang 0005
SIGCOMM4
2011 Latency inflation with MPLS-based traffic engineering
abstract
While MPLS has been extensively deployed in recent years, little is known about its behavior in practice. We examine the performance of MPLS in Microsoft's online service network (MSN), a well-provisioned multi-continent production network connecting tens of data centers. Using detailed traces collected over a 2-month period, we find that many paths experience significantly inflated latencies. We correlate occurrences of latency inflation with routers, links, and DC-pairs. This analysis sheds light on the causes of latency inflation and suggests several avenues for alleviating the problem.
Abhinav Pathak, Ming Zhang 0005, Y. Charlie Hu, Ratul Mahajan, David A. Maltz
Internet Measurement Conference5
2011 Public DNS system and Global Traffic Management
abstract
Cloud service providers operate data centers around the world, and they depend on Global Traffic Management systems to direct requests from clients to the most appropriate data center to serve the requests. While GTM systems have been in-use for years, they are attracting re-newed interests due to the rapid expansion of cloud service providers' networks, the introduction of public DNS systems, as well as new proposals to alter how they should work and what information local DNS servers (LDNS) should make available to drive the GTM systems. This paper uses large-scale measurements conducted from more than 5M clients to establish properties of the current Internet that affect the design of the GTM systems, such as the stretch between a client's actual position and its LDNS from GTM's perspective, the impact of public DNS systems, and the granularity at which GTM decisions should be made. The results can inform the debate over how GTM systems should be designed.
Cheng Huang 0002, David A. Maltz, Jin Li 0001, Albert G. Greenberg
INFOCOM2
2011 Profiling Network Performance for Multi-tier Data Center Applications
Minlan Yu, Albert G. Greenberg, David A. Maltz, Jennifer Rexford, Srikanth Kandula, Changhoon Kim
NSDI3
2011 Towards systematic design of enterprise networks
abstract
Enterprise networks are important, with size and complexity even surpassing carrier networks. Yet, the design of enterprise networks remains ad hoc and poorly understood. In this paper, we show how a systematic design approach can handle two key areas of enterprise design: virtual local area networks (VLANs) and reachability control. We focus on these tasks given their complexity, prevalence, and time-consuming nature. Our contributions are threefold. First, we show how these design tasks may be formulated in terms of network-wide performance, security, and resilience requirements. Our formulations capture the correctness and feasibility constraints on the design, and they model each task as one of optimizing desired criteria subject to the constraints. The optimization criteria may further be customized to meet operator-preferred design strategies. Second, we develop a set of algorithms to solve the problems that we formulate. Third, we demonstrate the feasibility and value of our systematic design approach through validation on a large-scale campus network with hundreds of routers and VLANs.
Yu-Wei Eric Sung, Xin Sun 0002, Sanjay G. Rao, Geoffrey G. Xie, David A. Maltz
IEEE/ACM Trans. Netw.5
2010 Network traffic characteristics of data centers in the wild
abstract
Although there is tremendous interest in designing improved networks for data centers, very little is known about the network-level traffic characteristics of data centers today. In this paper, we conduct an empirical study of the network traffic in 10 data centers belonging to three different categories, including university, enterprise campus, and cloud data centers. Our definition of cloud data centers includes not only data centers employed by large online service providers offering Internet-facing applications but also data centers used to host data-intensive (MapReduce style) applications). We collect and analyze SNMP statistics, topology and packet-level traces. We examine the range of applications deployed in these data centers and their placement, the flow-level and packet-level transmission properties of these applications, and their impact on network and link utilizations, congestion and packet drops. We describe the implications of the observed traffic patterns for data center internal traffic engineering as well as for recently proposed architectures for data center networks.
Theophilus Benson, Aditya Akella, David A. Maltz
Internet Measurement Conference3
2010 Data center TCP (DCTCP)
abstract
Cloud data centers host diverse applications, mixing workloads that require small predictable latency with others requiring large sustained throughput. In this environment, today's state-of-the-art TCP protocol falls short. We present measurements of a 6000 server production cluster and reveal impairments that lead to high application latencies, rooted in TCP's demands on the limited buffer space available in data center switches. For example, bandwidth hungry "background" flows build up queues at the switches, and thus impact the performance of latency sensitive "foreground" traffic.
Mohammad Alizadeh, Albert G. Greenberg, David A. Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, Murari Sridharan
SIGCOMM3
2010 Cloudward bound: planning for beneficial migration of enterprise applications to the cloud
abstract
In this paper, we tackle challenges in migrating enterprise services into hybrid cloud-based deployments, where enterprise operations are partly hosted on-premise and partly in the cloud. Such hybrid architectures enable enterprises to benefit from cloud-based architectures, while honoring application performance requirements, and privacy restrictions on what services may be migrated to the cloud. We make several contributions. First, we highlight the complexity inherent in enterprise applications today in terms of their multi-tiered nature, large number of application components, and interdependencies. Second, we have developed a model to explore the benefits of a hybrid migration approach. Our model takes into account enterprise-specific constraints, cost savings, and increased transaction delays and wide-area communication costs that may result from the migration. Evaluations based on real enterprise applications and Azure-based cloud deployments show the benefits of a hybrid migration approach, and the importance of planning which components to migrate. Third, we shed insight on security policies associated with enterprise applications in data centers. We articulate the importance of ensuring assurable reconfiguration of security policies as enterprise applications are migrated to the cloud. We present algorithms to achieve this goal, and demonstrate their efficacy on realistic migration scenarios.
Mohammad Y. Hajjat, Xin Sun 0002, Yu-Wei Eric Sung, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai, Mohit Tawarmalani
SIGCOMM4
2010 MMS: An autonomic network-layer foundation for network management
abstract
Networks cannot be managed without management plane communications among geographically distributed network devices and control agents. Unfortunately, the mechanisms used in commercial networks to support management plane communications are often hard to configure, insufficiently secured, and/or suboptimal in performance. This paper presents the design and implementation of the Meta-Management System (MMS), a network-layer subsystem that provides robust autonomic support for management plane communications. We demonstrate the practicality of the MMS via a fully functional implementation that runs on commodity hardware, and experimentally show that the MMS is efficient and scalable. The MMS software is freely available.
Hemant Gogineni, Albert G. Greenberg, David A. Maltz, T. S. Eugene Ng, Hong Yan 0002, Hui Zhang 0001
IEEE J. Sel. Areas Commun.3
2009 Mining policies from enterprise network configuration
abstract
Few studies so far have examined the nature of reachability policies in enterprise networks. A better understanding of reachability policies could both inform future approaches to network design as well as current network configuration mechanisms. In this paper, we introduce the notion of a policy unit, which is an abstract representation of how the policies implemented in a network apply to different network hosts. We develop an approach for reverse-engineering a network's policy units from its router configuration. We apply this approach to the configurations of five productions networks, including three university and two private enterprises. Through our empirical study, we validate that policy units capture useful characteristics of a network's policy. We also obtain insights into the nature of the policies implemented in modern enterprises. For example, we find most hosts in these networks are subject to nearly identical reachability policies at Layer 3.
Theophilus Benson, Aditya Akella, David A. Maltz
Internet Measurement Conference3
2009 Unraveling the Complexity of Network Management
Theophilus Benson, Aditya Akella, David A. Maltz
NSDI3
2009 VL2: a scalable and flexible data center network
abstract
To be agile and cost effective, data centers should allow dynamic resource allocation across large server pools. In particular, the data center network should enable any server to be assigned to any service. To meet these goals, we present VL2, a practical network architecture that scales to support huge data centers with uniform high capacity between servers, performance isolation between services, and Ethernet layer-2 semantics. VL2 uses (1) flat addressing to allow service instances to be placed anywhere in the network, (2) Valiant Load Balancing to spread traffic uniformly across network paths, and (3) end-system based address resolution to scale to large server pools, without introducing complexity to the network control plane. VL2's design is driven by detailed measurements of traffic and fault data from a large operational cloud service provider. VL2's implementation leverages proven network technologies, already available at low cost in high-speed hardware implementations, to build a scalable and reliable network architecture. As a result, VL2 networks can be deployed today, and we have built a working prototype. We evaluate the merits of the VL2 design using measurement, analysis, and experiments. Our VL2 prototype shuffles 2.7 TB of data among 75 servers in 395 seconds - sustaining a rate that is 94% of the maximum possible.
Albert G. Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A. Maltz, Parveen Patel, Sudipta Sengupta
SIGCOMM7
2009 Structure preserving anonymization of router configuration data
abstract
A repository of router configuration files from production networks would provide the research community with a treasure trove of data about network topologies, routing designs, and security policies. However, configuration files have been largely unobtainable precisely because they provide detailed information that could be exploited by competitors and attackers. This paper describes a method for anonymizing router configuration files by removing all information that connects the data to the identity of the underlying network, while still preserving the structure of information that makes the data valuable to networking researchers. Anonymizing configuration files has unusual requirements, including preserving relationships between elements of data, anonymizing regular expressions, and robustly coping with more than 200 versions of the configuration language. Conventional tools and techniques are poorly suited to the problem. Our anonymization method has been validated with a major carrier, earning unprivileged researchers access to the configuration files of thousands of routers in hundreds of networks. Through example analysis, we demonstrate that the anonymized data retains the key properties of the network design. The paper sets out techniques that could be used in an attempt to break the anonymization, and it concludes our anonymization techniques are most applicable to enterprise networks, because the large number of enterprises and the difficulty of probing them from the outside make it hard to recognize an anonymized network based solely on publicly-available information about its topology or configuration. When applied to backbone networks, which are few in number and many of whose properties can be publicly measured, the anonymization might be broken by fingerprinting techniques described in this paper.
David A. Maltz, Jibin Zhan, Gísli Hjálmtýsson, Albert G. Greenberg, Jennifer Rexford, Geoffrey G. Xie, Hui Zhang 0001
IEEE J. Sel. Areas Commun.1
2008 Towards systematic design of enterprise networks
abstract
Enterprise networks are important, with size and complexity even surpassing carrier networks. Yet, the design of enterprise networks is ad-hoc and poorly understood. In this paper, we show how a systematic design approach can handle two key areas of enterprise design: virtual local area networks (VLANs) and reachability control. We focus on these tasks given their complexity, prevalence, and time-consuming nature. Our contributions are three-fold. First, we show how these design tasks may be formulated in terms of network-wide performance, security, and resilience requirements. Our formulations capture the correctness and feasibility constraints on the design, and they model each task as one of optimizing desired criteria subject to the constraints. The optimization criteria may further be customized to meet operator-preferred design strategies. Second, we develop a set of algorithms to solve the problems that we formulate. Third, we demonstrate the feasibility and value of our systematic design approach through validation on a large-scale campus network with hundreds of routers and VLANs.
Yu-Wei Eric Sung, Sanjay G. Rao, Geoffrey G. Xie, David A. Maltz
CoNEXT4
2007 Fast Variational Inference for Large-scale Internet Diagnosis
abstract
Web servers on the Internet need to maintain high reliability, but the cause of intermittent failures of web transactions is non-obvious. We use Bayesian inference to diagnose problems with web services. This diagnosis problem is far larger than any previously attempted: it requires inference of 10^4 possible faults from 10^5 observations. Further, such inference must be performed in less than a second. Inference can be done at this speed by combining a variational approximation, a mean-field approximation, and the use of stochastic gradient descent to optimize a variational cost function. We use this fast inference to diagnose a time series of anomalous HTTP requests taken from a real web service. The inference is fast enough to analyze network logs with billions of entries in a matter of hours.
John C. Platt, Emre Kiciman, David A. Maltz
NIPS3
2007 Tesseract: A 4D Network Control Plane
Hong Yan 0002, David A. Maltz, T. S. Eugene Ng, Hemant Gogineni, Hui Zhang 0001, Zheng Cai
NSDI2
2007 Towards highly reliable enterprise network services via inference of multi-level dependencies
abstract
Localizing the sources of performance problems in large enterprise networks is extremely challenging. Dependencies are numerous, complex and inherently multi-level, spanning hardware and software components across the network and the computing infrastructure. To exploit these dependencies for fast, accurate problem localization, we introduce an Inference Graph model, which is well-adapted to user-perceptible problems rooted in conditions giving rise to both partial service degradation and hard faults. Further, we introduce the Sherlock system to discover Inference Graphs in the operational enterprise, infer critical attributes, and then leverage the result to automatically detect and localize problems. To illuminate strengths and limitations of the approach, we provide results from a prototype deployment in a large enterprise network, as well as from testbed emulations and simulations. In particular, we find that taking into account multi-level structure leads to a 30% improvement in fault localization, as compared to two-level approaches.
Paramvir Bahl, Ranveer Chandra, Albert G. Greenberg, Srikanth Kandula, David A. Maltz, Ming Zhang 0005
SIGCOMM5
2006 Discovering Dependencies for Network Management
Paramvir Bahl, Paul Barham 0001, Richard Black, Ranveer Chandra, Moisés Goldszmidt, Rebecca Isaacs, Srikanth Kandula, John MacCormick, David A. Maltz, Richard Mortier, Michal Wawrzoniak, Ming Zhang 0005
HotNets10
2005 On static reachability analysis of IP networks
abstract
The primary purpose of a network is to provide reachability between applications running on end hosts. In this paper, we describe how to compute the reachability a network provides from a snapshot of the configuration state from each of the routers. Our primary contribution is the precise definition of the potential reachability of a network and a substantial simplification of the problem through a unified modeling of packet filters and routing protocols. In the end, we reduce a complex, important practical problem to computing the transitive closure to set union and intersection operations on reachability set representations. We then extend our algorithm to model the influence of packet transformations (e.g., by NATs or ToS remapping) along the path. Our technique for static analysis of network reachability is valuable for verifying the intent of the network designer, troubleshooting reachability problems, and performing "what-if" analysis of failure scenarios.
Geoffrey G. Xie, Jibin Zhan, David A. Maltz, Hui Zhang 0001, Albert G. Greenberg, Gísli Hjálmtýsson, Jennifer Rexford
INFOCOM3
2005 Worm Origin Identification Using Random Moonwalks
abstract
We propose a novel technique that can determine both the host responsible for originating a propagating worm attack and the set of attack flows that make up the initial stages of the attack tree via which the worm infected successive generations of victims. We argue that knowledge of both is important for combating worms: knowledge of the origin supports law enforcement, and knowledge of the causal flows that advance the attack supports diagnosis of how network defenses were breached. Our technique exploits the "wide tree" shape of a worm propagation emanating from the source by performing random "moonwalks" backward in time along paths of flows. Correlating the repeated walks reveals the initial causal flows, thereby aiding in identifying the source. Using analysis, simulation, and experiments with real world traces, we show how the technique works against both today's fast propagating worms and stealthy worms that attempt to hide their attack flows among background traffic.
Yinglian Xie, Vyas Sekar, David A. Maltz, Michael K. Reiter, Hui Zhang 0001
S&P3
2004 Structure preserving anonymization of router configuration data
abstract
A repository of router configuration files from production networks would provide the research community with a treasure trove of data about network topologies, routing designs, and security policies. However, configuration files have been largely unobtainable precisely because they provide detailed information that could be exploited by competitors and attackers. This paper describes a method for anonymizing router configuration files by removing all information that connects the data to the identity of the originating network, while still preserving the structure of information that makes the data valuable to networking researchers. Anonymizing configuration files has unusual requirements, including preserving relationships between elements of data, anonymizing regular expressions, and robustly coping with more than 200 versions of the configuration language, that mean conventional tools and techniques are poorly suited to the problem. Our anonymization method has been validated with a major carrier, earning unprivileged researchers access to the configuration files of more than 7600 routers in 31 networks. Through example analysis, we demonstrate that the anonymized data retains the key properties of the network design. We believe that applying our single-blind methodology to a large number of production networks from different sources would be of tremendous value to both the research and operations communities.
David A. Maltz, Jibin Zhan, Geoffrey G. Xie, Hui Zhang 0001, Gísli Hjálmtýsson, Albert G. Greenberg, Jennifer Rexford
Internet Measurement Conference1
2004 Routing design in operational networks: a look from the inside
abstract
In any IP network, routing protocols provide the intelligence that takes a collection of physical links and transforms them into a network that enables packets to travel from one host to another. Though routing design is arguably the single most important design task for large IP networks, there has been very little systematic investigation into how routing protocols are actually used in production networks to implement the goals of network architects. We have developed a methodology for reverse engineering a coherent global view of a network's routing design from the static analysis of dumps of the local configuration state of each router. Starting with a set of 8,035 configuration files, we have applied this method to 31 production networks. In this paper we present a detailed examination of how routing protocols are used in operational networks. In particular, the results show the conventional model of interior and exterior gateway protocols is insufficient to describe the diverse set of mechanisms used by architects, and we provide examples of the more unusual designs and examine their trade-offs. We discuss the strengths and weaknesses of our methodology, and argue that it opens paths towards new understandings of network behavior and design.
Geoffrey G. Xie, Jibin Zhan, David A. Maltz, Hui Zhang 0001, Albert G. Greenberg, Gísli Hjálmtýsson
SIGCOMM3
2002 MSOCKS+: an architecture for transport layer mobility
Pravin Bhagwat, David A. Maltz, Adrian Segall
Comput. Networks2
2000 Quantitative lessons from a full-scale multi-hop wireless ad hoc network testbed
abstract
This paper presents preliminary quantitative results from data collected during runs of our multi-hop wireless ad hoc network testbed. The network successfully carried a composite workload including voice, bulk data, and real-time data. Careful analysis of recorded runs highlights radio propagation issues that network protocols will need to address in the future.
David A. Maltz, Josh Broch, David B. Johnson 0001
WCNC1
1999 The effects of on-demand behavior in routing protocols for multihop wireless ad hoc networks
abstract
A number of different routing protocols proposed for use in multihop wireless ad hoc networks are based in whole or in part on what can be described as on-demand behavior. By on-demand behavior, we mean approaches based only on reaction to the offered traffic being handled by the routing protocol. In this paper, we analyze the use of on-demand behavior in such protocols, focusing on its effect on the routing protocol's forwarding latency, overhead cost, and route caching correctness, drawing examples from detailed simulation of the dynamic source routing (DSR) protocol. We study the protocol's behavior and the changes introduced by variations on some of the mechanisms that make up the protocol, examining which mechanisms have the greatest impact and exploring the tradeoffs that exist between them.
David A. Maltz, Josh Broch, Jorjeta G. Jetcheva, David B. Johnson 0001
IEEE J. Sel. Areas Commun.1
1998 MSOCKS: An Architecture for Transport Layer Mobility
abstract
Mobile nodes of the future will be equiped with multiple network interfaces to take advantage of overlay networks, yet no current mobility systems provide full support for the simultaneous use of multiple interfaces. The need for such support arises when multiple connectivity options are available with different cost, coverage, latency and bandwidth characteristics, and applications want their data to flow over the interface that best matches the characteristics of the data. We present an architecture called transport layer mobility that allows mobile nodes to not only change their point of attachment to the Internet, but also to control which network interfaces are used for the different kinds of data leaving from and arriving at the mobile node. We implement our transport layer mobility scheme using a split connection proxy architecture and a new technique called TCP splice that gives split connection proxy systems the same end to-end semantics as normal TCP connections.
David A. Maltz, Pravin Bhagwat
INFOCOM1
1998 A Performance Comparison of Multi-Hop Wireless Ad Hoc Network Routing Protocols
abstract
An ad hoc networkis a collwtion of wirelessmobile nodes dynamically forminga temporarynetworkwithouttheuse of anyexistingnetworkirrfrastructureor centralizedadministration.Dueto the limitedtransmissionrange of ~vlreless nenvorkinterfaces,multiplenetwork"hops"maybe neededfor onenodeto exchangedata ivithanotheracrox the network.In recentyears, a ttiery of nelvroutingprotocols~geted specificallyat this environment have been developed.butlittle pcrfomrartwinformationon mch protocol and no ralistic performancecomparisonbehvwrrthem ISavailable.~Is paper presentsthe results of a derailedpacket-levelsimulationcomparing fourmulti-hopwirelessad hoc networkroutingprotocolsthatcovera range of design choices: DSDV,TORA, DSR and AODV.\Vehave extended the /~r-2networksimulatorto accuratelymodelthe MACand physical-layer behaviorof the IEEE 802.1I wirelessLAN standard,includinga realistic wtrelesstransmissionchannelmodel, and present the resultsof simulations of net(vorksof 50 mobilenodes.This work was supported in
Josh Broch, David A. Maltz, David B. Johnson 0001, Yih-Chun Hu, Jorjeta G. Jetcheva
MobiCom2
1995 Pointing the Way: Active Collaborative Filtering
abstract
Collaborative filtering is based on the premise that people looking for information should be able to make use of what others have already found and evaluated.Current collaborative filtoing systems provide tools for readers to filter documents based on aggregated ratings over a changing group of readers.Motivated by the results of a study of information sharing, we describe a different type of collaborative filtering system in which people who find interesting documents actively send "pointers" to those documents to their colleagues.A "pointer" contains a hypertext link to the source document as well as contextual information to help the recipient determine the interest and relevance of the document prior to accessing it.Preliminary data suggest that people are using the system in anticipated and unanticipated ways, as well as creating information "digests".
David A. Maltz, Kate Ehrlich
CHI1