EDBT 2026 Demo / reviewers in the wild / expert
David A. Maltz
dblp:65/1459
· DBLP profile ↗
45ranked-venue papers
7as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 40 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
31 papers |
Network management and operations · 27% Software-defined and programmable networks · 21% Network measurement and analytics · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Cloud and datacenter computing · 79% Storage systems · 10% Distributed systems · 9% |
Topics — the 30 heaviest of 89, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software-defined and programmable networks › programmable data plane
programmable switch |
1.0 | 1 | 2026 | Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch · NSDI 2026 |
Cloud and datacenter computing › cloud networking
cloud network services |
1.0 | 1 | 2026 | Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch · NSDI 2026 |
Cloud and datacenter computing › computation offloading
network function offloading |
1.0 | 1 | 2026 | Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch · NSDI 2026 |
Storage systems › networked storage › storage networking
RDMA storage |
0.7 | 1 | 2023 | Empowering Azure Storage with RDMA · NSDI 2023 |
Network management and operations › fault management
fault diagnosis |
0.5 | 4 | 2020 | Packet-Level Telemetry in Large Datacenter Networks · SIGCOMM 2015 Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing · SIGCOMM 2020 Towards highly reliable enterprise network services via inference of multi-level dependencies · SIGCOMM 2007 |
Cloud and datacenter computing
datacenter operations |
0.4 | 1 | 2020 | Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing · SIGCOMM 2020 |
Network management and operations
network configuration |
0.4 | 5 | 2013 | Unraveling the Complexity of Network Management · NSDI 2009 Mining policies from enterprise network configuration · Internet Measurement Conference 2009 Towards systematic design of enterprise networks · CoNEXT 2008 |
Cloud and datacenter computing › virtualization
network virtualization |
0.3 | 1 | 2018 | Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018 |
Cloud and datacenter computing › computation offloading › network function offloading
SmartNIC offload |
0.3 | 1 | 2018 | Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018 |
Cloud and datacenter computing
virtualization |
0.3 | 1 | 2018 | Azure Accelerated Networking: SmartNICs in the Public Cloud · NSDI 2018 |
Distributed systems › replication › replica management
replica selection |
0.3 | 2 | 2013 | Dynamic Request Splitting for Interactive Cloud Applications · IEEE J. Sel. Areas Commun. 2013 Dealer: application-aware request splitting for interactive cloud applications · CoNEXT 2012 |
Routing and switching › routing
anycast routing |
0.2 | 1 | 2015 | FastRoute: A Scalable Load-Aware Anycast Routing Architecture for Modern CDNs · NSDI 2015 |
Content delivery and video streaming
content delivery network |
0.2 | 1 | 2015 | FastRoute: A Scalable Load-Aware Anycast Routing Architecture for Modern CDNs · NSDI 2015 |
Network measurement and analytics
latency measurement |
0.2 | 1 | 2015 | Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis · SIGCOMM 2015 |
Internet of things and sensor networks › wireless sensor network
network diagnosis |
0.2 | 1 | 2015 | Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and Analysis · SIGCOMM 2015 |
Cloud and datacenter computing
cloud storage |
0.2 | 1 | 2023 | Empowering Azure Storage with RDMA · NSDI 2023 |
Software-defined and programmable networks › network update
congestion-free update |
0.2 | 1 | 2013 | zUpdate: updating data center networks with zero loss · SIGCOMM 2013 |
Datacenter networks
load balancing |
0.2 | 1 | 2013 | Per-packet load-balanced, low-latency routing for clos-based data center networks · CoNEXT 2013 |
Software-defined and programmable networks
network update |
0.2 | 1 | 2013 | zUpdate: updating data center networks with zero loss · SIGCOMM 2013 |
Cloud and datacenter computing › cloud networking
cloud load balancing |
0.2 | 1 | 2013 | Ananta: cloud scale load balancing · SIGCOMM 2013 |
Cloud and datacenter computing
datacenter infrastructure |
0.2 | 1 | 2013 | Challenges in cloud scale data centers · SIGMETRICS 2013 |
Cloud and datacenter computing › datacenter operations
datacenter workload characterization |
0.2 | 1 | 2013 | Challenges in cloud scale data centers · SIGMETRICS 2013 |
Cloud and datacenter computing
geo-distributed applications |
0.2 | 1 | 2013 | Dynamic Request Splitting for Interactive Cloud Applications · IEEE J. Sel. Areas Commun. 2013 |
Network optimization and economics › resource allocation
bandwidth optimization |
0.1 | 1 | 2012 | Surviving failures in bandwidth-constrained datacenters · SIGCOMM 2012 |
Network management and operations › network robustness
fault tolerance |
0.1 | 1 | 2012 | Surviving failures in bandwidth-constrained datacenters · SIGCOMM 2012 |
Distributed systems › distributed system architecture
geo-distributed systems |
0.1 | 1 | 2012 | Dealer: application-aware request splitting for interactive cloud applications · CoNEXT 2012 |
Privacy and data protection
anonymization |
0.1 | 2 | 2009 | Structure preserving anonymization of router configuration data · IEEE J. Sel. Areas Commun. 2009 Structure preserving anonymization of router configuration data · Internet Measurement Conference 2004 |
Internet architecture and protocols
domain name system |
0.1 | 1 | 2011 | Public DNS system and Global Traffic Management · INFOCOM 2011 |
Network measurement and analytics
internet measurement |
0.1 | 1 | 2011 | Public DNS system and Global Traffic Management · INFOCOM 2011 |
Routing and switching
MPLS |
0.1 | 1 | 2011 | Latency inflation with MPLS-based traffic engineering · Internet Measurement Conference 2011 |
Methods — techniques the papers use, named apart from their topics
machine learning · 0.9optimization · 0.6large-scale measurement · 0.5traffic trace analysis · 0.2packet-level telemetry · 0.2latency analysis · 0.2algorithms · 0.2replica selection · 0.2performance monitoring · 0.2network modeling · 0.2digit-reversal bouncing · 0.2component graph modeling · 0.1automation · 0.1packet-level trace analysis · 0.1SNMP statistics analysis · 0.1fingerprinting analysis · 0.1experiments · 0.1analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitch
Shaofeng Wu, Zhixiong Niu, Riff Jiang, Lawrence Lee, Junhua Zhai, Ze Gan, Vasundhara Volam, Prabhat Aravind, Prince Sunny, Prince George, Evan Langlais, Soumya Tiwari, Venkat Satish Katta, Weixi Chen, Rishiraj Hazarika, Sachin Jain, Deven Jagasia, Michal Zygmunt, Avijit Gupta, Neeraj Motwani, Pranjal Shrivastava, Anil Reddy Pannala, Kristina Moore, James Grantham, Anupam Pandey, Guohan Lu, Gerald DeGrace, Rishabh Tewari, Erica Lan, Deepak Bansal, David A. Maltz, Yongqiang Xiong, Hong Xu 0001 |
NSDI | 35 |
| 2023 | Empowering Azure Storage with RDMA
Wei Bai 0001, Shanim Sainul Abdeen, Ankit Agrawal 0013, Krishan Kumar Attre, Paramvir Bahl, Ameya Bhagat, Gowri Bhaskara, Tanya Brokhman, Ahmad Cheema, Rebecca Chow, Jeff Cohen, Mahmoud Elhaddad, Vivek Ette, Igal Figlin, Daniel Firestone, Mathew George, Ilya German, Lakhmeet Ghai, Eric Green, Albert G. Greenberg, Randy Haagens, Matthew Hendel, Ridwan Howlader, Neetha John, Julia Johnstone, Tom Jolly, Greg Kramer, David Kruse, Erica Lan, Avi Levy, Marina Lipshteyn, Guohan Lu, Yuemin Lu, Xiakun Lu, Vadim Makhervaks, Ulad Malashanka, David A. Maltz, Ilias Marinos, Rohan Mehta, Sharda Murthi, Anup Namdhari, Aaron Ogus, Jitendra Padhye, Madhav Pandya, Douglas Phillips, Adrian Power, Suraj Puri, Shachar Raindel, Jordan Rhee, Anthony Russo, Maneesh Sah, Ali Sheriff, Chris Sparacino, Ashutosh Srivastava, Weixiang Sun, Nick Swanson, Fuhou Tian, Lukasz Tomczyk, Vamsi Vadlamuri, Alec Wolman, Joyce Yom, Yanzhao Zhang, Brian Zill |
NSDI | 43 |
| 2020 | Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingabstractIncident routing is critical for maintaining service level objectives in the cloud: the time-to-diagnosis can increase by 10x due to mis-routings. Properly routing incidents is challenging because of the complexity of today's data center (DC) applications and their dependencies. For instance, an application running on a VM might rely on a functioning host-server, remote-storage service, and virtual and physical network components. It is hard for any one team, rule-based system, or even machine learning solution to fully learn the complexity and solve the incident routing problem. We propose a different approach using per-team Scouts. Each teams' Scout acts as its gate-keeper --- it routes relevant incidents to the team and routes-away unrelated ones. We solve the problem through a collection of these Scouts. Our PhyNet Scout alone --- currently deployed in production --- reduces the time-to-mitigation of 65% of mis-routed incidents in our dataset. Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri, Vincent Liu 0001, Ricardo Bianchini, Ramaswamy Aditya, Xiaohang Wang 0008, Henry Lee, David A. Maltz, Minlan Yu, Behnaz Arzani |
SIGCOMM | 10 |
| 2018 | Azure Accelerated Networking: SmartNICs in the Public Cloud
Daniel Firestone, Andrew Putnam, Sambrama Mundkur, Derek Chiou, Alireza Dabagh, Mike Andrewartha, Hari Angepat, Vivek Bhanu, Adrian M. Caulfield, Eric S. Chung, Harish Kumar Chandrappa, Somesh Chaturmohta, Matt Humphrey, Jack Lavier, Norman Lam, Fengfen Liu, Kalin Ovtcharov, Jitendra Padhye, Gautham Popuri, Shachar Raindel, Tejas Sapre, Mark Shaw 0001, Gabriel Silva, Madhan Sivakumar, Nisheeth Srivastava, Anshuman Verma, Qasim Zuhair, Deepak Bansal, Doug Burger, Kushagra Vaid, David A. Maltz, Albert G. Greenberg |
NSDI | 31 |
| 2017 | Distributed load management algorithms in anycast-based CDNs
Abhishek Sinha, Pradeepkumar Mani, Jie Liu 0001, Ashley Flavel, David A. Maltz |
Comput. Networks | 5 |
| 2015 | FastRoute: A Scalable Load-Aware Anycast Routing Architecture for Modern CDNs
Ashley Flavel, Pradeepkumar Mani, David A. Maltz, Nick Holt, Jie Liu 0001, Oleg Surmachev |
NSDI | 3 |
| 2015 | Pingmesh: A Large-Scale System for Data Center Network Latency Measurement and AnalysisabstractCan we get network latency between any two servers at any time in large-scale data center networks? The collected latency data can then be used to address a series of challenges: telling if an application perceived latency issue is caused by the network or not, defining and tracking network service level agreement (SLA), and automatic network troubleshooting. We have developed the Pingmesh system for large-scale data center network latency measurement and analysis to answer the above question affirmatively. Pingmesh has been running in Microsoft data centers for more than four years, and it collects tens of terabytes of latency data per day. Pingmesh is widely used by not only network software developers and engineers, but also application and service developers and operators. Chuanxiong Guo, Yingnong Dang, Ray Huang, David A. Maltz, Vin Wang, Varugis Kurien |
SIGCOMM | 6 |
| 2015 | Packet-Level Telemetry in Large Datacenter NetworksabstractDebugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task. Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 7 |
| 2014 | Theia: Simple and Cheap Networking for Ultra-Dense Data CentersabstractRecent trends to pack data centers with more CPUs per rack have led to a scenario in which each individual rack may contain hundreds, or even thousands, of compute nodes using system-on-chip (SoC) architectures. At this increased scale, traditional rack-level star topologies with a top-of-rack (ToR) switch as the hub and servers as the leaves are no longer feasible in terms of monetary cost, physical space, and oversubscription. We propose Theia, an architecture to connect hundreds of SoC nodes within a rack, using inexpensive, low-latency, hardware elements to group the rack's servers into subsets which we term SubRacks. We then replace the traditional per-rack ToR with a low-latency, passive, circuit-style patch panel that interconnects these SubRacks. We explore alternatives for the rack-level topology implemented by this patch panel, and we consider approaches for interconnecting racks within a data center. Finally, we investigate options for routing over these new topologies. Our proposal of Theia is unique in that it offers the flexibility of a packet-switched networking over a fixed circuit topology. Meg Walraed-Sullivan, Jitendra Padhye, David A. Maltz |
HotNets | 3 |
| 2014 | Re-evaluating the responsiveness of DNS-based network controlabstractFor more than a decade, special purpose DNS [1] servers have been used to direct users to different IP addresses based on characteristics such as proximity, load balancing and failover. In this paper, we examine two of the mechanisms these special purpose DNS servers may use to control user traffic:-short time-to-lives (TTL) and time-splicing multiple responses. We find that although 2% of users are still using a stale DNS response after 15 minutes (despite a 20 second TTL), 90% of user requests will have switched to a new answer within 60 seconds. Further, when trying to direct users to two different locations in the ratio 90:10, we find time-splicing CNAME and A responses (corresponding to the two locations, respectively) is incompatible with a commonly deployed recursive resolver, and results in a ratio policy closer to 10:90. Overall we found that although DNS TTLs are not always respected, we are able manage load at a coarse granularity - provided time-spliced responses are of the same response type. Ashley Flavel, Pradeepkumar Mani, David A. Maltz |
LANMAN | 3 |
| 2013 | Per-packet load-balanced, low-latency routing for clos-based data center networksabstractClos-based networks including Fat-tree and VL2 are being built in data centers, but existing per-flow based routing causes low network utilization and long latency tail. In this paper, by studying the structural properties of Fat-tree and VL2, we propose a per-packet round-robin based routing algorithm called Digit-Reversal Bouncing (DRB). DRB achieves perfect packet interleaving. Our analysis and simulations show that, compared with random-based load-balancing algorithms, DRB results in smaller and bounded queues even when traffic load approaches 100%, and it uses smaller re-sequencing buffer for absorbing out-of-order packet arrivals. Our implementation demonstrates that our design can be readily implemented with commodity switches. Experiments on our testbed, a Fat-tree with 54 servers, confirm our analysis and simulations, and further show that our design handles network failures in 1-2 seconds and has the desirable graceful performance degradation property. Jiaxin Cao, Pengkun Yang, Chuanxiong Guo, Guohan Lu, Yixin Zheng, Yongqiang Xiong, David A. Maltz |
CoNEXT | 10 |
| 2013 | zUpdate: updating data center networks with zero lossabstractDatacenter networks (DCNs) are constantly evolving due to various updates such as switch upgrades and VM migrations. Each update must be carefully planned and executed in order to avoid disrupting many of the mission-critical, interactive applications hosted in DCNs. The key challenge arises from the inherent difficulty in synchronizing the changes to many devices, which may result in unforeseen transient link load spikes or even congestions. We present one primitive, zUpdate, to perform congestion-free network updates under asynchronous switch and traffic matrix changes. We formulate the update problem using a network model and apply our model to a variety of representative update scenarios in DCNs. We develop novel techniques to handle several practical challenges in realizing zUpdate as well as implement the zUpdate prototype on OpenFlow switches and deploy it on a testbed that resembles real DCN topology. Our results, from both real-world experiments and large-scale trace-driven simulations, show that zUpdate can effectively perform congestion-free updates in production DCNs. Hongqiang Harry Liu, Ming Zhang 0005, Roger Wattenhofer, David A. Maltz |
SIGCOMM | 6 |
| 2013 | Ananta: cloud scale load balancingabstractLayer-4 load balancing is fundamental to creating scale-out web services. We designed and implemented Ananta, a scale-out layer-4 load balancer that runs on commodity hardware and meets the performance, reliability and operational requirements of multi-tenant cloud computing environments. Ananta combines existing techniques in routing and distributed systems in a unique way and splits the components of a load balancer into a consensus-based reliable control plane and a decentralized scale-out data plane. A key component of Ananta is an agent in every host that can take over the packet modification function from the load balancer, thereby enabling the load balancer to naturally scale with the size of the data center. Due to its distributed architecture, Ananta provides direct server return (DSR) and network address translation (NAT) capabilities across layer-2 boundaries. Multiple instances of Ananta have been deployed in the Windows Azure public cloud with combined bandwidth capacity exceeding 1Tbps. It is serving traffic needs of a diverse set of tenants, including the blob, table and relational storage services. With its scale-out data plane we can easily achieve more than 100Gbps throughput for a single public IP address. In this paper, we describe the requirements of a cloud-scale load balancer, the design of Ananta and lessons learnt from its implementation and operation in the Windows Azure public cloud. Parveen Patel, Deepak Bansal, Ashwin Murthy, Albert G. Greenberg, David A. Maltz, Randy Kern, Marios Zikos, Changhoon Kim, Naveen Karri |
SIGCOMM | 6 |
| 2013 | Challenges in cloud scale data centersabstractData centers are fascinating places, where the massive scale required to deliver on-line services like web search and cloud hosting turns minor issues into major challenges that must be addressed in the design of the physical infrastructure and the software platform. In this talk, I'll briefly overview the kinds of applications that run in mega-data centers and the workloads they place on the infrastructure. I'll then describe a number of challenges seen in Microsoft's data centers, with the goals of posing questions more than describing solutions and explaining how economic factors, technology issues, and software design interact when creating low-latency, low-cost, high availability services. David A. Maltz |
SIGMETRICS | 1 |
| 2013 | Dynamic Request Splitting for Interactive Cloud ApplicationsabstractDeploying interactive applications in the cloud is a challenge due to the high variability in performance of cloud services. In this paper, we present Dealer — a system that helps geo-distributed, interactive and multi-tier applications meet their stringent requirements on response time despite such variability. Our approach is motivated by the fact that, at any time, only a small number of application components of large multi-tier applications experience poor performance. Dealer continually monitors the performance of individual components and communication latencies between them to build a global view of the application. In serving any given request, Dealer seeks to minimize user response times by picking the best combination of replicas (potentially located across different data centers). While Dealer requires modifications to application code, we show the changes required are modest. Our evaluations on two multi-tier applications using real cloud deployments indicate the 90%ile of response times could be reduced by more than a factor of 6 under natural cloud dynamics. Our results indicate the cost of inter-data-center traffic with Dealer is minor, and that Dealer can in fact be used to reduce the overall operational costs of applications by up to 15% by leveraging the difference in billing plans of cloud instances. Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai |
IEEE J. Sel. Areas Commun. | 3 |
| 2012 | Dealer: application-aware request splitting for interactive cloud applicationsabstractDeploying interactive applications in the cloud is a challenge due to the high variability in performance of cloud services. In this paper, we present Dealer-- a system that helps geo-distributed, interactive and multi-tier applications meet their stringent requirements on response time despite such variability. Our approach is motivated by the fact that, at any time, only a small number of application components of large multi-tier applications experience poor performance. Dealer abstracts application structure as a component graph, with nodes being application components and edges capturing inter-component communication patterns. Dealer continually monitors the performance of individual component replicas and communication latencies between replica pairs. In serving any given user request, Dealer seeks to minimize user response times by picking the best combination of replicas (potentially located across different data-centers). While Dealer does require modifications to application code, we show through integration with two multi-tier applications that the changes required are modest. Our evaluations on two multi-tier applications using real cloud deployments indicate the 90%ile of application response times could be reduced by a factor of 3 under natural cloud dynamics compared to conventional data-center redirection techniques which are agnostic of application structure. Mohammad Y. Hajjat, Shankaranarayanan Puzhavakath Narayanan, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai |
CoNEXT | 3 |
| 2012 | Surviving failures in bandwidth-constrained datacentersabstractDatacenter networks have been designed to tolerate failures of network equipment and provide sufficient bandwidth. In practice, however, failures and maintenance of networking and power equipment often make tens to thousands of servers unavailable, and network congestion can increase service latency. Unfortunately, there exists an inherent tradeoff between achieving high fault tolerance and reducing bandwidth usage in network core; spreading servers across fault domains improves fault tolerance, but requires additional bandwidth, while deploying servers together reduces bandwidth usage, but also decreases fault tolerance. We present a detailed analysis of a large-scale Web application and its communication patterns. Based on that, we propose and evaluate a novel optimization framework that achieves both high fault tolerance and significantly reduces bandwidth usage in the network core by exploiting the skewness in the observed communication patterns. Peter Bodík, Ishai Menache, Mosharaf Chowdhury, Pradeepkumar Mani, David A. Maltz, Ion Stoica |
SIGCOMM | 5 |
| 2012 | NetPilot: automating datacenter network failure mitigationabstractDriven by the soaring demands for always-on and fast-response online services, modern datacenter networks have recently undergone tremendous growth. These networks often rely on commodity hardware to reach immense scale while keeping capital expenses under check. The downside is that commodity devices are prone to failures, raising a formidable challenge for network operators to promptly handle these failures with minimal disruptions to the hosted services. Daniel Turner, Chao-Chih Chen, David A. Maltz, Xiaowei Yang 0001, Ming Zhang 0005 |
SIGCOMM | 4 |
| 2011 | Latency inflation with MPLS-based traffic engineeringabstractWhile MPLS has been extensively deployed in recent years, little is known about its behavior in practice. We examine the performance of MPLS in Microsoft's online service network (MSN), a well-provisioned multi-continent production network connecting tens of data centers. Using detailed traces collected over a 2-month period, we find that many paths experience significantly inflated latencies. We correlate occurrences of latency inflation with routers, links, and DC-pairs. This analysis sheds light on the causes of latency inflation and suggests several avenues for alleviating the problem. Abhinav Pathak, Ming Zhang 0005, Y. Charlie Hu, Ratul Mahajan, David A. Maltz |
Internet Measurement Conference | 5 |
| 2011 | Public DNS system and Global Traffic ManagementabstractCloud service providers operate data centers around the world, and they depend on Global Traffic Management systems to direct requests from clients to the most appropriate data center to serve the requests. While GTM systems have been in-use for years, they are attracting re-newed interests due to the rapid expansion of cloud service providers' networks, the introduction of public DNS systems, as well as new proposals to alter how they should work and what information local DNS servers (LDNS) should make available to drive the GTM systems. This paper uses large-scale measurements conducted from more than 5M clients to establish properties of the current Internet that affect the design of the GTM systems, such as the stretch between a client's actual position and its LDNS from GTM's perspective, the impact of public DNS systems, and the granularity at which GTM decisions should be made. The results can inform the debate over how GTM systems should be designed. Cheng Huang 0002, David A. Maltz, Jin Li 0001, Albert G. Greenberg |
INFOCOM | 2 |
| 2011 | Profiling Network Performance for Multi-tier Data Center Applications
Minlan Yu, Albert G. Greenberg, David A. Maltz, Jennifer Rexford, Srikanth Kandula, Changhoon Kim |
NSDI | 3 |
| 2011 | Towards systematic design of enterprise networksabstractEnterprise networks are important, with size and complexity even surpassing carrier networks. Yet, the design of enterprise networks remains ad hoc and poorly understood. In this paper, we show how a systematic design approach can handle two key areas of enterprise design: virtual local area networks (VLANs) and reachability control. We focus on these tasks given their complexity, prevalence, and time-consuming nature. Our contributions are threefold. First, we show how these design tasks may be formulated in terms of network-wide performance, security, and resilience requirements. Our formulations capture the correctness and feasibility constraints on the design, and they model each task as one of optimizing desired criteria subject to the constraints. The optimization criteria may further be customized to meet operator-preferred design strategies. Second, we develop a set of algorithms to solve the problems that we formulate. Third, we demonstrate the feasibility and value of our systematic design approach through validation on a large-scale campus network with hundreds of routers and VLANs. Yu-Wei Eric Sung, Xin Sun 0002, Sanjay G. Rao, Geoffrey G. Xie, David A. Maltz |
IEEE/ACM Trans. Netw. | 5 |
| 2010 | Network traffic characteristics of data centers in the wildabstractAlthough there is tremendous interest in designing improved networks for data centers, very little is known about the network-level traffic characteristics of data centers today. In this paper, we conduct an empirical study of the network traffic in 10 data centers belonging to three different categories, including university, enterprise campus, and cloud data centers. Our definition of cloud data centers includes not only data centers employed by large online service providers offering Internet-facing applications but also data centers used to host data-intensive (MapReduce style) applications). We collect and analyze SNMP statistics, topology and packet-level traces. We examine the range of applications deployed in these data centers and their placement, the flow-level and packet-level transmission properties of these applications, and their impact on network and link utilizations, congestion and packet drops. We describe the implications of the observed traffic patterns for data center internal traffic engineering as well as for recently proposed architectures for data center networks. Theophilus Benson, Aditya Akella, David A. Maltz |
Internet Measurement Conference | 3 |
| 2010 | Data center TCP (DCTCP)abstractCloud data centers host diverse applications, mixing workloads that require small predictable latency with others requiring large sustained throughput. In this environment, today's state-of-the-art TCP protocol falls short. We present measurements of a 6000 server production cluster and reveal impairments that lead to high application latencies, rooted in TCP's demands on the limited buffer space available in data center switches. For example, bandwidth hungry "background" flows build up queues at the switches, and thus impact the performance of latency sensitive "foreground" traffic. Mohammad Alizadeh, Albert G. Greenberg, David A. Maltz, Jitendra Padhye, Parveen Patel, Balaji Prabhakar, Sudipta Sengupta, Murari Sridharan |
SIGCOMM | 3 |
| 2010 | Cloudward bound: planning for beneficial migration of enterprise applications to the cloudabstractIn this paper, we tackle challenges in migrating enterprise services into hybrid cloud-based deployments, where enterprise operations are partly hosted on-premise and partly in the cloud. Such hybrid architectures enable enterprises to benefit from cloud-based architectures, while honoring application performance requirements, and privacy restrictions on what services may be migrated to the cloud. We make several contributions. First, we highlight the complexity inherent in enterprise applications today in terms of their multi-tiered nature, large number of application components, and interdependencies. Second, we have developed a model to explore the benefits of a hybrid migration approach. Our model takes into account enterprise-specific constraints, cost savings, and increased transaction delays and wide-area communication costs that may result from the migration. Evaluations based on real enterprise applications and Azure-based cloud deployments show the benefits of a hybrid migration approach, and the importance of planning which components to migrate. Third, we shed insight on security policies associated with enterprise applications in data centers. We articulate the importance of ensuring assurable reconfiguration of security policies as enterprise applications are migrated to the cloud. We present algorithms to achieve this goal, and demonstrate their efficacy on realistic migration scenarios. Mohammad Y. Hajjat, Xin Sun 0002, Yu-Wei Eric Sung, David A. Maltz, Sanjay G. Rao, Kunwadee Sripanidkulchai, Mohit Tawarmalani |
SIGCOMM | 4 |
| 2010 | MMS: An autonomic network-layer foundation for network managementabstractNetworks cannot be managed without management plane communications among geographically distributed network devices and control agents. Unfortunately, the mechanisms used in commercial networks to support management plane communications are often hard to configure, insufficiently secured, and/or suboptimal in performance. This paper presents the design and implementation of the Meta-Management System (MMS), a network-layer subsystem that provides robust autonomic support for management plane communications. We demonstrate the practicality of the MMS via a fully functional implementation that runs on commodity hardware, and experimentally show that the MMS is efficient and scalable. The MMS software is freely available. Hemant Gogineni, Albert G. Greenberg, David A. Maltz, T. S. Eugene Ng, Hong Yan 0002, Hui Zhang 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2009 | Mining policies from enterprise network configurationabstractFew studies so far have examined the nature of reachability policies in enterprise networks. A better understanding of reachability policies could both inform future approaches to network design as well as current network configuration mechanisms. In this paper, we introduce the notion of a policy unit, which is an abstract representation of how the policies implemented in a network apply to different network hosts. We develop an approach for reverse-engineering a network's policy units from its router configuration. We apply this approach to the configurations of five productions networks, including three university and two private enterprises. Through our empirical study, we validate that policy units capture useful characteristics of a network's policy. We also obtain insights into the nature of the policies implemented in modern enterprises. For example, we find most hosts in these networks are subject to nearly identical reachability policies at Layer 3. Theophilus Benson, Aditya Akella, David A. Maltz |
Internet Measurement Conference | 3 |
| 2009 | Unraveling the Complexity of Network Management
Theophilus Benson, Aditya Akella, David A. Maltz |
NSDI | 3 |
| 2009 | VL2: a scalable and flexible data center networkabstractTo be agile and cost effective, data centers should allow dynamic resource allocation across large server pools. In particular, the data center network should enable any server to be assigned to any service. To meet these goals, we present VL2, a practical network architecture that scales to support huge data centers with uniform high capacity between servers, performance isolation between services, and Ethernet layer-2 semantics. VL2 uses (1) flat addressing to allow service instances to be placed anywhere in the network, (2) Valiant Load Balancing to spread traffic uniformly across network paths, and (3) end-system based address resolution to scale to large server pools, without introducing complexity to the network control plane. VL2's design is driven by detailed measurements of traffic and fault data from a large operational cloud service provider. VL2's implementation leverages proven network technologies, already available at low cost in high-speed hardware implementations, to build a scalable and reliable network architecture. As a result, VL2 networks can be deployed today, and we have built a working prototype. We evaluate the merits of the VL2 design using measurement, analysis, and experiments. Our VL2 prototype shuffles 2.7 TB of data among 75 servers in 395 seconds - sustaining a rate that is 94% of the maximum possible. Albert G. Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A. Maltz, Parveen Patel, Sudipta Sengupta |
SIGCOMM | 7 |
| 2009 | Structure preserving anonymization of router configuration dataabstractA repository of router configuration files from production networks would provide the research community with a treasure trove of data about network topologies, routing designs, and security policies. However, configuration files have been largely unobtainable precisely because they provide detailed information that could be exploited by competitors and attackers. This paper describes a method for anonymizing router configuration files by removing all information that connects the data to the identity of the underlying network, while still preserving the structure of information that makes the data valuable to networking researchers. Anonymizing configuration files has unusual requirements, including preserving relationships between elements of data, anonymizing regular expressions, and robustly coping with more than 200 versions of the configuration language. Conventional tools and techniques are poorly suited to the problem. Our anonymization method has been validated with a major carrier, earning unprivileged researchers access to the configuration files of thousands of routers in hundreds of networks. Through example analysis, we demonstrate that the anonymized data retains the key properties of the network design. The paper sets out techniques that could be used in an attempt to break the anonymization, and it concludes our anonymization techniques are most applicable to enterprise networks, because the large number of enterprises and the difficulty of probing them from the outside make it hard to recognize an anonymized network based solely on publicly-available information about its topology or configuration. When applied to backbone networks, which are few in number and many of whose properties can be publicly measured, the anonymization might be broken by fingerprinting techniques described in this paper. David A. Maltz, Jibin Zhan, Gísli Hjálmtýsson, Albert G. Greenberg, Jennifer Rexford, Geoffrey G. Xie, Hui Zhang 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2008 | Towards systematic design of enterprise networksabstractEnterprise networks are important, with size and complexity even surpassing carrier networks. Yet, the design of enterprise networks is ad-hoc and poorly understood. In this paper, we show how a systematic design approach can handle two key areas of enterprise design: virtual local area networks (VLANs) and reachability control. We focus on these tasks given their complexity, prevalence, and time-consuming nature. Our contributions are three-fold. First, we show how these design tasks may be formulated in terms of network-wide performance, security, and resilience requirements. Our formulations capture the correctness and feasibility constraints on the design, and they model each task as one of optimizing desired criteria subject to the constraints. The optimization criteria may further be customized to meet operator-preferred design strategies. Second, we develop a set of algorithms to solve the problems that we formulate. Third, we demonstrate the feasibility and value of our systematic design approach through validation on a large-scale campus network with hundreds of routers and VLANs. Yu-Wei Eric Sung, Sanjay G. Rao, Geoffrey G. Xie, David A. Maltz |
CoNEXT | 4 |
| 2007 | Fast Variational Inference for Large-scale Internet DiagnosisabstractWeb servers on the Internet need to maintain high reliability, but the cause of intermittent failures of web transactions is non-obvious. We use Bayesian inference to diagnose problems with web services. This diagnosis problem is far larger than any previously attempted: it requires inference of 10^4 possible faults from 10^5 observations. Further, such inference must be performed in less than a second. Inference can be done at this speed by combining a variational approximation, a mean-field approximation, and the use of stochastic gradient descent to optimize a variational cost function. We use this fast inference to diagnose a time series of anomalous HTTP requests taken from a real web service. The inference is fast enough to analyze network logs with billions of entries in a matter of hours. John C. Platt, Emre Kiciman, David A. Maltz |
NIPS | 3 |
| 2007 | Tesseract: A 4D Network Control Plane
Hong Yan 0002, David A. Maltz, T. S. Eugene Ng, Hemant Gogineni, Hui Zhang 0001, Zheng Cai |
NSDI | 2 |
| 2007 | Towards highly reliable enterprise network services via inference of multi-level dependenciesabstractLocalizing the sources of performance problems in large enterprise networks is extremely challenging. Dependencies are numerous, complex and inherently multi-level, spanning hardware and software components across the network and the computing infrastructure. To exploit these dependencies for fast, accurate problem localization, we introduce an Inference Graph model, which is well-adapted to user-perceptible problems rooted in conditions giving rise to both partial service degradation and hard faults. Further, we introduce the Sherlock system to discover Inference Graphs in the operational enterprise, infer critical attributes, and then leverage the result to automatically detect and localize problems. To illuminate strengths and limitations of the approach, we provide results from a prototype deployment in a large enterprise network, as well as from testbed emulations and simulations. In particular, we find that taking into account multi-level structure leads to a 30% improvement in fault localization, as compared to two-level approaches. Paramvir Bahl, Ranveer Chandra, Albert G. Greenberg, Srikanth Kandula, David A. Maltz, Ming Zhang 0005 |
SIGCOMM | 5 |
| 2006 | Discovering Dependencies for Network Management
Paramvir Bahl, Paul Barham 0001, Richard Black, Ranveer Chandra, Moisés Goldszmidt, Rebecca Isaacs, Srikanth Kandula, John MacCormick, David A. Maltz, Richard Mortier, Michal Wawrzoniak, Ming Zhang 0005 |
HotNets | 10 |
| 2005 | On static reachability analysis of IP networksabstractThe primary purpose of a network is to provide reachability between applications running on end hosts. In this paper, we describe how to compute the reachability a network provides from a snapshot of the configuration state from each of the routers. Our primary contribution is the precise definition of the potential reachability of a network and a substantial simplification of the problem through a unified modeling of packet filters and routing protocols. In the end, we reduce a complex, important practical problem to computing the transitive closure to set union and intersection operations on reachability set representations. We then extend our algorithm to model the influence of packet transformations (e.g., by NATs or ToS remapping) along the path. Our technique for static analysis of network reachability is valuable for verifying the intent of the network designer, troubleshooting reachability problems, and performing "what-if" analysis of failure scenarios. Geoffrey G. Xie, Jibin Zhan, David A. Maltz, Hui Zhang 0001, Albert G. Greenberg, Gísli Hjálmtýsson, Jennifer Rexford |
INFOCOM | 3 |
| 2005 | Worm Origin Identification Using Random MoonwalksabstractWe propose a novel technique that can determine both the host responsible for originating a propagating worm attack and the set of attack flows that make up the initial stages of the attack tree via which the worm infected successive generations of victims. We argue that knowledge of both is important for combating worms: knowledge of the origin supports law enforcement, and knowledge of the causal flows that advance the attack supports diagnosis of how network defenses were breached. Our technique exploits the "wide tree" shape of a worm propagation emanating from the source by performing random "moonwalks" backward in time along paths of flows. Correlating the repeated walks reveals the initial causal flows, thereby aiding in identifying the source. Using analysis, simulation, and experiments with real world traces, we show how the technique works against both today's fast propagating worms and stealthy worms that attempt to hide their attack flows among background traffic. Yinglian Xie, Vyas Sekar, David A. Maltz, Michael K. Reiter, Hui Zhang 0001 |
S&P | 3 |
| 2004 | Structure preserving anonymization of router configuration dataabstractA repository of router configuration files from production networks would provide the research community with a treasure trove of data about network topologies, routing designs, and security policies. However, configuration files have been largely unobtainable precisely because they provide detailed information that could be exploited by competitors and attackers. This paper describes a method for anonymizing router configuration files by removing all information that connects the data to the identity of the originating network, while still preserving the structure of information that makes the data valuable to networking researchers. Anonymizing configuration files has unusual requirements, including preserving relationships between elements of data, anonymizing regular expressions, and robustly coping with more than 200 versions of the configuration language, that mean conventional tools and techniques are poorly suited to the problem. Our anonymization method has been validated with a major carrier, earning unprivileged researchers access to the configuration files of more than 7600 routers in 31 networks. Through example analysis, we demonstrate that the anonymized data retains the key properties of the network design. We believe that applying our single-blind methodology to a large number of production networks from different sources would be of tremendous value to both the research and operations communities. David A. Maltz, Jibin Zhan, Geoffrey G. Xie, Hui Zhang 0001, Gísli Hjálmtýsson, Albert G. Greenberg, Jennifer Rexford |
Internet Measurement Conference | 1 |
| 2004 | Routing design in operational networks: a look from the insideabstractIn any IP network, routing protocols provide the intelligence that takes a collection of physical links and transforms them into a network that enables packets to travel from one host to another. Though routing design is arguably the single most important design task for large IP networks, there has been very little systematic investigation into how routing protocols are actually used in production networks to implement the goals of network architects. We have developed a methodology for reverse engineering a coherent global view of a network's routing design from the static analysis of dumps of the local configuration state of each router. Starting with a set of 8,035 configuration files, we have applied this method to 31 production networks. In this paper we present a detailed examination of how routing protocols are used in operational networks. In particular, the results show the conventional model of interior and exterior gateway protocols is insufficient to describe the diverse set of mechanisms used by architects, and we provide examples of the more unusual designs and examine their trade-offs. We discuss the strengths and weaknesses of our methodology, and argue that it opens paths towards new understandings of network behavior and design. Geoffrey G. Xie, Jibin Zhan, David A. Maltz, Hui Zhang 0001, Albert G. Greenberg, Gísli Hjálmtýsson |
SIGCOMM | 3 |
| 2002 | MSOCKS+: an architecture for transport layer mobility
Pravin Bhagwat, David A. Maltz, Adrian Segall |
Comput. Networks | 2 |
| 2000 | Quantitative lessons from a full-scale multi-hop wireless ad hoc network testbedabstractThis paper presents preliminary quantitative results from data collected during runs of our multi-hop wireless ad hoc network testbed. The network successfully carried a composite workload including voice, bulk data, and real-time data. Careful analysis of recorded runs highlights radio propagation issues that network protocols will need to address in the future. David A. Maltz, Josh Broch, David B. Johnson 0001 |
WCNC | 1 |
| 1999 | The effects of on-demand behavior in routing protocols for multihop wireless ad hoc networksabstractA number of different routing protocols proposed for use in multihop wireless ad hoc networks are based in whole or in part on what can be described as on-demand behavior. By on-demand behavior, we mean approaches based only on reaction to the offered traffic being handled by the routing protocol. In this paper, we analyze the use of on-demand behavior in such protocols, focusing on its effect on the routing protocol's forwarding latency, overhead cost, and route caching correctness, drawing examples from detailed simulation of the dynamic source routing (DSR) protocol. We study the protocol's behavior and the changes introduced by variations on some of the mechanisms that make up the protocol, examining which mechanisms have the greatest impact and exploring the tradeoffs that exist between them. David A. Maltz, Josh Broch, Jorjeta G. Jetcheva, David B. Johnson 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 1998 | MSOCKS: An Architecture for Transport Layer MobilityabstractMobile nodes of the future will be equiped with multiple network interfaces to take advantage of overlay networks, yet no current mobility systems provide full support for the simultaneous use of multiple interfaces. The need for such support arises when multiple connectivity options are available with different cost, coverage, latency and bandwidth characteristics, and applications want their data to flow over the interface that best matches the characteristics of the data. We present an architecture called transport layer mobility that allows mobile nodes to not only change their point of attachment to the Internet, but also to control which network interfaces are used for the different kinds of data leaving from and arriving at the mobile node. We implement our transport layer mobility scheme using a split connection proxy architecture and a new technique called TCP splice that gives split connection proxy systems the same end to-end semantics as normal TCP connections. David A. Maltz, Pravin Bhagwat |
INFOCOM | 1 |
| 1998 | A Performance Comparison of Multi-Hop Wireless Ad Hoc Network Routing ProtocolsabstractAn ad hoc networkis a collwtion of wirelessmobile nodes dynamically forminga temporarynetworkwithouttheuse of anyexistingnetworkirrfrastructureor centralizedadministration.Dueto the limitedtransmissionrange of ~vlreless nenvorkinterfaces,multiplenetwork"hops"maybe neededfor onenodeto exchangedata ivithanotheracrox the network.In recentyears, a ttiery of nelvroutingprotocols~geted specificallyat this environment have been developed.butlittle pcrfomrartwinformationon mch protocol and no ralistic performancecomparisonbehvwrrthem ISavailable.~Is paper presentsthe results of a derailedpacket-levelsimulationcomparing fourmulti-hopwirelessad hoc networkroutingprotocolsthatcovera range of design choices: DSDV,TORA, DSR and AODV.\Vehave extended the /~r-2networksimulatorto accuratelymodelthe MACand physical-layer behaviorof the IEEE 802.1I wirelessLAN standard,includinga realistic wtrelesstransmissionchannelmodel, and present the resultsof simulations of net(vorksof 50 mobilenodes.This work was supported in Josh Broch, David A. Maltz, David B. Johnson 0001, Yih-Chun Hu, Jorjeta G. Jetcheva |
MobiCom | 2 |
| 1995 | Pointing the Way: Active Collaborative FilteringabstractCollaborative filtering is based on the premise that people looking for information should be able to make use of what others have already found and evaluated.Current collaborative filtoing systems provide tools for readers to filter documents based on aggregated ratings over a changing group of readers.Motivated by the results of a study of information sharing, we describe a different type of collaborative filtering system in which people who find interesting documents actively send "pointers" to those documents to their colleagues.A "pointer" contains a hypertext link to the source document as well as contextual information to help the recipient determine the interest and relevance of the document prior to accessing it.Preliminary data suggest that people are using the system in anticipated and unanticipated ways, as well as creating information "digests". David A. Maltz, Kate Ehrlich |
CHI | 1 |