Aman Shaikh

dblp:s/AmanShaikh · DBLP profile ↗
← Back
34ranked-venue papers
7as first author
1since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 26 · 7 first-author · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 2Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
25 papers
Routing and switching · 38% Network management and operations · 29% Network measurement and analytics · 16%
Databases, data mining, and information retrieval
2 papers
Graph data management · 83% Spatial and temporal data management · 17%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network management and operations › performance management
performance diagnosis
0.522020
Microscope: Queue-based Performance Diagnosis for Network Functions · SIGCOMM 2020
Towards automated performance diagnosis in a large IPTV network · SIGCOMM 2009
Routing and switching
inter-domain routing
0.562011
Address-based route reflection · CoNEXT 2011
Impact of prefix-match changes on IP reachability · Internet Measurement Conference 2009
Impact of hot-potato routing changes in IP networks · IEEE/ACM Trans. Netw. 2008
Software-defined and programmable networks
network function virtualization
0.412020
Microscope: Queue-based Performance Diagnosis for Network Functions · SIGCOMM 2020
Graph data management
graph database
0.422018
A Graph Database for a Virtualized Network Infrastructure · SIGMOD Conference 2018
Virtualized Network Service Topology Exploration Using Nepal · SIGMOD Conference 2017
Network measurement and analytics
topology discovery
0.422018
Virtualized Network Service Topology Exploration Using Nepal · SIGMOD Conference 2017
A Graph Database for a Virtualized Network Infrastructure · SIGMOD Conference 2018
Network management and operations › fault management
fault diagnosis
0.442011
Demystifying configuration challenges and trade-offs in network-based ISP services · SIGCOMM 2011
Towards automated performance diagnosis in a large IPTV network · SIGCOMM 2009
Troubleshooting chronic conditions in large IP networks · CoNEXT 2008
Routing and switching › inter-domain routing
BGP
0.342011
Address-based route reflection · CoNEXT 2011
Where's Waldo? practical searches for stability in iBGP · ICNP 2008
Wresting Control from BGP: Scalable Fine-Grained Route Control · USENIX ATC 2007
Network management and operations
cellular network management
0.212016
Libra: Impact assessment of cellular load balancing · INFOCOM 2016
Routing and switching › routing protocol
OSPF
0.242006
Avoiding instability during graceful shutdown of multiple OSPF routers · IEEE/ACM Trans. Netw. 2006
OSPF Monitoring: Architecture, Design, and Deployment Experience · NSDI 2004
An Efficient Algorithm for OSPF Subnet Aggregation · ICNP 2003
Routing and switching › adaptive routing
deflection routing
0.232008
Impact of hot-potato routing changes in IP networks · IEEE/ACM Trans. Netw. 2008
Dynamics of hot-potato routing in IP networks · SIGMETRICS 2004
Network sensitivity to hot-potato disruptions · SIGCOMM 2004
Routing and switching
forwarding table
0.222011
SMALTA: practical and near-optimal FIB aggregation · CoNEXT 2011
Avoiding Instability during Graceful Shutdown of OSPF · INFOCOM 2002
Routing and switching
routing protocol
0.132006
Avoiding instability during graceful shutdown of multiple OSPF routers · IEEE/ACM Trans. Netw. 2006
An Efficient Algorithm for OSPF Subnet Aggregation · ICNP 2003
An OSPF topology server: design and evaluation · IEEE J. Sel. Areas Commun. 2002
Processor architecture and microarchitecture
resource contention
0.112020
Microscope: Queue-based Performance Diagnosis for Network Functions · SIGCOMM 2020
Network management and operations
configuration management
0.112011
Demystifying configuration challenges and trade-offs in network-based ISP services · SIGCOMM 2011
Routing and switching › forwarding table
FIB aggregation
0.112011
SMALTA: practical and near-optimal FIB aggregation · CoNEXT 2011
Routing and switching
route reflection
0.112011
Address-based route reflection · CoNEXT 2011
Routing and switching › routing protocol
intra-domain routing
0.142006
Dynamics of hot-potato routing in IP networks · SIGMETRICS 2004
A case study of OSPF behavior in a large enterprise network · Internet Measurement Workshop 2002
Placing Relay Nodes for Intra-Domain Path Diversity · INFOCOM 2006
Network measurement and analytics › internet measurement › routing measurement
BGP prefix reachability
0.112009
Impact of prefix-match changes on IP reachability · Internet Measurement Conference 2009
Network measurement and analytics › internet measurement › routing measurement
internet routing measurement
0.112009
Impact of prefix-match changes on IP reachability · Internet Measurement Conference 2009
Routing and switching › IP lookup
longest prefix matching
0.112009
Impact of prefix-match changes on IP reachability · Internet Measurement Conference 2009
Spatial and temporal data management › temporal query processing
time-travel query
0.112017
Virtualized Network Service Topology Exploration Using Nepal · SIGMOD Conference 2017
Network management and operations
network monitoring
0.122004
OSPF Monitoring: Architecture, Design, and Deployment Experience · NSDI 2004
An OSPF topology server: design and evaluation · IEEE J. Sel. Areas Commun. 2002
Network management and operations › network monitoring
routing oscillation detection
0.112008
Where's Waldo? practical searches for stability in iBGP · ICNP 2008
Network management and operations
network robustness
0.122006
Placing Relay Nodes for Intra-Domain Path Diversity · INFOCOM 2006
Network sensitivity to hot-potato disruptions · SIGCOMM 2004
Internet architecture and protocols › overlay networks
overlay routing
0.112006
Placing Relay Nodes for Intra-Domain Path Diversity · INFOCOM 2006
Routing and switching › multipath routing
path diversity
0.112006
Placing Relay Nodes for Intra-Domain Path Diversity · INFOCOM 2006
Internet of things and sensor networks › topology control
relay node placement
0.112006
Placing Relay Nodes for Intra-Domain Path Diversity · INFOCOM 2006
Routing and switching
routing dynamics
0.012004
Dynamics of hot-potato routing in IP networks · SIGMETRICS 2004
Routing and switching
routing
0.012011
SMALTA: practical and near-optimal FIB aggregation · CoNEXT 2011
Network measurement and analytics › flow monitoring
netflow analysis
0.012009
Impact of prefix-match changes on IP reachability · Internet Measurement Conference 2009

Methods — techniques the papers use, named apart from their topics

queuing analysis · 0.9load balancing metrics · 0.2measurement · 0.2topology-based route reflection · 0.1configuration analysis · 0.1statistical data mining · 0.1online algorithm · 0.1multi-resolution analysis · 0.1BGP update analysis · 0.1intermittent event analysis · 0.1
YearPublicationVenuePosition
2026 GGN: Experiences in Designing and Deploying the Next-Generation Google Global Network
abstract
Cloud and AI/ML workloads are posing unprecedented new requirements on the wide-area network: it must combine strict availability, massive growth, and feature agility. It became increasingly clear that traditional WAN designs were ill-equipped to adapt to these requirements.
Mohammad Al-Fares, Richard Alimi, Arda Balkanay, Dennis Fetterly, Chi-Yao Hong, Nachikethas A. Jagadeesan, Bikash Koley, Priya Mahadevan, Subhasree Mandal, Warren Martins, Arjun Muralidharan, Namrata Pralhad Kadam, Aman Shaikh, Anees Shaikh, Rob Shakir, Sankalp Singh, Charith Wickramaarachchi, Jonathan Zolla
SIGCOMM13
2020 Microscope: Queue-based Performance Diagnosis for Network Functions
abstract
By moving monolithic network appliances to software running on commodity hardware, network function virtualization allows flexible resource sharing among network functions and achieves scalability with low cost. However, due to resource contention, network functions can suffer from performance problems that are hard to diagnose. In particular, when many flows traverse a complex topology of NF instances, it is hard to pinpoint root causes for a flow experiencing performance issues such as low throughput or high latency. Simply maintaining resource counters at individual NFs is not sufficient since the effect of resource contention can propagate across NFs and over time. In this paper, we introduce Microscope, a performance diagnosis tool, for network functions that leverages queuing information at NFs to identify the root causes (i.e., resources, NFs, traffic patterns of flows etc.). Our evaluation on realistic NF chains and traffic shows that we can correctly capture root causes behind 89.7% of performance impairments, up to 2.5 times more than the state-of-the-art tools with low overhead.
Junzhi Gong, Muhammad Bilal Anwer, Aman Shaikh, Minlan Yu
SIGCOMM4
2018 A Graph Database for a Virtualized Network Infrastructure
abstract
Modern communication networks are large, dynamic, complex, and increasingly use virtualized network infrastructure. To deploy, maintain, and troubleshoot such networks, it is essential to understand how network elements - such as servers, switches, virtual machines, and virtual network functions - are connected to one another, and to be able to discover communication paths between them. For network maintenance applications such as troubleshooting and service quality management, it is also essential to understand how connections change over time, and be able to pose time-travel queries to retrieve information about past network states. With the industry-wide move to Software Defined Networks and Virtualized Network Functions (VNFs) [26][24], maintaining these inventory and topology databases becomes a critical issue.
Pramod A. Jamkhedkar, Theodore Johnson, Yaron Kanza, Aman Shaikh, N. K. Shankaranarayanan, Vladislav Shkapenyuk
SIGMOD Conference4
2017 Virtualized Network Service Topology Exploration Using Nepal
abstract
Modern communication networks are large, dynamic, and complex. To deploy, maintain, and troubleshoot such networks, it is essential to understand how network elements such as servers, switches, virtual machines, and virtual network functions are connected to one another, and to be able to discover communication paths between them. For network maintenance applications such as troubleshooting and service quality management it is also essential to understand how connections change over time, and be able to pose time-travel queries to retrieve information about past network states. With the industry-wide move to SDNs and virtualized network functions [13], maintaining these inventory databases becomes a critical issue.
Pramod A. Jamkhedkar, Theodore Johnson, Yaron Kanza, Aman Shaikh, N. K. Shankaranarayanan, Vladislav Shkapenyuk, Gordon Woodhull
SIGMOD Conference4
2016 Libra: Impact assessment of cellular load balancing
abstract
Load on cellular towers is one of the key metrics that cellular service providers monitor as part of their operational and management tasks. Increased load on the towers can lead to congestion, which in turn can severely degrade quality of service perceived by users. Hence, it is of great interest to cellular service providers to minimize the maximum load at cell towers, and thereby minimize chances of congestion in the event of a sudden increase in load due to user demand changes. This goal can be achieved by proactive load balancing among neighboring cell towers, i.e., proactively identify opportunities to balance the load through re-binding of users from heavily loaded cell towers to lightly loaded neighboring towers. In this paper, we propose a new tool Libra to effectively assess the impact of load balancing related parameter changes. Libra provides an objective measure of the degree of load imbalance across multiple network locations and identifies if the measure improves or degrades after parameter changes. Our evaluation of Libra using real-world data collected from a large cellular provider demonstrates its effectiveness in accurately capturing the degree of imbalance at multiple cell towers.
Kanthi Nagaraj, Ajay Mahimkar, Zihui Ge, Aman Shaikh, Jia Wang 0001, Kevin Mohr, Mark Stockert
INFOCOM4
2012 Path inference in data center networks
Kyriaki Levanti, Vijay Gopalakrishnan, Hyong S. Kim 0001, Seungjoon Lee, Emmanuil Mavrogiorgis, Aman Shaikh
CNSM6
2012 Practical Network-Wide Compression of IP Routing Tables
abstract
The memory Internet routers use to store paths to destinations is expensive, andmustbecontinuallyupgradedinthefaceofsteadilyincreasingrouting table size. Unfortunately, routing protocols are not designed to gracefully handle cases where memory becomes full, which arises increasingly often due to misconfigurations and routing table growth. Hence router memory must typically be heavily overprovisioned by network operators, inflating operating costs and administrative effort. The research community has primarily focused on clean-slate solutions that cannot interoperate with the deployed base of protocols. This paper presents an incrementally-deployable Memory Management System (MMS) that reduces associated router state by up to 70%. The MMS coalesces prefixes to reduce memory consumption and can be deployed locally on each router or centrally on a route server. The system can operate transparently, without requiring changes in other ASes. Our memory manager can extend router lifetimes up to seven years, given current prefix growth trends. 1.
Elliott Karpilovsky, Matthew Caesar 0001, Jennifer Rexford, Aman Shaikh, Jacobus E. van der Merwe
IEEE Trans. Netw. Serv. Manag.4
2011 Address-based route reflection
abstract
BGP Route Reflectors (RR), which are commonly used to help scale Internal BGP (iBGP), can produce oscillations, forwarding loops, and path inefficiencies. ISPs avoid these pitfalls through careful topology design, RR placement, and link-metric assignment. This paper presents Address-Based Route Reflection (ABRR): the first iBGP solution that completely solves all oscillation and looping problems, has no path inefficiencies, and puts no constraints on RR placement. ABRR does this by emulating the semantics of full-mesh iBGP, and thereby adopting the correctness and path efficiency properties of full-mesh iBGP. Both traditional Topology-Based Route Reflection (TBRR) and ABRR take a divide-and-conquer approach. While TBRR scales by making each RR responsible for all prefixes from some fraction of routers, ABRR scales by making each RR responsible for some fraction of prefixes from all routers. We have implemented a fully functional ABRR prototype. Using BGP data from a Tier-1 ISP, our analytical and implementation results show that ABRR's scaling and convergence properties compare positively with traditional TBRR.
Ruichuan Chen, Aman Shaikh, Jia Wang 0001, Paul Francis
CoNEXT2
2011 SMALTA: practical and near-optimal FIB aggregation
abstract
IP Routers use sophisticated forwarding table (FIB) lookup algorithms that minimize lookup time, storage, and update time. This paper presents SMALTA, a practical, near-optimal FIB aggregation scheme that shrinks forwarding table size without modifying routing semantics or the external behavior of routers, and without requiring changes to FIB lookup algorithms and associated hardware and software. On typical IP routers using the FIB lookup algorithm Tree Bitmap, SMALTA shrinks FIB storage by at least 50%, representing roughly four years of routing table growth at current rates. SMALTA also reduces average lookup time by 25% for a uniform traffic matrix. Besides the benefits this brings to future routers, SMALTA provides a critical easy-to-deploy one-time benefit to the installed base should IPv4 address depletion result in increased routing table growth rate. The effective cost of this improvement is a sub-second delay in inserting updates into the FIB once every few hours. We describe SMALTA, prove its correctness, measure its performance using data from a Tier-1 provider as well as Route-Views. We also describe an implementation in Quagga that demonstrates its ease of implementation.
Zartash Afzal Uzmi, Markus E. Nebel, Ahsan Tariq, Sana Jawad, Ruichuan Chen, Aman Shaikh, Jia Wang 0001, Paul Francis
CoNEXT6
2011 Demystifying configuration challenges and trade-offs in network-based ISP services
abstract
ISPs are increasingly offering a variety of network-based services such as VPN, VPLS, VoIP, Virtual-Wire and DDoS protection. Although both enterprise and residential networks are rapidly adopting these services, there is little systematic work on the design challenges and trade-offs ISPs face in providing them. The goal of our paper is to understand the complexity underlying the layer-3 design of services and to highlight potential factors that hinder their introduction, evolution and management. Using daily snapshots of configuration and device metadata collected from a tier-1 ISP, we examine the logical dependencies and special cases in device configurations for five different network-based services. We find: (1) the design of the core data-plane is usually service-agnostic and simple, but the control-planes for different services become more complex as services evolve; (2) more crucially, the configuration at the service edge inevitably becomes more complex over time, potentially hindering key management issues such as service upgrades and troubleshooting; and (3) there are key service-specific issues that also contribute significantly to the overall design complexity. Thus, the high prevalent complexity could impede the adoption and growth of network-based services. We show initial evidence that some of the complexity can be mitigated systematically.
Theophilus Benson, Aditya Akella, Aman Shaikh
SIGCOMM3
2010 Detecting the performance impact of upgrades in large operational networks
abstract
Networks continue to change to support new applications, improve reliability and performance and reduce the operational cost. The changes are made to the network in the form of upgrades such as software or hardware upgrades, new network or service features and network configuration changes. It is crucial to monitor the network when upgrades are made because they can have a significant impact on network performance and if not monitored may lead to unexpected consequences in operational networks. This can be achieved manually for a small number of devices, but does not scale to large networks with hundreds or thousands of routers and extremely large number of different upgrades made on a regular basis.
Ajay Mahimkar, Han Hee Song, Zihui Ge, Aman Shaikh, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Joanne Emmons
SIGCOMM4
2010 BGP route prediction within ISPs
Ashley Flavel, Jeremy McMahon, Aman Shaikh, Matthew Roughan, Nigel G. Bean
Comput. Commun.3
2009 Analyzing the Hold Time Schemes to Limit the Routing Table Calculations in OSPF Protocol
abstract
OSPF is a popular interior gateway routing protocol. Commercial OSPF routers limit their processing load by using a hold time between successive routing table calculations as new link state advertisements (LSAs) arrive following a topology change. A large hold time value limits the frequency of routing table calculations but also causes large delays in convergence to the topology change. Hence, commercial routers now use an exponential back off scheme, where the hold time is initially set to a small value that is expected to rapidly increase, and hence limit the frequency of routing table calculations, in face of continuous LSA arrivals. In this paper, we analyze the ability of different hold time schemes to limit the frequency of routing table calculations under continuous LSA arrivals starting with a small value for the hold time. This analysis is performed using Markov regenerative process based stochastic models as well as simulations using an extensively modified OSPFD simulator.
Mukul Goyal, Mohd Soperi, Seyed Hossein Hosseini 0001, Kishor S. Trivedi, Aman Shaikh, G. Choudhury
AINA5
2009 Impact of prefix-match changes on IP reachability
abstract
Although most studies of Internet routing treat each IP address block (or prefix) independently, the relationship between prefixes is important because routers ultimately forward packets based on the "longest-matching prefix." In fact, the most-specific prefix for a given destination address may change over time, as BGP routes are announced and withdrawn. Even if the most-specific route is withdrawn, routers may still be able to deliver packets to the destination using a less-specific route. In this paper, we analyze BGP update messages and Netflow traffic traces from a large ISP to characterize both the changes to the longest-matching prefix over time and the resulting effects on end-to-end reachability of the destination hosts. To drive our analysis, we design and implement an efficient online algorithm for tracking changes in the longest-matching prefix for each IP address. We analyze the BGP message traces to identify the reasons for prefix-match changes, including failures, route flapping, sub-prefix hijacking, and load-balancing policies. Our preliminary analysis of the Netflow data suggests that the relationship between BGP updates and IP reachability is sometimes counterintuitive.
Jennifer Rexford, Subhabrata Sen, Aman Shaikh
Internet Measurement Conference4
2009 Humpty Dumpty: Putting iBGP Back Together Again
Ashley Flavel, Jeremy McMahon, Aman Shaikh, Matthew Roughan, Nigel G. Bean
Networking3
2009 Quantifying the Extent of IPv6 Deployment
Elliott Karpilovsky, Alexandre Gerber, Dan Pei, Jennifer Rexford, Aman Shaikh
PAM5
2009 Towards automated performance diagnosis in a large IPTV network
abstract
IPTV is increasingly being deployed and offered as a commercial service to residential broadband customers. Compared with traditional ISP networks, an IPTV distribution network (i) typically adopts a hierarchical instead of mesh-like structure, (ii) imposes more stringent requirements on both reliability and performance, (iii) has different distribution protocols (which make heavy use of IP multicast) and traffic patterns, and (iv) faces more serious scalability challenges in managing millions of network elements. These unique characteristics impose tremendous challenges in the effective management of IPTV network and service. In this paper, we focus on characterizing and troubleshooting performance issues in one of the largest IPTV networks in North America. We collect a large amount of measurement data from a wide range of sources, including device usage and error logs, user activity logs, video quality alarms, and customer trouble tickets. We develop a novel diagnosis tool called Giza that is specifically tailored to the enormous scale and hierarchical structure of the IPTV network. Giza applies multi-resolution data analysis to quickly detect and localize regions in the IPTV distribution hierarchy that are experiencing serious performance problems. Giza then uses several statistical data mining techniques to troubleshoot the identified problems and diagnose their root causes. Validation against operational experiences demonstrates the effectiveness of Giza in detecting important performance issues and identifying interesting dependencies. The methodology and algorithms in Giza promise to be of great use in IPTV network operations.
Ajay Mahimkar, Zihui Ge, Aman Shaikh, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Qi Zhao 0006
SIGCOMM3
2009 Efficient and scalable provisioning of always-on multicast streaming services
Meeyoung Cha, W. Art Chaovalitwongse, Jennifer Yates, Aman Shaikh, Sue B. Moon
Comput. Networks4
2008 Troubleshooting chronic conditions in large IP networks
abstract
Chronic network conditions are caused by performance impairing events that occur intermittently over an extended period of time. Such conditions can cause repeated performance degradation to customers, and sometimes can even turn into serious hard failures. It is therefore critical to troubleshoot and repair chronic network conditions in a timely fashion in order to ensure high reliability and performance in large IP networks. Today, troubleshooting chronic conditions is often performed manually, making it a tedious, time-consuming and error-prone process.
Ajay Mahimkar, Jennifer Yates, Yin Zhang 0001, Aman Shaikh, Jia Wang 0001, Zihui Ge, Cheng Tien Ee
CoNEXT4
2008 Where's Waldo? practical searches for stability in iBGP
abstract
What does a childpsilas search of a large, complex cartoon for the eponymous character (Waldo) have to do with Internet routing? Network operators also search complex datasets, but Waldo is the least of their worries. Routing oscillation is a much greater concern. Networks can be designed to avoid routing oscillation, but the approaches so far proposed unnecessarily reduce the configuration flexibility. More importantly, apparently minor changes to a configuration can lead to instability. Verification of network stability is therefore an important task, but unlike the childpsilas search, this problem is NP hard. Until now, no practical method was available for large networks. In this paper, we present an efficient algorithm for proving stability of iBGP, or finding the potential oscillatory modes, and demonstrate its efficacy by applying it to the iBGP configuration of a large Tier-2 AS.
Ashley Flavel, Matthew Roughan, Nigel G. Bean, Aman Shaikh
ICNP4
2008 Impact of hot-potato routing changes in IP networks
Renata Teixeira, Aman Shaikh, Timothy G. Griffin, Jennifer Rexford
IEEE/ACM Trans. Netw.2
2007 Measurement Informed Route Selection
Nick G. Duffield, Kartik Gopalan, Michael R. Hines, Aman Shaikh, Jacobus E. van der Merwe
PAM4
2007 Wresting Control from BGP: Scalable Fine-Grained Route Control
Patrick Verkaik, Dan Pei, Tom Scholl, Aman Shaikh, Alex C. Snoeren, Jacobus E. van der Merwe
USENIX ATC4
2006 Placing Relay Nodes for Intra-Domain Path Diversity
abstract
Abstract — To increase reliability and robustness of missioncritical services in the face of routing changes, it is often desirable and beneficial to take advantage of path diversity provided by the network topology. One way of achieving this inside a single Autonomous System (AS) is to use two paths between every Origin-Destination (OD) pair. One path is the default path defined by the intra-domain routing protocol; the other path is defined as an overlay path that passes through a strategically placed relay node. The key question then is how to place such relay nodes inside an AS, which is the focus of this paper. We propose two heuristic algorithms to find the positions of relay nodes such that every OD pair has an overlay path, going through a relay node, that is disjoint from the default path. When it is not possible to find completely disjoint overlay paths, we allow overlay paths to have overlapped links with default paths. Since overlapped links diminish the robustness of overlay paths against a single point of failure, we introduce the notion of penalty for partially disjoint paths. We apply our algorithms on three different types of topology data – real, inferred, and synthetic – and show that our algorithms find relay nodes of close-to-minimum penalty. Using daily topology snapshots and network event log, we also show that our choices for relay nodes are relatively insensitive to network dynamics; which is very important for a placement algorithm to be viable and practical. Index Terms — Routing protocols, Overlay networks, Path diversity. I.
Meeyoung Cha, Sue B. Moon, Chong-Dae Park, Aman Shaikh
INFOCOM4
2006 Avoiding instability during graceful shutdown of multiple OSPF routers
Aman Shaikh, Rohit Dube, Anujan Varma
IEEE/ACM Trans. Netw.1
2005 Design and Implementation of a Routing Control Platform
Matthew Caesar 0001, Donald F. Caldwell, Nick Feamster, Jennifer Rexford, Aman Shaikh, Jacobus E. van der Merwe
NSDI5
2004 OSPF Monitoring: Architecture, Design, and Deployment Experience
Aman Shaikh, Albert G. Greenberg
NSDI1
2004 Network sensitivity to hot-potato disruptions
abstract
Hot-potato routing is a mechanism employed when there are multiple (equally good) interdomain routes available for a given destination. In this scenario, the Border Gateway Protocol (BGP) selects the interdomain route associated with the closest egress point based upon intradomain path costs. Consequently, intradomain routing changes can impact interdomain routing and cause abrupt swings of external routes, which we call hot-potato disruptions. Recent work has shown that hot-potato disruptions can have a substantial impact on large ISP backbones and thereby jeopardize the network robustness. As a result, there is a need for guidelines and tools to assist in the design of networks that minimize hot-potato disruptions. However, developing these tools is challenging due to the complex and subtle nature of the interactions between exterior and interior routing. In this paper, we address these challenges using an analytic model of hot-potato routing that incorporates metrics to evaluate network sensitivity to hot-potato disruptions. We then present a methodology for computing these metrics using measurements of real ISP networks. We demonstrate the utility of our model by analyzing the sensitivity of a large AS in a tier~1 ISP network.
Renata Teixeira, Aman Shaikh, Timothy G. Griffin, Geoffrey M. Voelker
SIGCOMM2
2004 Dynamics of hot-potato routing in IP networks
abstract
Despite the architectural separation between intradomain and interdomain routing in the Internet, intradomain protocols do influence the path-selection process in the Border Gateway Protocol (BGP). When choosing between multiple equally-good BGP routes, a router selects the one with the closest egress point, based on the intradomain path cost. Under such hot-potato routing, an intradomain event can trigger BGP routing changes. To characterize the influence of hot-potato routing, we conduct controlled experiments with a commercial router. Then, we propose a technique for associating BGP routing changes with events visible in the intradomain protocol, and apply our algorithm to AT&T's backbone network. We show that (i) hot-potato routing can be a significant source of BGP updates, (ii) BGP updates can lag 60 seconds or more behind the intradomain event, (iii) the number of BGP path changes triggered by hot-potato routing has a nearly uniform distribution across destination prefixes, and (iv) the fraction of BGP messages triggered by intradomain changes varies significantly across time and router locations. We show that hot-potato routing changes lead to longer delays in forwarding-plane convergence, shifts in the flow of traffic to neighboring domains, extra externally-visible BGP update messages, and inaccuracies in Internet performance measurements.
Renata Teixeira, Aman Shaikh, Timothy G. Griffin, Jennifer Rexford
SIGMETRICS2
2003 An Efficient Algorithm for OSPF Subnet Aggregation
abstract
Multiple addresses within an OSPF area can be aggregated and advertised together to other areas. This process is known as address aggregation and is used to reduce router computational overheads and memory requirements and to reduce the network bandwidth consumed by OSPF messages. The downside of address aggregation is that it leads to information loss and consequently sub-optimal (non-shortest path) routing of data packets. The resulting difference (path selection error) between the length of the actual forwarding path and the shortest path varies between different sources and destinations. This paper proves that the path selection error from any source to any destination can be bounded using only parameters describing the destination area. Based on this, the paper presents an efficient algorithm that generates the minimum number of aggregates subject to a maximum allowed path selection error. A major operational benefit of our algorithm is that network administrators can select aggregates for an area based solely on the topology of the area without worrying about remaining areas of the OSPF network. The other benefit is that the algorithm enables trade-offs between the number of aggregates and the bound on the path selection error. The paper also evaluates the algorithm's performance on random topologies. Our results show that in some cases, the algorithm is capable of reducing the number of aggregates by as much as 50% with only a relatively small introduction of maximum path selection error.
Aman Shaikh, Dongmei Wang, Guangzhi Li, Jennifer Yates, Charles R. Kalmanek
ICNP1
2002 A case study of OSPF behavior in a large enterprise network
abstract
Open Shortest Path First (OSPF) is widely deployed in IP networks to manage intra-domain routing. OSPF is a link-state protocol, in which routers reliably flood "Link State Advertisements" (LSAs), enabling each to build a consistent, global view of the routing topology. Reliable performance hinges on routing stability, yet the behavior of large operational OSPF networks is not well understood. In this paper, we provide a case study on the eharacteristics and dynamics of LSA traffic for a large enterprise network. This network consists of several hundred routers, distributed in tens of OSPF areas, and connected by LANs and private lines. For this network, we focus on LSA traffic and analyze: (a) the class of LSAs triggered by OSPF's soft-state refresh, (b) the class of LSAs triggered by events that change the status of the network, and (c) a class of "duplicate" LSAs received due to redundancy in OSPF's reliable LSA flooding mechanism. We derive the baseline rate of refresh-triggered LSAs automatically from network configuration information. We also investigate finer time scale statistical properties of this traffic, including burstiness, periodicity, and synchronization. We discuss root causes of event-triggered and duplicate LSA traffic, as well as steps identified to reduce this traffic (e.g., localizing a failing router or changing the OSPF configuration).
Aman Shaikh, Chris Isett, Albert G. Greenberg, Matthew Roughan, Joel Gottlieb
Internet Measurement Workshop1
2002 Avoiding Instability during Graceful Shutdown of OSPF
abstract
In this paper, we describe an enhancement to OSPF, called the IBB (I'll Be Back) capability, that enables other routers to use a router whose OSPF process is inactive for forwarding traffic for a certain period of time. The IBB capability can be used for avoiding route flaps that occur when the OSPF process is brought down in a router to facilitate protocol software upgrade, operating system upgrade, router ID change, AS and interface renumbering, etc. When the OSPF process in an IBB-capable router is inactive, it cannot adapt its forwarding table to reflect changes in network topology. This can lead to routing loops and/or black holes. We provide a detailed analysis of how and when loops or black holes are formed and propose solutions to prevent them. Using the GateD platform, we have developed an IBB extension to OSPF incorporating these solutions. Using this system in an experimental setup, we demonstrate that the overhead of the IBB extension is modest compared to the benefit it offers, and has good scaling behavior in terms of network size and the number of routers with inactive OSPF processes.
Aman Shaikh, Rohit Dube, Anujan Varma
INFOCOM1
2002 An OSPF topology server: design and evaluation
abstract
In large scale, operational Internet protocol networks, creating timely, accurate and network-wide views of the intradomain topology is a fundamental problem. Topical network backbones consist of hundreds of routers, which establish routing adjacencies with one another through static configuration and dynamic routing protocols, such as open shortest path first (OSPF). We describe the design of an OSPF topology server which tracks intradomain topology, by passively and safely listening into OSPFs reliable flooding mechanism, or by pushing and pulling information from the routers via the simple network management protocol. We provide a detailed evaluation and comparison of the two approaches in terms of operational issues, reliability and timeliness of information.
Aman Shaikh, Mukul Goyal, Albert G. Greenberg, Raju Rajan, K. K. Ramakrishnan
IEEE J. Sel. Areas Commun.1
2000 Routability stability in congested networks: experimentation and analysis
abstract
Loss of the routing protocol messages due to network congestion can cause peering session failures in routers, leading to route flaps and routing instabilities. We study the effects of traffic overload on routing protocols by quantifying the stability and robustness properties of two common Internet routing protocols, OSPF and BGP, when the routing control traffic is not isolated from data traffic. We develop analytical models to quantify the effect of congestion on the robustness of OSPF and BGP as a function of the traffic overload factor, queueing delays, and packet sizes. We perform extensive measurements in an experimental network of routers to validate the analytical results. Subsequently we use the analytical framework to investigate the effect of factors that are difficult to incorporate into an experimental setup, such as a wide range of link propagation delays and packet dropping policies. Our results show that increased queueing and propagation delays adversely affect BGP's resilience to congestion, in spite of its use of a reliable transport protocol. Our findings demonstrate the importance of selective treatment of routing protocol messages from other traffic, by using scheduling and utilizing buffer management policies in the routers, to achieve stable and robust network operation.
Aman Shaikh, Anujan Varma, Lampros Kalampoukas, Rohit Dube
SIGCOMM1