Ankit Singla

dblp:78/8146 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-8971-8101ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 6 first-author · 2 since 2021Systems, architecture and hardware · 11 · 2 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
23 papers
Internet architecture and protocols · 22% Routing and switching · 19% Vehicular, aerial and satellite networks · 12%
Computer architecture, parallel and distributed computing, and storage systems
11 papers
Cloud and datacenter computing · 47% Interconnection networks and networks-on-chip · 24% Electronic design automation · 23%
Artificial intelligence
3 papers
Efficient and distributed learning · 100%

Topics — the 30 heaviest of 54, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
serverless computing
1.322024
A systematic evaluation of machine learning on serverless infrastructure · VLDB J. 2024
Towards Demystifying Serverless Machine Learning Training · SIGMOD Conference 2021
Cloud and datacenter computing › serverless computing
serverless distributed training
1.322024
A systematic evaluation of machine learning on serverless infrastructure · VLDB J. 2024
Towards Demystifying Serverless Machine Learning Training · SIGMOD Conference 2021
Electronic design automation › physical design
routing
0.922021
High-Performance Routing With Multipathing and Path Diversity in Ethernet and HPC Networks · IEEE Trans. Parallel Distributed Syst. 2021
FatPaths: routing in supercomputers and data centers when shortest paths fall short · SC 2020
Internet architecture and protocols
network topology
0.942017
Beyond fat-trees without antennae, mirrors, and disco-balls · SIGCOMM 2017
Measuring and understanding throughput of network topologies · SC 2016
OSA: An Optical Switching Architecture for Data Center Networks With Unprecedented Flexibility · IEEE/ACM Trans. Netw. 2014
Machine learning › Efficient and distributed learning
distributed training
0.922021
Towards Demystifying Serverless Machine Learning Training · SIGMOD Conference 2021
Distributed Learning over Unreliable Networks · ICML 2019
Software-defined and programmable networks › SDN control plane
distributed SDN
0.812024
A Decentralized SDN Architecture for the WAN · SIGCOMM 2024
Routing and switching › routing protocol
routing convergence
0.812024
A Decentralized SDN Architecture for the WAN · SIGCOMM 2024
Interconnection networks and networks-on-chip
network topology
0.752021
High-Performance Routing With Multipathing and Path Diversity in Ethernet and HPC Networks · IEEE Trans. Parallel Distributed Syst. 2021
Measuring and understanding throughput of network topologies · SC 2016
Measuring throughput of data center network topologies · SIGMETRICS 2014
Internet architecture and protocols
internet service provider
0.612022
cISP: A Speed-of-Light Internet Service Provider · NSDI 2022
Routing and switching
routing
0.522020
FatPaths: routing in supercomputers and data centers when shortest paths fall short · SC 2020
Scalable routing on flat names · CoNEXT 2010
Internet architecture and protocols
low-latency network design
0.522020
Network topology design at 27, 000 km/hour · CoNEXT 2019
A Bird's Eye View of the World's Fastest Networks · Internet Measurement Conference 2020
Interconnection networks and networks-on-chip › network topology
low-diameter topology
0.512021
High-Performance Routing With Multipathing and Path Diversity in Ethernet and HPC Networks · IEEE Trans. Parallel Distributed Syst. 2021
Electronic design automation › physical design › routing
multipath routing
0.512021
High-Performance Routing With Multipathing and Path Diversity in Ethernet and HPC Networks · IEEE Trans. Parallel Distributed Syst. 2021
Network measurement and analytics
latency measurement
0.412020
A Bird's Eye View of the World's Fastest Networks · Internet Measurement Conference 2020
Vehicular, aerial and satellite networks › satellite networks
LEO satellite constellation
0.412020
Exploring the "Internet from space" with Hypatia · Internet Measurement Conference 2020
Routing and switching
multipath routing
0.412020
FatPaths: routing in supercomputers and data centers when shortest paths fall short · SC 2020
Vehicular, aerial and satellite networks › satellite constellation
satellite constellation networks
0.412019
Network topology design at 27, 000 km/hour · CoNEXT 2019
Vehicular, aerial and satellite networks › satellite communication
satellite internet
0.412019
Network topology design at 27, 000 km/hour · CoNEXT 2019
Network measurement and analytics
traffic characterization
0.412019
Is advance knowledge of flow sizes a plausible assumption? · NSDI 2019
Distributed systems
fault tolerance
0.412019
Distributed Learning over Unreliable Networks · ICML 2019
Network optimization and economics › network design
network topology design
0.322014
High Throughput Data Center Topology Design · NSDI 2014
Jellyfish: Networking Data Centers Randomly · NSDI 2012
Datacenter networks
optical datacenter network
0.322014
OSA: An Optical Switching Architecture for Data Center Networks With Unprecedented Flexibility · IEEE/ACM Trans. Netw. 2014
OSA: An Optical Switching Architecture for Data Center Networks with Unprecedented Flexibility · NSDI 2012
Optical networks › optical switch
optical switch architecture
0.322014
OSA: An Optical Switching Architecture for Data Center Networks With Unprecedented Flexibility · IEEE/ACM Trans. Netw. 2014
OSA: An Optical Switching Architecture for Data Center Networks with Unprecedented Flexibility · NSDI 2012
Network performance modeling
topology comparison
0.212016
Measuring and understanding throughput of network topologies · SC 2016
Network performance modeling › traffic modeling
traffic matrix
0.212016
Measuring and understanding throughput of network topologies · SC 2016
Internet architecture and protocols
wide area network
0.212024
A Decentralized SDN Architecture for the WAN · SIGCOMM 2024
Datacenter networks
lossless ethernet
0.212014
Practical DCB for improved data center networks · INFOCOM 2014
Transport protocols and congestion control
TCP variants
0.212014
Practical DCB for improved data center networks · INFOCOM 2014
Network performance modeling
throughput analysis
0.212014
Measuring throughput of data center network topologies · SIGMETRICS 2014
Internet of things and sensor networks › topology control
connectivity guarantee
0.212013
Ensuring Connectivity via Data Plane Mechanisms · NSDI 2013

Methods — techniques the papers use, named apart from their topics

benchmarking · 1.5taxonomy · 1.0path diversity analysis · 1.0analytic cost-performance modeling · 1.0transport layer redesign · 0.9flowlet switching · 0.9parameter server · 0.8convergence analysis · 0.8network design · 0.6traffic matrix generation · 0.5benchmarking framework · 0.5network reconstruction from public information · 0.4network topology design · 0.4flow size analysis · 0.4
YearPublicationVenuePosition
2024 A Decentralized SDN Architecture for the WAN
abstract
Motivated by our experiences operating a global WAN, we argue that SDN's reliance on infrastructure external to the data plane has substantially complicated the challenge of maintaining high availability. We propose a new decentralized SDN (dSDN) architecture in which SDN control logic instead runs within routers, eliminating the control plane's reliance on external infrastructure and restoring fate-sharing between control and data planes. We present dSDN as a simpler approach to realizing the benefits of SDN in the WAN. Despite its much simpler design, we show that dSDN is practical from an implementation viewpoint, and outperforms centralized SDN in terms of routing convergence and SLO impact.
Alexander Krentsel, Nitika Saran, Bikash Koley, Subhasree Mandal, Ashok Narayanan, Sylvia Ratnasamy, Ali Al-Shabibi, Anees Shaikh, Rob Shakir, Ankit Singla, Hakim Weatherspoon
SIGCOMM10
2024 A systematic evaluation of machine learning on serverless infrastructure
Jiawei Jiang 0001, Shaoduo Gan, Bo Du 0001, Gustavo Alonso, Ana Klimovic, Ankit Singla, Wentao Wu 0001, Sheng Wang 0007, Ce Zhang 0001
VLDB J.6
2022 cISP: A Speed-of-Light Internet Service Provider
Debopam Bhattacherjee, Waqar Aqeel, Sangeetha Abdu Jyothi, Ilker Nadi Bozkurt, William Sentosa, Muhammad Tirmazi, Anthony Aguirre, Balakrishnan Chandrasekaran 0002, Brighten Godfrey, Gregory Laughlin, Bruce M. Maggs, Ankit Singla
NSDI12
2021 Towards Demystifying Serverless Machine Learning Training
abstract
The appeal of serverless (FaaS) has triggered a growing interest on how to use it in data-intensive applications such as ETL, query processing, or machine learning (ML). Several systems exist for training large-scale ML models on top of serverless infrastructures (e.g., AWS Lambda) but with inconclusive results in terms of their performance and relative advantage over "serverful" infrastructures (IaaS). In this paper we present a systematic, comparative study of distributed ML training over FaaS and IaaS. We present a design space covering design choices such as optimization algorithms and synchronization protocols, and implement a platform, LambdaML, that enables a fair comparison between FaaS and IaaS. We present experimental results using LambdaML, and further develop an analytic model to capture cost/performance tradeoffs that must be considered when opting for a serverless infrastructure. Our results indicate that ML training pays off in serverless only for models with efficient (i.e., reduced) communication and that quickly converge. In general, FaaS can be much faster but it is never significantly cheaper than IaaS.
Jiawei Jiang 0001, Shaoduo Gan, Fanlin Wang, Gustavo Alonso, Ana Klimovic, Ankit Singla, Wentao Wu 0001, Ce Zhang 0001
SIGMOD Conference7
2021 ICARUS: Attacking low Earth orbit satellite networks
Giacomo Giuliari, Tommaso Ciussani, Adrian Perrig, Ankit Singla
USENIX ATC4
2021 High-Performance Routing With Multipathing and Path Diversity in Ethernet and HPC Networks
abstract
The recent line of research into topology design focuses on lowering network diameter. Many low-diameter topologies such as Slim Fly or Jellyfish that substantially reduce cost, power consumption, and latency have been proposed. A key challenge in realizing the benefits of these topologies is routing. On one hand, these networks provide shorter path lengths than established topologies such as Clos or torus, leading to performance improvements. On the other hand, the number of shortest paths between each pair of endpoints is much smaller than in Clos, but there is a large number of non-minimal paths between router pairs. This hampers or even makes it impossible to use established multipath routing schemes such as ECMP. In this article, to facilitate high-performance routing in modern networks, we analyze existing routing protocols and architectures, focusing on how well they exploit the diversity of minimal and non-minimal paths. We first develop a taxonomy of different forms of support for multipathing and overall path diversity. Then, we analyze how existing routing schemes support this diversity. Among others, we consider multipathing with both shortest and non-shortest paths, support for disjoint paths, or enabling adaptivity. To address the ongoing convergence of HPC and “Big Data” domains, we consider routing protocols developed for both HPC systems and for data centers as well as general clusters. Thus, we cover architectures and protocols based on Ethernet, InfiniBand, and other HPC networks such as Myrinet. Our review will foster developing future high-performance multipathing routing protocols in supercomputers and data centers.
Maciej Besta, Jens Domke, Marcel Schneider, Marek Konieczny, Salvatore Di Girolamo, Timo Schneider, Ankit Singla, Torsten Hoefler
IEEE Trans. Parallel Distributed Syst.7
2020 Specializing the network for scatter-gather workloads
abstract
Data processing and distributed querying workloads often involve a "scatter-gather" or "partition-aggregate" architectural pattern, whereby one application queries hundreds or even thousands of workers. Network communication is often a bottleneck in this pattern, especially when the compute task at each worker is small, such as for Web queries and interactive analytics. The network bottleneck can result in low throughput, high CPU utilization, and cause job completion time to increase by orders of magnitude. To overcome these inefficiencies, we explore hardware-offload of the scatter-gather primitive, whereby a smart NIC takes on the responsibility of sending out queries and collecting responses. We show that this approach not only virtually eliminates CPU usage, but with suitable scheduling of responses, it also speeds up scatter by allowing parallel queries, and gather by preventing throughput collapse due to excessive congestion. Besides response scheduling, we use a careful design at the NIC to limit FPGA resource usage: our approach uses about 25% of on-chip logic and 33% of on-chip memory on a mid-sized FPGA, leaving enough room for implementing other functions on the smart NIC.
Catalina Álvarez, Zhenhao He, Gustavo Alonso, Ankit Singla
SoCC4
2020 Photons: lambdas on a diet
abstract
Serverless computing allows users to create short, stateless functions and invoke hundreds of them concurrently to tackle massively parallel workloads. We observe that even though most of the footprint of a serverless function is fixed across its invocations --- language runtime, libraries, and other application state --- today's serverless platforms do not exploit this redundancy. Such an inefficiency has cascading negative impacts: longer startup times, lower throughput, higher latency, and higher cost. To mitigate these problems, we have built Photons, a framework leveraging workload parallelism to co-locate multiple instances of the same function within the same runtime. Concurrent invocations can then share the runtime and application state transparently, without compromising execution safety. Photons reduce function's memory consumption by 25% to 98% per invocation, with no performance degradation compared to today's serverless platforms. We also show that our approach can reduce the overall memory utilization by 30%, and the total number of cold starts by 52%.
Vojislav Dukic, Rodrigo Bruno, Ankit Singla, Gustavo Alonso
SoCC3
2020 In-orbit Computing: An Outlandish thought Experiment?
abstract
Space industry upstarts are deploying thousands of satellites to offer global Internet service. These plans promise large improvements in coverage and latency, and could fundamentally transform the Internet. But what if this transformation extends beyond network transit into a new type of computing service? What if each satellite, in addition to serving as a network router, also offers cloud-like compute, making the new constellations not just global Internet service providers, but at the same time, a new breed of cloud providers offering "compute where you need it"?
Debopam Bhattacherjee, Simon Kassing, Melissa Licciardello, Ankit Singla
HotNets4
2020 "Internet from Space" without Inter-satellite Links
abstract
Buoyed by advances in space technology, several firms are planning satellite constellations to offer broadband Internet service. While these developments are happening quickly, there are also many uncertainties about the design of these networks. A key open question is whether or not they will incorporate direct connectivity between satellites, instead of only ground-satellite connections. We compare the network behavior resulting from the two outcomes of that question. Our analysis shows that inter-satellite links substantially reduce the temporal variations in latency, add greater resilience to weather, and could yield more than 3x the throughput achieved without such links. Thus, whether this one design element pans out could have a large bearing on the performance, reliability, and economics of these networks.
Yannick Hauri, Debopam Bhattacherjee, Manuel Grossmann, Ankit Singla
HotNets4
2020 A Bird's Eye View of the World's Fastest Networks
abstract
Low latency is of interest for a variety of applications. The most stringent latency requirements arise in financial trading, where sub-microsecond differences matter. As a result, firms in the financial technology sector are pushing networking technology to its limits, giving a peek into the future of consumer-grade terrestrial microwave networks. Here, we explore the world's most competitive network design race, which has played out over the past decade on the Chicago-New Jersey trading corridor. We systematically reconstruct licensed financial trading networks from publicly available information, and examine their latency, path redundancy, wireless link lengths, and operating frequencies.
Debopam Bhattacherjee, Waqar Aqeel, Gregory Laughlin, Bruce M. Maggs, Ankit Singla
Internet Measurement Conference5
2020 Exploring the "Internet from space" with Hypatia
abstract
SpaceX, Amazon, and others plan to put thousands of satellites in low Earth orbit to provide global low-latency broadband Internet. SpaceX's plans have matured quickly, such that their underdeployment satellite constellation is already the largest in history, and may start offering service in 2020.
Simon Kassing, Debopam Bhattacherjee, André Baptista Águas, Jens Eirik Saethre, Ankit Singla
Internet Measurement Conference5
2020 Understanding Video Streaming Algorithms in the Wild
Melissa Licciardello, Maximilian Grüner, Ankit Singla
PAM3
2020 FatPaths: routing in supercomputers and data centers when shortest paths fall short
abstract
We introduce FatPaths: a simple, generic, and robust routing architecture that enables state-of-the-art low-diameter topologies such as Slim Fly to achieve unprecedented performance. FatPaths targets Ethernet stacks in both HPC supercomputers as well as cloud data centers and clusters. FatPaths exposes and exploits the rich (“fat”) diversity of both minimal and non-minimal paths for high-performance multi-pathing. Moreover, FatPaths uses a redesigned “purified” transport layer that removes virtually all TCP performance issues (e.g., the slow start), and incorporates flowlet switching, a technique used to prevent packet reordering in TCP networks, to enable very simple and effective load balancing. Our design enables recent low-diameter topologies to outperform powerful Clos designs, achieving 15% higher net throughput at 2” lower latency for comparable cost. FatPaths will significantly accelerate Ethernet clusters that form more than 50% of the Top500 list and it may become a standard routing scheme for modern topologies.Extended paper version: https://arxiv.org/abs/1906.10885
Maciej Besta, Marcel Schneider, Marek Konieczny, Karolina Cynk, Erik Henriksson, Salvatore Di Girolamo, Ankit Singla, Torsten Hoefler
SC7
2020 Beyond the mega-data center: networking multi-data center regions
abstract
The difficulty of building large data centers in dense metro areas is pushing big cloud providers towards a different approach to scaling: multiple smaller data centers within tens of kilometers of each other, comprising a "region". We show that networking this small number of nearby sites with each other is a surprisingly challenging and multi-faceted problem. We draw out the operational goals and constraints of such networks, and highlight the design trade-offs involved using data from Microsoft Azure's regions.
Vojislav Dukic, Ginni Khanna, Christos Gkantsidis, Thomas Karagiannis, Francesca Parmigiani, Ankit Singla, Mark Filer, Jeffrey L. Cox, Anna Ptasznik, Nick Harland, Winston Saunders, Christian Belady
SIGCOMM6
2020 Reconstructing proprietary video streaming algorithms
Maximilian Grüner, Melissa Licciardello, Ankit Singla
USENIX ATC3
2019 Network topology design at 27, 000 km/hour
abstract
Upstart space companies are actively developing massive constellations of low-flying satellites to provide global Internet service. We examine the problem of designing the inter-satellite network for low latency and high capacity. We posit that the high density of these new constellations and the high-velocity nature of such systems render traditional approaches for network design ineffective, motivating new methods specialized for this problem setting.
Debopam Bhattacherjee, Ankit Singla
CoNEXT2
2019 (Self) Driving Under the Influence: Intoxicating Adversarial Network Inputs
abstract
Traditional network control planes can be slow and require manual tinkering from operators to change their behavior. There is thus great interest in a faster, data-driven approach that uses signals from real-time traffic instead. However, the promise of fast and automatic reaction to data comes with new risks: malicious inputs designed towards negative outcomes for the network, service providers, users, and operators.
Roland Meier, Thomas Holterbach, Stephan Keck, Matthias Stähli, Vincent Lenders, Ankit Singla, Laurent Vanbever
HotNets6
2019 Distributed Learning over Unreliable Networks
abstract
Most of today’s distributed machine learning systems assume reliable networks: whenever two machines exchange information (e.g., gradients or models), the network should guarantee the delivery of the message. At the same time, recent work exhibits the impressive tolerance of machine learning algorithms to errors or noise arising from relaxed communication or synchronization. In this paper, we connect these two trends, and consider the following question: Can we design machine learning systems that are tolerant to network unreliability during training? With this motivation, we focus on a theoretical problem of independent interest—given a standard distributed parameter server architecture, if every communication between the worker and the server has a non-zero probability $p$ of being dropped, does there exist an algorithm that still converges, and at what speed? In the context of prior art, this problem can be phrased as distributed learning over random topologies. The technical contribution of this paper is a novel theoretical analysis proving that distributed learning over random topologies can achieve comparable convergence rate to centralized or distributed learning over reliable networks. Further, we prove that the influence of the packet drop rate diminishes with the growth of the number of parameter servers. We map this theoretical result onto a real-world scenario, training deep neural networks over an unreliable network layer, and conduct network simulation to validate the system improvement by allowing the networks to be unreliable.
Chen Yu 0009, Hanlin Tang 0002, Cédric Renggli, Simon Kassing, Ankit Singla, Dan Alistarh, Ce Zhang 0001, Ji Liu 0002
ICML5
2019 Is advance knowledge of flow sizes a plausible assumption?
Vojislav Dukic, Sangeetha Abdu Jyothi, Bojan Karlas, Muhsen Owaida, Ce Zhang 0001, Ankit Singla
NSDI6
2018 Network Scheduling in the Dark
abstract
No abstract available.
Vojislav Dukic, Sangeetha Abdu Jyothi, Bojan Karlas, Muhsen Owaida, Ce Zhang 0001, Ankit Singla
SoCC6
2018 Providing Multi-tenant Services with FPGAs: Case Study on a Key-Value Store
abstract
FPGAs can be used to speed up computation and data management tasks in various application domains. In cloud settings, however, high utilization is as important as high performance. In software it is common to co-locate different tenants' workloads on the same servers to increase utilization. Sharing an FPGA is more complex because applications take up physical space on the chip. Even though it is possible to physically partition the FPGA, tenants can have widely different requirements and their needs can also fluctuate over time. In this paper, we take a different approach and provide flexibility to the tenants who are interested in the same type of application but have different workloads and quality of service requirements. We demonstrate our approach of multi-tenant design using a key-value store service but the ideas generalize to other network-facing services as well. A key challenge of multi-tenancy is to efficiently share the underlying hardware while enforcing strict data and performance isolation between tenants. In this paper we demonstrate that, by following a single-pipeline design principle, it is possible to control each tenant's share of network bandwidth and computational resources even for complex, distributed operations. Furthermore, we show how state-machine based logic on the FPGA can be made tenant-aware without introducing significant context-switching overhead. Finally, our hardware design provides flexibility for changing per-tenant shares, allowing the same circuit to be used by one or multiple tenants without performance loss.
Zsolt István, Gustavo Alonso, Ankit Singla
FPL3
2018 Gearing up for the 21st century space race
abstract
A new space race is imminent, with several industry players working towards satellite-based Internet connectivity. While satellite networks are not themselves new, these recent proposals are aimed at orders of magnitude higher bandwidth and much lower latency, with constellations planned to comprise thousands of satellites. These are not merely far future plans --- the first satellite launches have already commenced, and substantial planned capacity has already been sold. It is thus critical that networking researchers engage actively with this research space, instead of missing what may be one of the most significant modern developments in networking.
Debopam Bhattacherjee, Waqar Aqeel, Ilker Nadi Bozkurt, Anthony Aguirre, Balakrishnan Chandrasekaran 0002, Brighten Godfrey, Gregory Laughlin, Bruce M. Maggs, Ankit Singla
HotNets9
2017 Why Is the Internet so Slow?!
Ilker Nadi Bozkurt, Anthony Aguirre, Balakrishnan Chandrasekaran 0002, Brighten Godfrey, Gregory Laughlin, Bruce M. Maggs, Ankit Singla
PAM7
2017 Beyond fat-trees without antennae, mirrors, and disco-balls
abstract
Recent studies have observed that large data center networks often have a few hotspots while most of the network is underutilized. Consequently, numerous data center network designs have explored the approach of identifying these communication hotspots in real-time and eliminating them by leveraging flexible optical or wireless connections to dynamically alter the network topology. These proposals are based on the premise that statically wired network topologies, which lack the opportunity for such online optimization, are fundamentally inefficient, and must be built at uniform full capacity to handle unpredictably skewed traffic.
Simon Kassing, Asaf Valadarsky, Gal Shahaf, Michael Schapira, Ankit Singla
SIGCOMM5
2016 Fat-FREE Topologies
abstract
With the growing size of data center networks, full-bandwidth connectivity between all pairs of servers is becoming difficult and expensive to scale. Thus, numerous recent topology proposals incorporate reconfigurable wireless and optical connectivity, allowing the topology to adapt to the traffic demands --- only servers that require bandwidth at any given time receive such dynamic connections. Implicitly, this work has suggested that statically wired topologies are fundamentally inflexible, and would need to be built at full capacity to handle unpredictably skewed traffic.
Ankit Singla
HotNets1
2016 Measuring and understanding throughput of network topologies
abstract
High throughput is of particular interest in data center and HPC networks. Although myriad network topologies have been proposed, a broad head-to-head comparison across topologies and across traffic patterns is absent, and the right way to compare worst-case throughput performance is a subtle problem. In this paper, we develop a framework to benchmark the throughput of network topologies, using a two-pronged approach. First, we study performance on a variety of synthetic and experimentally-measured traffic matrices (TMs). Second, we show how to measure worst-case throughput by generating a near-worst-case TM for any given topology. We apply the framework to study the performance of these TMs in a wide range of network topologies, revealing insights into the performance of topologies with scaling, robustness of performance across TMs, and the effect of scattered workload placement. Our evaluation code is freely available.
Sangeetha Abdu Jyothi, Ankit Singla, Brighten Godfrey, Alexandra Kolla
SC2
2014 The Internet at the Speed of Light
abstract
For many Internet services, reducing latency improves the user experience and increases revenue for the service provider. While in principle latencies could nearly match the speed of light, we find that infrastructural inefficiencies and protocol overheads cause today's Internet to be much slower than this bound: typically by more than one, and often, by more than two orders of magnitude. Bridging this large gap would not only add value to today's Internet applications, but could also open the door to exciting new applications. Thus, we propose a grand challenge for the networking research community: a speed-of-light Internet. To inform this research agenda, we investigate the causes of latency inflation in the Internet across the network stack. We also discuss a few broad avenues for latency improvement.
Ankit Singla, Balakrishnan Chandrasekaran 0002, Brighten Godfrey, Bruce M. Maggs
HotNets1
2014 Practical DCB for improved data center networks
abstract
Storage area networking is driving commodity data center switches to support lossless Ethernet (DCB). Unfortunately, to enable DCB for all traffic on arbitrary network topologies, we must address several problems that can arise in lossless networks, e.g., large buffering delays, unfairness, head of line blocking, and deadlock. We propose TCP-Bolt, a TCP variant that not only addresses the first three problems but reduces flow completion times by as much as 70%. We also introduce a simple, practical deadlock-free routing scheme that eliminates deadlock while achieving aggregate network throughput within 15% of ECMP routing. This small compromise in potential routing capacity is well worth the gains in flow completion time. We note that our results on deadlock-free routing are also of independent interest to the storage area networking community. Further, as our hardware testbed illustrates, these gains are achievable today, without hardware changes to switches or NICs.
Brent E. Stephens, Alan L. Cox, Ankit Singla, John B. Carter, Colin Dixon, Wes Felter
INFOCOM3
2014 High Throughput Data Center Topology Design
Ankit Singla, Brighten Godfrey, Alexandra Kolla
NSDI1
2014 Measuring throughput of data center network topologies
abstract
High throughput is a fundamental goal of network design. While myriad network topologies have been proposed to meet this goal, particularly in data center and HPC networking, a consistent and accurate method of evaluating a design's throughput performance and comparing it to past proposals is conspicuously absent. In this work, we develop a framework to benchmark the throughput of network topologies and apply this methodology to reveal insights about network structure. We show that despite being commonly used, cut-based metrics such as bisection bandwidth are the wrong metrics: they yield incorrect conclusions about the throughput performance of networks. We therefore measure flow-based throughput directly and show how to evaluate topologies with nearly-worst-case traffic matrices. We use the flow-based throughput metric to compare the throughput performance of a variety of computer networks. We have made our evaluation framework freely available to facilitate future work on design and evaluation of networks.
Sangeetha Abdu Jyothi, Ankit Singla, Brighten Godfrey, Alexandra Kolla
SIGMETRICS2
2014 OSA: An Optical Switching Architecture for Data Center Networks With Unprecedented Flexibility
abstract
A detailed examination of evolving traffic characteristics, operator requirements, and network technology trends suggests a move away from nonblocking interconnects in data center networks (DCNs). As a result, recent efforts have advocated oversubscribed networks with the capability to adapt to traffic requirements on-demand. In this paper, we present the design, implementation, and evaluation of OSA, a novel Optical Switching Architecture for DCNs. Leveraging runtime reconfigurable optical devices, OSA dynamically changes its topology and link capacities, thereby achieving unprecedented flexibility to adapt to dynamic traffic patterns. Extensive analytical simulations using both real and synthetic traffic patterns demonstrate that OSA can deliver high bisection bandwidth (60%-100% of the nonblocking architecture). Implementation and evaluation of a small-scale functional prototype further demonstrate the feasibility of OSA.
Kai Chen 0005, Ankit Singla, Atul Singh, Kishore Ramachandran, Lei Xu 0017, Yueping Zhang, Xitao Wen, Yan Chen 0004
IEEE/ACM Trans. Netw.2
2013 Ensuring Connectivity via Data Plane Mechanisms
Junda Liu, Aurojit Panda, Ankit Singla, Brighten Godfrey, Michael Schapira, Scott Shenker
NSDI3
2012 Jellyfish: Networking Data Centers Randomly
Ankit Singla, Chi-Yao Hong, Lucian Popa 0002, Brighten Godfrey
NSDI1
2012 OSA: An Optical Switching Architecture for Data Center Networks with Unprecedented Flexibility
Kai Chen 0005, Ankit Singla, Atul Singh, Kishore Ramachandran, Lei Xu 0017, Yueping Zhang, Xitao Wen, Yan Chen 0004
NSDI2
2012 Brief announcement: on the resilience of routing tables
abstract
Many modern network designs incorporate "failover" paths into routers' forwarding tables. We initiate the theoretical study of such resilient routing tables.
Joan Feigenbaum, Brighten Godfrey, Aurojit Panda, Michael Schapira, Scott Shenker, Ankit Singla
PODC6
2011 Information-centric networking: seeing the forest for the trees
abstract
There have been many recent papers on data-oriented or content-centric network architectures. Despite the voluminous literature, surprisingly little clarity is emerging as most papers focus on what differentiates them from other proposals. We begin this paper by identifying the existing commonalities and important differences in these designs, and then discuss some remaining research issues. After our review, we emerge skeptical (but open-minded) about the value of this approach to networking.
Ali Ghodsi 0002, Scott Shenker, Teemu Koponen, Ankit Singla, Barath Raghavan, James R. Wilcox
HotNets4
2011 Intelligent design enables architectural evolution
abstract
What does it take for an Internet architecture to be evolvable? Despite our ongoing frustration with today's rigid IP-based architecture and the research community's extensive research on clean-slate designs, it remains unclear how to best design for architectural evolvability. We argue here that evolvability is far from mysterious. In fact, we claim that only a few "intelligent" design changes are needed to support evolvability. While these changes are definitely nonincremental (i.e., cannot be deployed in an incremental fashion starting with today's architecture), they follow directly from the well-known engineering principles of indirection, modularity, and extensibility.
Ali Ghodsi 0002, Scott Shenker, Teemu Koponen, Ankit Singla, Barath Raghavan, James R. Wilcox
HotNets4
2010 Verifiable network-performance measurements
abstract
In the current Internet, there is no clean way for affected parties to react to poor forwarding performance: when a domain violates its Service Level Agreement (SLA) with a contractual partner, the partner must resort to ad-hoc probing-based monitoring to determine the existence and extent of the violation. Instead, we propose a new, systematic approach to the problem of forwarding-performance verification. Our mechanism relies on voluntary reporting, allowing each domain to disclose its loss and delay performance to its neighbors; it does not disclose any information regarding the participating domains' topology or routing policies beyond what is already publicly available. Most importantly, it enables verifiable performance measurements, i.e., domains cannot abuse it to significantly exaggerate their performance. Finally, our mechanism is tunable, allowing each participating domain to determine how many resources to devote to it independently (i.e., without any inter-domain coordination), exposing a controllable trade-off between performance-verification quality and resource consumption. Our mechanism comes at the cost of deploying modest functionality at the participating domains' border routers; we show that it requires reasonable processing and memory resources within modern network capabilities.
Katerina J. Argyraki, Petros Maniatis, Ankit Singla
CoNEXT3
2010 Scalable routing on flat names
abstract
We introduce a protocol which routes on flat, location-independent identifiers with guaranteed scalability and low stretch. Our design builds on theoretical advances in the area of compact routing, and is the first to realize these guarantees in a dynamic distributed setting.
Ankit Singla, Brighten Godfrey, Kevin R. Fall, Gianluca Iannaccone, Sylvia Ratnasamy
CoNEXT1
2010 Proteus: a topology malleable data center network
abstract
Full-bandwidth connectivity between all servers of a data center may be necessary for all-to-all traffic patterns, but such interconnects suffer from high cost, complexity, and energy consumption. Recent work has argued that if all-to-all traffic is uncommon, oversubscribed network architectures that can adapt the topology to meet traffic demands, are sufficient. In line with this work, we propose Proteus, an all-optical architecture targeting unprecedented topology-flexibility, lower complexity and higher energy efficiency.
Ankit Singla, Atul Singh, Kishore Ramachandran, Lei Xu 0017, Yueping Zhang
HotNets1