VLDB 2026 Research / reviewers in the wild / expert
Srikanth Sundaresan
dblp:47/8093
· DBLP profile ↗
30ranked-venue papers
11as first author
9since 2021 · last 2026
0009-0003-6230-9396ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 21 · 7 first-author · 9 since 2021Security and privacy · 6 · 2 first-authorSystems, architecture and hardware · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Truly Burst-Aware Evaluation of Data Center Congestion Control
Pragna Mamidipaka, Srikanth Sundaresan, Theophilus Benson |
APNet | 2 |
| 2026 | Detection and Localization of End-to-end Bitflip Errors in Data CentersabstractPacket corrupting devices can cause bitflips and data corruption, and, depending on exactly where the corruption occurs, can evade even strong link-level checks. Even though such instances are probabilistically rare, they are all too common at hyper-scale. At Meta, they have been causing end-to-end errors for years in our datacenter fleet, resulting in dozens of instances of performance impact to services, and rare instances of data corruption. Over the years, we have developed three tools that have progressively improved our ability to localize this problem to specific hosts, rack switches, links, and even specific ports on those switches. They include a fleet wide passive monitoring service to track Layer-4 checksum errors detected by hosts, an active monitoring and fault localization system to probe for corruption on devices and localize them to individual links at any layer of the topology; and a loopback test to reproduce and confirm packet corruption on specific ports on localized switches. Together, they comprise Flipperino, a system that detects and localizes packet corrupting devices, making it easy to identify them for remediation. We share our experience from investigations, and recent results from Flipperino, which has localized devices causing 97% of Layer-4 checksum errors. We have also reproduced corruption on 8 switches in the last 3 months, including 7 rack switches and a fabric switch. Abhigyan Sharma, Neil Spring, Srikanth Sundaresan, Francesco Caggioni, Laurent Virot, Arber Mancaj |
SIGCOMM | 3 |
| 2026 | Connecting 100K+ GPUs: Building the Communication Stack for Large-Scale LLM TrainingabstractThe arrival of 100K+ GPU clusters marks a new frontier in AI infrastructure. Standard communication stack meets new challenges as physical topologies span multiple datacenter buildings, introducing high bandwidth-delay product links where latency increases by up to 30× compared to intra-rack traffic. Furthermore, the transition toward Mixture-of-Experts architectures generating bursty all-to-all patterns that create transient congestion hotspots. These constraints, combined with an operational environment where hardware failures shift from anomalies to frequent occurrences, renders traditionally lightweight operations like initialization and resource management challenging. Hongyi Zeng, Min Si, Pavan Balaji, Yongzhou Chen, Ching-Hsiang Chu, Adithya Gangidi, Prashanth Kannan, Bingzhe Liu, Saif Hasan, Deep Shah, Ashmitha Jeevaraj Shetty, Gregory R. Steinbrecher, Srikanth Sundaresan, Yulun Wang, Yexin Wu, Mingran Yang, Kenny Yu, Minlan Yu, Cen Zhao, Shengbao Zheng, Wesley Bland, Denis Boyda, Suman Gumudavelli, Subodh Iyengar, Cristian Lumezanu, Rui Miao 0001, Venkat Ramesh, Jingliang Ren, Maxim Samoylov, Jan Seidel, Qiye Tan, Xinfeng Xie, Yimeng Zhao, Shuqiang Zhang, Art Zhu |
SIGCOMM | 14 |
| 2025 | Congestion Patterns in a Large-scale RDMA DatacenterabstractRDMA datacenters are proliferating to meet the demand of emerging workloads such as AI training and inference as well as distributed storage. This trend has opened up a critical knowledge gap: the traffic characteristics of congestion in these networks remain unknown. We do not know, for example, which layers of the network are the most congested, if the network is load balanced effectively, how long congestion events last, and how accurate existing telemetry systems are in capturing congestion. This paper bridges this gap by investigating congestion in a large-scale RDMA datacenter dedicated to distributed AI training. We provide insights into three specific congestion patterns: (a) location and distribution in the network, (b) burstiness, e.g., the duration and synchrony of bursts, and (c) observability using existing telemetry methods. We show, for instance, that the deployment of Priority Flow Control (PFC) in RDMA networks has shifted the location of congestion one level up: from the edge-host in legacy TCP/IP datacenters to the network core in RDMA datacenters. At the same time, we show that the same protocol enables us to observe and understand congestion better, even bursty events. The findings of this research reveal open challenges for measuring, characterizing, and managing congestion in RDMA networks, paving the way for future research. Soudeh Ghorbani, Yimeng Zhao, Srikanth Sundaresan, Ying Zhang 0022, Yijing Zeng, Abhigyan Sharma, Prashanth Kannan, Cristian Lumezanu |
IMC | 3 |
| 2024 | Understanding Incast Bursts in Modern DatacentersabstractIn datacenters, common incast traffic patterns are challenging because they violate the basic premise of bandwidth stability on which TCP congestion control convergence is built, overwhelming shallow switch buffers and causing packet losses and high latency. To understand why these challenges remain despite decades of research on datacenter congestion control, we conduct an in-depth investigation into high-degree incasts both in production workloads at Meta and in simulation. In addition to characterizing the bursty nature of these incasts and their impacts on the network, our findings demonstrate the shortcomings of widely deployed window-based congestion control techniques used to address incast problems. Furthermore, we find that hosts associated with a specific application or service exhibit similar and predictable incast traffic properties across hours, pointing the way toward solutions that predict and prevent incast bursts, instead of reacting to them. Christopher Canel, Balasubramanian Madhavan, Srikanth Sundaresan, Neil Spring, Prashanth Kannan, Ying Zhang 0022, Srinivasan Seshan |
IMC | 3 |
| 2024 | A large-scale deployment of DCTCP
Abhishek Dhamija, Balasubramanian Madhavan, Hechao Li, Shrikrishna Khare, Madhavi Rao, Lawrence Brakmo, Neil Spring, Prashanth Kannan, Srikanth Sundaresan, Soudeh Ghorbani |
NSDI | 10 |
| 2024 | Netcastle: Network Infrastructure Testing At Scale
Rob Sherwood, Jinghao Shi, Ying Zhang 0022, Neil Spring, Srikanth Sundaresan, Jasmeet Bagga, Prathyusha Peddi, Vineela Kukkadapu, Rashmi Shrivastava, Manikantan K., Pavan Patil, Srikrishna Gopu, Varun Varadan, Ethan Shi, Hany Morsy, Yuting Bu, Renjie Yang, Rasmus Jonsson, Jesus Jussepen Arredondo, Diana Saha, Sean Choi |
NSDI | 5 |
| 2024 | NetEdit: An Orchestration Platform for eBPF Network Functions at ScaleabstractManaging the performance of thousands of services across millions of servers demands a networking stack that can dynamically adjust protocol settings to match diverse priorities and network characteristics. Moreover, given the constantly evolving nature of services and their requirements, the set of configurable protocols must remain adaptable. However, current host networking stacks lack the necessary flexibility and adaptability. Although eBPF shows promise in this regard, it lacks essential primitives for efficient development and safe deployment of multiple co-existing services. Theophilus Benson, Prashanth Kannan, Prankur Gupta, Balasubramanian Madhavan, Kumar Saurabh Arora, Martin Lau, Abhishek Dhamija, Rajiv Krishnamurthy, Srikanth Sundaresan, Neil Spring, Ying Zhang 0022 |
SIGCOMM | 10 |
| 2022 | A microscopic view of bursts, buffer contention, and loss in data centersabstractManaging data center networks with low loss requires understanding traffic dynamics at short (millisecond) time-scales, especially the burstiness of traffic, and to what extent bursts contend for switch buffer resources. Yet, monitoring traffic over such intervals is a challenge at scale. Ehab Ghabashneh, Yimeng Zhao, Cristian Lumezanu, Neil Spring, Srikanth Sundaresan, Sanjay G. Rao |
IMC | 5 |
| 2019 | Internet Performance from Facebook's EdgeabstractWe examine the current state of user network performance and opportunities to improve it from the vantage point of Facebook, a global content provider. Facebook serves over 2 billion users distributed around the world using a network of PoPs and interconnections spread across 6 continents. In this paper, we execute a large-scale, 10-day measurement study of metrics at the TCP and HTTP layers for production user traffic at all of Facebook's PoPs worldwide, collecting performance measurements for hundreds of trillions of sampled HTTP sessions. We discuss our approach to collecting and analyzing measurements, including a novel approach to characterizing user achievable goodput from the server side. We find that most user sessions have MinRTT less than 39ms and can support HD video. We investigate if it is possible to improve performance by incorporating performance information into Facebook's routing decisions; we find that default routing by Facebook is largely optimal. To our knowledge, our measurement study is the first characterization of user performance on today's Internet from the vantage point of a global content provider. Brandon Schlinker, Ítalo S. Cunha, Yi-Ching Chiu, Srikanth Sundaresan, Ethan Katz-Bassett |
Internet Measurement Conference | 4 |
| 2019 | Service Traceroute: Tracing Paths of Application Flows
Ivan Morandi, Francesco Bronzino, Renata Teixeira, Srikanth Sundaresan |
PAM | 4 |
| 2018 | Apps, Trackers, Privacy, and Regulators: A Global Study of the Mobile Tracking Ecosystem
Abbas Razaghpanah, Rishab Nithyanand, Narseo Vallina-Rodriguez, Srikanth Sundaresan, Mark Allman, Christian Kreibich, Phillipa Gill |
NDSS | 4 |
| 2017 | IoT S&P 2017: First Workshop on Internet of Things Security and PrivacyabstractThe First Workshop on Internet of Things Security and Privacy is held in Dallas, TX, USA on November 3, 2017, co-located with the ACM Conference on Computer and Communications Security (CCS). The workshop aims to address the security and privacy challenges of the emerging Internet-of-Things landscape. The workshop aims to bring together academic and industrial researchers, and to that end, we have put together an exciting program offering a a mix of current and potential challenges. The workshop will also features 12 papers, 4 posters, and an invited keynote. Theophilus Benson, Peng Liu 0005, Srikanth Sundaresan, Yuqing Zhang 0001 |
CCS | 3 |
| 2017 | Studying TLS Usage in Android AppsabstractTransport Layer Security (TLS), has become the de-facto standard for secure Internet communication. When used correctly, it provides secure data transfer, but used incorrectly, it can leave users vulnerable to attacks while giving them a false sense of security. Numerous efforts have studied the adoption of TLS (and its predecessor, SSL) and its use in the desktop ecosystem, attacks, and vulnerabilities in both desktop clients and servers. However, there is a dearth of knowledge of how TLS is used in mobile platforms. In this paper we use data collected by Lumen, a mobile measurement platform, to analyze how 7,258 Android apps use TLS in the wild. We analyze and fingerprint handshake messages to characterize the TLS APIs and libraries that apps use, and also evaluate weaknesses. We see that about 84% of apps use default OS APIs for TLS. Many apps use third-party TLS libraries; in some cases they are forced to do so because of restricted Android capabilities. Our analysis shows that both approaches have limitations, and that improving TLS security in mobile is not straightforward. Apps that use their own TLS configurations may have vulnerabilities due to developer inexperience, but apps that use OS defaults are vulnerable to certain attacks if the OS is out of date, even if the apps themselves are up to date. We also study certificate verification, and see low prevalence of security measures such as certificate pinning, even among high-risk apps such as those providing financial services, though we did observe major third-party tracking and advertisement services deploying certificate pinning. Abbas Razaghpanah, Arian Akhavan Niaki, Narseo Vallina-Rodriguez, Srikanth Sundaresan, Johanna Amann, Phillipa Gill |
CoNEXT | 4 |
| 2017 | TCP congestion signaturesabstractWe develop and validate Internet path measurement techniques to distinguish congestion experienced when a flow self-induces congestion in the path from when a flow is affected by an already congested path. One application of this technique is for speed tests, when the user is affected by congestion either in the last mile or in an interconnect link. This difference is important because in the latter case, the user is constrained by their service plan (i.e., what they are paying for), and in the former case, they are constrained by forces outside of their control. We exploit TCP congestion control dynamics to distinguish these cases for Internet paths that are predominantly TCP traffic. In TCP terms, we re-articulate the question: was a TCP flow bottlenecked by an already congested (possibly interconnect) link, or did it induce congestion in an otherwise idle (possibly a last-mile) link? Srikanth Sundaresan, Mark Allman, Amogh Dhamdhere, K. C. Claffy |
Internet Measurement Conference | 1 |
| 2017 | Challenges in inferring internet congestion using throughput measurementsabstractWe revisit the use of crowdsourced throughput measurements to infer and localize congestion on end-to-end paths, with particular focus on points of interconnections between ISPs. We analyze three challenges with this approach. First, accurately identifying which link on the path is congested requires fine-grained network tomography techniques not supported by existing throughput measurement platforms. Coarse-grained network tomography can perform this link identification under certain topological conditions, but we show that these conditions do not always hold on the global Internet. Second, existing measurement platforms provide limited visibility of paths to popular web content sources, and only capture a small fraction of interconnections between ISPs. Third, crowdsourcing measurements inherently risks sample bias: using measurements from volunteers across the Internet leads to uneven distribution of samples across time of day, access link speeds, and home network conditions. Finally, it is not clear how large a drop in throughput to interpret as evidence of congestion. We investigate these challenges in detail, and offer guidelines for deployment of measurement infrastructure, strategies, and technologies that can address empirical gaps in our understanding of congestion on the Internet. Srikanth Sundaresan, Xiaohong Deng, Danny Lee, Amogh Dhamdhere |
Internet Measurement Conference | 1 |
| 2016 | Do You See What I See? Differential Treatment of Anonymous Users
Sheharbano Khattak, David Fifield, Sadia Afroz 0001, Mobin Javed, Srikanth Sundaresan, Damon McCoy, Vern Paxson, Steven J. Murdoch |
NDSS | 5 |
| 2016 | Home Network or Access Link? Locating Last-Mile Downstream Throughput Bottlenecks
Srikanth Sundaresan, Nick Feamster, Renata Teixeira |
PAM | 1 |
| 2015 | uCap: An Internet Data Management Tool For The HomeabstractInternet Service Providers (ISPs) have introduced "data caps", or quotas on the amount of data that a customer can download during a billing cycle. Under this model, Internet users who reach a data cap can be subject to degraded performance, extra fees, or even temporary interruption of Internet service. For this reason, users need better visibility into and control over their Internet usage to help them understand what uses up data and control how these quotas are reached. In this paper, we present the design and implementation of a tool, called uCap, to help home users manage Internet data. We conducted a field trial of uCap in 21 home networks in three countries and performed an in-depth qualitative study of ten of these homes. We present the results of the evaluation and implications for the design of future Internet data management tools. Marshini Chetty, Hyojoon Kim, Srikanth Sundaresan, Sam Burnett, Nick Feamster, W. Keith Edwards |
CHI | 3 |
| 2015 | Beyond the Radio: Illuminating the Higher Layers of Mobile NetworksabstractCellular network performance is often viewed as primarily dominated by the radio technology. However, reality proves more complex: mobile operators deploy and configure their networks in different ways, and sometimes establish network sharing agreements with other mobile carriers. Moreover, regulators have encouraged newer operational models such as Mobile Virtual Network Operators (MVNOs) to promote competition. In this paper we draw upon data collected by the ICSI Netalyzr app for Android to characterize how operational decisions, such as network configurations, business models, and relationships between operators introduce diversity in service quality and affect user security and privacy. We delve in detail beyond the radio link and into network configuration and business relationships in six countries. We identify the widespread use of transparent middleboxes such as HTTP and DNS proxies, analyzing how they actively modify user traffic, compromise user privacy, and potentially undermine user security. In addition, we identify network sharing agreements between operators, highlighting the implications of roaming and characterizing the properties of MVNOs, including that a majority are simply rebranded versions of major operators. More broadly, our findings highlight the importance of considering higher-layer relationships when seeking to analyze mobile traffic in a sound fashion. Narseo Vallina-Rodriguez, Srikanth Sundaresan, Christian Kreibich, Nicholas Weaver, Vern Paxson |
MobiSys | 2 |
| 2015 | Measuring the Performance of User Traffic in Home Wireless Networks
Srikanth Sundaresan, Nick Feamster, Renata Teixeira |
PAM | 1 |
| 2014 | Locating throughput bottlenecks in home networksabstractWe present a demonstration of WTF (Where's The Fault?), a system that localizes performance problems in home and access networks. We implement WTF as custom firmware that runs in an off-the-shelf home router. WTF uses timing and buffering information from passively monitored traffic at home routers to detect both access link and wireless network bottlenecks. Srikanth Sundaresan, Nick Feamster, Renata Teixeira |
SIGCOMM | 1 |
| 2014 | BISmark: A Testbed for Deploying Measurements and Applications in Broadband Access Networks
Srikanth Sundaresan, Sam Burnett, Nick Feamster, Walter de Donato |
USENIX ATC | 1 |
| 2013 | Peeking behind the NAT: an empirical study of home networksabstractWe present the first empirical study of home network availability, infrastructure, and usage, using data collected from home networks around the world. In each home, we deploy a router with custom firmware to collect information about the availability of home broadband network connectivity, the home network infrastructure (including the wireless connectivity in each home network and the number of devices connected to the network), and how people in each home network use the network. Downtime is more frequent and longer in developing countries---sometimes due to the network, and in other cases because they simply turn their home router off. We also find that some portions of the wireless spectrum are extremely crowded, that diurnal patterns are more pronounced during the week, and that most traffic in home networks is exchanged over a few connections to a small number of domains. Our study is both a preliminary view into many home networks and an illustration of how measurements from a home router can yield significant information about home networks. Sarthak Grover, Mi Seon Park, Srikanth Sundaresan, Sam Burnett, Hyojoon Kim, Bharath Ravi, Nick Feamster |
Internet Measurement Conference | 3 |
| 2013 | Community contribution award - Measuring and mitigating web performance bottlenecks in broadband access networksabstractWe measure Web performance bottlenecks in home broadband access networks and evaluate ways to mitigate these bottlenecks with caching within home networks. We first measure Web performance bottlenecks to nine popular Web sites from more than 5,000 broadband access networks and demonstrate that when the downstream throughput of the access link exceeds about 16 Mbits/s, latency is the main bottleneck for Web page load time. Next, we use a router-based Web measurement tool, Mirage, to deconstruct Web page load time into its constituent components (DNS lookup, TCP connection setup, object download) and show that simple latency optimizations can yield significant improvements in overall page load times. We then present a case for placing a cache in the home network and deploy three common optimizations: DNS caching, TCP connection caching, and content caching. We show that caching only DNS and TCP connections yields significant improvements in page load time, even when the user's browser is already performing similar independent optimizations. Finally, we use traces from real homes to demonstrate how prefetching DNS and TCP connections for popular sites in a home-router cache can achieve faster page load times. Srikanth Sundaresan, Nick Feamster, Renata Teixeira, Nazanin Magharei |
Internet Measurement Conference | 1 |
| 2013 | Web performance bottlenecks in broadband access networksabstractWe present the first large-scale analysis of Web performance bottlenecks as measured from broadband access networks, using data collected from extensive home router deployments. We analyze the limits of throughput on improving Web performance and identify the contribution of critical factors such as DNS lookups and TCP connection establishment to Web page load times. We find that, as broadband speeds continue to increase, other factors such as TCP connection setup time, server response time, and network latency are often dominant performance bottlenecks. Thus, realizing a "faster Web" requires not only higher download throughput, but also optimizations to reduce both client and server-side latency. Srikanth Sundaresan, Nazanin Magharei, Nick Feamster, Renata Teixeira, Sam Crawford |
SIGMETRICS | 1 |
| 2012 | Accelerating last-mile web performance with popularity-based prefetchingabstractNo abstract available. Srikanth Sundaresan, Nazanin Magharei, Nick Feamster, Renata Teixeira |
SIGCOMM | 1 |
| 2011 | Communicating with caps: managing usage caps in home networksabstractAs Internet service providers increasingly implement and impose "usage caps", consumers need better ways to help them understand and control how devices in the home use up the available network resources or available capacity. Towards this goal, we will demonstrate a system that allows users to monitor and manage their usage caps. The system uses the BISMark firmware running on network gateways to collect usage statistics and report them to a logically centralized controller, which displays usage information. The controller allows users to specify policies about how different people, devices, and applications should consume the usage cap; it implements and enforces these policies via a secure OpenFlow control channel to each gateway device. The demonstration will show various use cases, such as limiting the usage of a particular application, visualizing usage statistics, and allowing users within a single household to "trade" caps with one another. Hyojoon Kim, Srikanth Sundaresan, Marshini Chetty, Nick Feamster, W. Keith Edwards |
SIGCOMM | 2 |
| 2011 | Broadband internet performance: a view from the gatewayabstractWe present the first study of network access link performance measured directly from home gateway devices. Policymakers, ISPs, and users are increasingly interested in studying the performance of Internet access links. Because of many confounding factors in a home network or on end hosts, however, thoroughly understanding access network performance requires deploying measurement infrastructure in users' homes as gateway devices. In conjunction with the Federal Communication Commission's study of broadband Internet access in the United States, we study the throughput and latency of network access links using longitudinal measurements from nearly 4,000 gateway devices across 8 ISPs from a deployment of over 4,200 devices. We study the performance users achieve and how various factors ranging from the user's choice of modem to the ISP's traffic shaping policies can affect performance. Our study yields many important findings about the characteristics of existing access networks. Our findings also provide insights into the ways that access network performance should be measured and presented to users, which can help inform ongoing broader efforts to benchmark the performance of access networks. Srikanth Sundaresan, Walter de Donato, Nick Feamster, Renata Teixeira, Sam Crawford, Antonio Pescapè |
SIGCOMM | 1 |
| 2010 | Autonomous traffic engineering with self-configuring topologiesabstractNetwork operators use traffic engineering (TE) to control the flow of traffic across their networks. Existing TE methods require manual configuration of link weights or tunnels, which is difficult to get right, or prior knowledge of traffic demands and hence may not be robust to link failures or traffic fluctuations. We present a self-configuring TE scheme, SculpTE, which automatically adapts the network-layer topology to changing traffic demands. SculpTE is responsive, stable, and achieves excellent load balancing. Srikanth Sundaresan, Cristian Lumezanu, Nick Feamster, Pierre François |
SIGCOMM | 1 |