VLDB 2026 Research / reviewers in the wild / expert
Abhigyan Sharma
dblp:92/9826
· DBLP profile ↗
9ranked-venue papers
5as first author
4since 2021 · last 2026
0009-0006-0368-9829ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detection and Localization of End-to-end Bitflip Errors in Data CentersabstractPacket corrupting devices can cause bitflips and data corruption, and, depending on exactly where the corruption occurs, can evade even strong link-level checks. Even though such instances are probabilistically rare, they are all too common at hyper-scale. At Meta, they have been causing end-to-end errors for years in our datacenter fleet, resulting in dozens of instances of performance impact to services, and rare instances of data corruption. Over the years, we have developed three tools that have progressively improved our ability to localize this problem to specific hosts, rack switches, links, and even specific ports on those switches. They include a fleet wide passive monitoring service to track Layer-4 checksum errors detected by hosts, an active monitoring and fault localization system to probe for corruption on devices and localize them to individual links at any layer of the topology; and a loopback test to reproduce and confirm packet corruption on specific ports on localized switches. Together, they comprise Flipperino, a system that detects and localizes packet corrupting devices, making it easy to identify them for remediation. We share our experience from investigations, and recent results from Flipperino, which has localized devices causing 97% of Layer-4 checksum errors. We have also reproduced corruption on 8 switches in the last 3 months, including 7 rack switches and a fabric switch. Abhigyan Sharma, Neil Spring, Srikanth Sundaresan, Francesco Caggioni, Laurent Virot, Arber Mancaj |
SIGCOMM | 1 |
| 2025 | Congestion Patterns in a Large-scale RDMA DatacenterabstractRDMA datacenters are proliferating to meet the demand of emerging workloads such as AI training and inference as well as distributed storage. This trend has opened up a critical knowledge gap: the traffic characteristics of congestion in these networks remain unknown. We do not know, for example, which layers of the network are the most congested, if the network is load balanced effectively, how long congestion events last, and how accurate existing telemetry systems are in capturing congestion. This paper bridges this gap by investigating congestion in a large-scale RDMA datacenter dedicated to distributed AI training. We provide insights into three specific congestion patterns: (a) location and distribution in the network, (b) burstiness, e.g., the duration and synchrony of bursts, and (c) observability using existing telemetry methods. We show, for instance, that the deployment of Priority Flow Control (PFC) in RDMA networks has shifted the location of congestion one level up: from the edge-host in legacy TCP/IP datacenters to the network core in RDMA datacenters. At the same time, we show that the same protocol enables us to observe and understand congestion better, even bursty events. The findings of this research reveal open challenges for measuring, characterizing, and managing congestion in RDMA networks, paving the way for future research. Soudeh Ghorbani, Yimeng Zhao, Srikanth Sundaresan, Ying Zhang 0022, Yijing Zeng, Abhigyan Sharma, Prashanth Kannan, Cristian Lumezanu |
IMC | 6 |
| 2023 | Modeling and Generating Control-Plane Traffic for Cellular NetworksabstractWith 5G deployment gaining momentum, the control-plane traffic volume of cellular networks is escalating. Such rapid traffic growth motivates the need to study the mobile core network (MCN) control-plane design and performance optimization. Doing so requires realistic, large control-plane traffic traces in order to profile and debug the mobile network performance under real workload. However, large-scale control-plane traffic traces are not made available to the public by mobile operators due to business and privacy concerns. As such, it is critically important to develop accurate, scalable, versatile, and open-to-innovation control traffic generators, which in turn critically rely on an accurate traffic model for the control plane. Developing an accurate model of control-plane traffic faces several challenges: (1) how to capture the dependence among the control events generated by each User Equipment (UE), (2) how to model the inter-arrival time and sojourn time of control events of individual UEs, and (3) how to capture the diversity of control-plane traffic across UEs. We present a novel two-level hierarchical state-machine-based control-plane traffic model. We further show how our model can be easily adjusted from LTE to NextG networks (e.g., 5G) to support modeling future control-plane traffic. We experimentally validate that the proposed model can generate large realistic control-plane traffic traces. We have open-sourced our traffic generator to the public to foster MCN research. Jiayi Meng, Jingqi Huang, Y. Charlie Hu, Yaron Koral, Xiaojun Lin 0001, Muhammad Shahbaz 0001, Abhigyan Sharma |
IMC | 7 |
| 2021 | Demo: SkyRoute, a Fast and Realistic UAV Cellular Simulation FrameworkabstractThere is a growing interest in reusing cellular base stations on the ground to provide long range, high-speed wireless connectivity to UAVs. Towards this goal, we present SkyRoute – a novel and powerful simulation platform for rapid and realistic assessment of UAV cellular connectivity. SkyRoute combines real base station locations and antenna data with a lightweight version of the widely-used ns-3 simulation platform for full-stack wireless channel and cellular network simulation. As an exemplary application, we demonstrate realistic coverage and cell selection prediction in a large metropolitan area. Mingsheng Yin, Tuyen X. Tran, Abhigyan Sharma, Marco Mezzavilla, Sundeep Rangan |
ICNP | 3 |
| 2020 | A Study of Network-Side 5G User Localization Using Angle-Based FingerprintsabstractThis paper explores network-side cellular user localization using fingerprints created from the angle measurements enabled by 5G. Our key idea is a binning-based fingerprinting technique that leverages multipath propagation to create fingerprint vectors based on angles of arrival of signals along multiple paths at each user. In network simulations that recreate urban environments with 3D building geometry and base station locations for a major city, our binning-based fingerprinting for 5G achieves significantly lower localization errors with a single base station than signal strength-based fingerprinting for LTE. Jiayi Meng, Abhigyan Sharma, Tuyen X. Tran, Bharath Balasubramanian, Gueyoung Jung, Matti A. Hiltunen, Y. Charlie Hu |
LANMAN | 2 |
| 2019 | Switchboard: A Middleware for Wide-Area Service ChainingabstractProduction networks are transitioning from the use of physical middleboxes to virtual network functions (VNFs), which makes it easy to construct highly-customized service chains of VNFs dynamically using software. Wide-area service chains are increasingly important given the emergence of heterogeneous execution platforms consisting of customer premise equipment (CPE), small edge cloud sites, and large centralized cloud data centers, since only part of the service chain can be deployed at the CPE and even the closest edge site may not always be able to process all the customers' traffic. Switchboard is a middleware for realizing and managing such an ecosystem of diverse VNFs and cloud platforms. It exploits principles from service-oriented architectures to treat VNFs as independent services, and provides a traffic routing platform shared by all VNFs. Moreover, Switchboard's global controller optimizes wide-area routes based on a holistic view of customer traffic as well as the resources available at VNFs and the underlying network. Switchboard globally optimized routes achieve up to 57% higher throughput and 49% lower latency than a distributed load balancing approach in a wide-area testbed. Its routing platform supports line-rate traffic with millions of concurrent flows. Abhigyan Sharma, Yoji Ozawa, Matti A. Hiltunen, Kaustubh R. Joshi, Richard D. Schlichting, Zhaoyu Gao |
Middleware | 1 |
| 2014 | A global name service for a highly mobile internetworkabstractMobile devices dominate the Internet today, however the Internet rooted in its tethered origins continues to provide poor infrastructure support for mobility. Our position is that in order to address this problem, a key challenge that must be addressed is the design of a massively scalable global name service that rapidly resolves identities to network locations under high mobility. Our primary contribution is the design, implementation, and evaluation of auspice, a next-generation global name service that addresses this challenge. A key insight underlying auspice is a demand-aware replica {placement engine} that intelligently replicates name records to provide low lookup latency, low update cost, and high availability. We have implemented a prototype of auspice and compared it against several commercial managed DNS providers as well as state-of-the-art research alternatives, and shown that auspice significantly outperforms both. We demonstrate proof-of-concept that auspice can serve as a complete end-to-end mobility solution as well as enable novel context-based communication primitives that generalize name- or address-based communication in today's Internet. Abhigyan Sharma, Xiaozheng Tie, Hardeep Uppal, Arun Venkataramani, David Westbrook, Aditya Yadav |
SIGCOMM | 1 |
| 2013 | Distributing content simplifies ISP traffic engineeringabstractSeveral major Internet service providers today also offer content distribution services. The emergence of such "network-CDNs" (NCDNs) is driven both by market forces as well as the cost of carrying ever-increasing volumes of traffic across their backbones. An NCDN has the flexibility to determine both where content is placed and how traffic is routed within the network. However NCDNs today continue to treat traffic engineering independently from content placement and request redirection decisions. In this paper, we investigate the interplay between content distribution strategies and traffic engineering and ask whether or how an NCDN should address these concerns in a joint manner. Our experimental analysis, based on traces from a large content distribution network and real ISP topologies, shows that realistic (i.e., history-based) joint optimization strategies offer little benefit (and often significantly underperform) compared to simple and "unplanned" strategies for routing and placement such as InverseCap and LRU. We also find that the simpler strategies suffice to achieve network cost and user-perceived latencies close to those of a joint-optimal strategy with future knowledge. Abhigyan Sharma, Arun Venkataramani, Ramesh K. Sitaraman |
SIGMETRICS | 1 |
| 2011 | Beyond MLU: An application-centric comparison of traffic engineering schemesabstractTraffic engineering (TE) has been long studied as a network optimization problem, but its impact on user-perceived application performance has received little attention. Our paper takes a first step to address this disparity. Using real traffic matrices and topologies from three ISPs, we conduct very large-scale experiments simulating ISP traffic as an aggregate of a large number of TCP flows. Our application-centric, empirical approach yields two rather unexpected findings. First, link utilization metrics, and MLU in particular, are poor predictors of application performance. Despite significant differences in MLU, all TE schemes and even a static shortest-path routing scheme achieve nearly identical application performance. Second, application adaptation in the form of location diversity, i.e., the ability to download content from multiple potential locations, significantly improves the capacity achieved by all schemes. Even the ability to download from just 2-4 locations enables all TE schemes to achieve near-optimal capacity, and even static routing to be within 30% of optimal. Our findings call into question the value of TE as practiced today, and compel us to significantly rethink the TE problem in the light of application adaptation. Abhigyan Sharma, Aditya Kumar Mishra, Arun Venkataramani |
INFOCOM | 1 |