EDBT 2026 Demo / reviewers in the wild / expert
Srinivas Narayana
dblp:03/1245
· DBLP profile ↗
25ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-1128-477XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Lamport, Towards Probabilistic Fair OrderingabstractA growing class of applications demands fair ordering of events, which ensures that events generated earlier are processed before later events. However, achieving such sequencing is challenging due to the inherent errors in clock synchronization: two events at two clients generated close together may have timestamps that cannot be compared confidently. We advocate for an approach that embraces, rather than eliminates, clock synchronization errors. Instead of attempting to remove the error from a timestamp, Tommy, our proposed system, leverages a statistical model to compare two noisy timestamps probabilistically by learning per-clock synchronization error distributions. Our preliminary statistical model computes the probability that one event precedes another by only relying on local clocks of clients. This serves as a foundation for a new relation: likely-happened-before denoted by →p where p represents the probability that an event happened before another. The →p relation provides a basis for ordering multiple events which are otherwise considered concurrent by Lamport's happened-before (→) relation. We highlight various related challenges including the intransitivity of the →p relation as opposed to the transitive → relation. We outline several research directions: online fair sequencing, stochastically fair total ordering, and handling byzantine clients. Jinkun Geng, Radhika Mittal, Aurojit Panda, Srinivas Narayana, Anirudh Sivaraman |
HotNets | 5 |
| 2025 | State-Compute Replication: Parallelizing High-Speed Stateful Packet Processing
Qiongwen Xu, Sebastiano Miano, Tao Wang 0088, Adithya Murugadass, Songyuan Zhang, Anirudh Sivaraman, Gianni Antichi, Srinivas Narayana |
NSDI | 9 |
| 2025 | Comparing the Precision of Abstract Operators in the eBPF Verifier Using Differential Synthesis
Matan Shachnai, Harishankar Vishwanathan, Srinivas Narayana, Santosh Nagarakatte |
SAS | 3 |
| 2025 | ParserHawk: Hardware-aware parser generator using program synthesisabstractParser programs are becoming increasingly complex to accommodate intricate network packet formats and advanced protocols. Existing parser compilers incorporate predefined program rewrite rules to output the low-level parser implementation. Yet, these rules are often brittle and sensitive to how the input parser program is written. As a result, generated implementations could consume more hardware resources than necessary. In some cases, these compilers unnecessarily reject valid parser programs that could have fit within the target device parser's resource constraints. Karan Kumar G., Ennan Zhai, Bili Dong, Joseph Tassarotti, Srinivas Narayana, Anirudh Sivaraman |
SIGCOMM | 8 |
| 2025 | Network Support For Scalable And High Performance Cloud ExchangesabstractFinancial exchanges are migrating to the public cloud, but the best-effort nature of the cloud fabric is at odds with the stringent networking requirements of the exchanges. We present Onyx, a system for meeting such requirements which uses many well-studied techniques in a new context as well as introduces new techniques that enable a scalable cloud financial exchange. An overlay multicast tree is used to disseminate data to 1000 participants with ≤ 1 μs difference in data reception time between any two participants, crucial for maintaining fair competition. Several techniques for mitigating latency variance are introduced. Onyx also presents a scheduling policy for trade orders that enhances an exchange's performance and gracefully services bursty traffic. Onyx achieves ≈50% lower latency than the AWS multicast service [1]. Onyx outperforms an existing system, CloudEx [2] in terms of supported number of participants, exchange's throughput and multicast latency. Onyx's techniques can be applied to other existing systems (e.g., DBO) to enhance their performance. Jinkun Geng, Daniel Duclos-Cavalcanti, Xiyu Hao, Ulysses Butler, Radhika Mittal, Srinivas Narayana, Anirudh Sivaraman |
SIGCOMM | 7 |
| 2024 | Cross-Platform Transpilation of Packet-Processing Programs using Program SynthesisabstractThe proliferation of programmable network devices offers a wide range of device options for developers of packet processing programs. However, there are several differences in programming language usage, hardware resource constraints, and hardware architecture across these devices. Programmers must understand multiple programming languages and hardware designs to write programs for various devices. Karan Kumar G., Ennan Zhai, Srinivas Narayana, Anirudh Sivaraman |
APNet | 5 |
| 2024 | Fixing Latent Unsound Abstract Operators in the eBPF Verifier of the Linux Kernel
Matan Shachnai, Harishankar Vishwanathan, Srinivas Narayana, Santosh Nagarakatte |
SAS | 3 |
| 2023 | CaT: A Solver-Aided Compiler for Packet-Processing PipelinesabstractCompiling high-level programs to high-speed packet-processing pipelines is a challenging combinatorial optimization problem. The compiler must configure the pipeline’s resources to match the semantics of the program’s high-level specification, while packing all of the program’s computation into the pipeline’s limited resources. State of the art approaches tackle individual aspects of this problem. Yet, they miss opportunities to produce globally high-quality outcomes within reasonable compilation times. We develop a framework to decompose the compilation problem for such pipelines into three phases—making extensive use of solver engines (e.g., ILP, SMT, and program synthesis) to simplify the development of these phases. Transformation rewrites programs to use more abundant pipeline resources, avoiding scarce ones. Synthesis breaks complex transactional code into configurations of pipelined compute units. Allocation maps the program’s compute and memory to the pipeline’s hardware resources. We prototype these ideas in a compiler, CaT, which targets (1) the Tofino programmable switch pipeline and (2) Menshen, a cycle-accurate simulator of a Verilog description of the RMT pipeline. CaT can handle programs that existing compilers cannot currently run on pipelines and generates code faster than existing compilers, where the generated code uses fewer pipeline resources. Divya Raghunathan, Ruijie Fang, Tao Wang 0088, Xiaotong Zhu, Anirudh Sivaraman, Srinivas Narayana, Aarti Gupta |
ASPLOS (3) | 7 |
| 2023 | Verifying the Verifier: eBPF Range Analysis VerificationabstractAbstract This paper proposes an automated method to check the correctness of range analysis used in the Linux kernel ’s eBPF verifier. We provide the specification of soundness for range analysis performed by the eBPF verifier. We automatically generate verification conditions that encode the operation of the eBPF verifier directly from the Linux kernel ’s C source code and check it against our specification. When we discover instances where the eBPF verifier is unsound, we propose a method to generate an eBPF program that demonstrates the mismatch between the abstract and the concrete semantics. Our prototype automatically checks the soundness of 16 versions of the eBPF verifier in the Linux kernel versions ranging from 4.14 to 5.19. In this process, we have discovered new bugs in older versions and proved the soundness of range analysis in the latest version of the Linux kernel. Harishankar Vishwanathan, Matan Shachnai, Srinivas Narayana, Santosh Nagarakatte |
CAV (3) | 3 |
| 2022 | Sound, Precise, and Fast Abstract Interpretation with Tristate NumbersabstractExtended Berkeley Packet Filter (BPF) is a language and run-time system that allows non-superusers to extend the Linux and Windows operating systems by downloading user code into the kernel. To ensure that user code is safe to run in kernel context, BPF relies on a static analyzer that proves properties about the code, such as bounded memory access and the absence of operations that crash. The BPF static analyzer checks safety using abstract interpretation with several abstract domains. Among these, the domain of tnums (tristate numbers) is a key domain used to reason about the bitwise uncertainty in program values. This paper formally specifies the tnum abstract domain and its arithmetic operators. We provide the first proofs of soundness and optimality of the abstract arithmetic operators for tnum addition and subtraction used in the BPF analyzer. Further, we describe a novel sound algorithm for multiplication of tnums that is more precise and efficient (runs 33% faster on average) than the Linux kernel’s algorithm. Our tnum multiplication is now merged in the Linux kernel. Harishankar Vishwanathan, Matan Shachnai, Srinivas Narayana, Santosh Nagarakatte |
CGO | 3 |
| 2022 | Load balancers need in-band feedback controlabstractServer load balancers (LBs) are critical components of interactive services, routing client requests to servers in a pool. LBs improve service performance and increase availability by spreading the request load evenly across servers. Bhavana Vannarth Shobhana, Srinivas Narayana, B. R. Badrinath |
HotNets | 2 |
| 2022 | Privid: Practical, Privacy-Preserving Video Analytics Queries
Frank Cangialosi, Neil Agarwal, Venkat Arun, Junchen Jiang, Srinivas Narayana, Anand D. Sarwate, Ravi Netravali |
NSDI | 5 |
| 2022 | Elasticity detection: a building block for internet congestion controlabstractThis paper introduces a new metric, "elasticity," which characterizes the nature of cross-traffic competing with a flow. Elasticity captures whether the cross traffic reacts to changes in available bandwidth. We show that it is possible to robustly detect the elasticity of cross traffic at a sender without router support, and that elasticity detection can reduce delays in the Internet by enabling delay-controlling congestion control protocols to be deployed without hurting flow throughput. Our results show that the proposed method achieves more than 85% accuracy under a variety of network conditions, and that congestion control using elasticity detection achieves throughput comparable to Cubic but with delays that are 50--70 ms lower when cross traffic is inelastic. Prateesh Goyal, Akshay Narayan 0001, Frank Cangialosi, Srinivas Narayana, Mohammad Alizadeh, Hari Balakrishnan |
SIGCOMM | 4 |
| 2021 | Snicket: Query-Driven Distributed TracingabstractIncreasing application complexity has caused applications to be refactored into smaller components known as microservices that communicate with each other using RPCs. Distributed tracing has emerged as an important debugging tool for such microservice-based applications. Distributed tracing follows the journey of a user request from its starting point at the application's front-end, through RPC calls made by the front-end to different microservices recursively, all the way until a response is constructed and sent back to the user. To reduce storage costs, distributed tracing systems sample traces before collecting them for subsequent querying, affecting the accuracy of queries on the collected traces. Jessica Berg, Fabian Ruffy, Khanh Nguyen 0001, Nicholas Yang, Anirudh Sivaraman, Ravi Netravali, Srinivas Narayana |
HotNets | 8 |
| 2021 | Synthesizing safe and efficient kernel extensions for packet processingabstractExtended Berkeley Packet Filter (BPF) has emerged as a powerful method to extend packet-processing functionality in the Linux operating system. BPF allows users to write code in high-level languages (like C or Rust) and execute them at specific hooks in the kernel, such as the network device driver. To ensure safe execution of a user-developed BPF program in kernel context, Linux uses an in-kernel static checker. The checker allows a program to execute only if it can prove that the program is crash-free, always accesses memory within safe bounds, and avoids leaking kernel data. Qiongwen Xu, Michael D. Wong, Tanvi Wagle, Srinivas Narayana, Anirudh Sivaraman |
SIGCOMM | 4 |
| 2021 | Automated SmartNIC Offloading Insights for Network FunctionsabstractThe gap between CPU and networking speeds has motivated the development of SmartNICs for NF (network functions) offloading. However, offloading performance is predicated upon intricate knowledge about SmartNIC hardware and careful hand-tuning of the ported programs. Today, developers cannot easily reason about the offloading performance or the effectiveness of different porting strategies without resorting to a trial-and-error approach. Yiming Qiu 0001, Jiarong Xing, Kuo-Feng Hsu, Qiao Kang, Ming Liu 0027, Srinivas Narayana, Ang Chen 0001 |
SOSP | 6 |
| 2020 | Switch Code Generation Using Program SynthesisabstractWriting packet-processing programs for programmable switch pipelines is challenging because of their all-or-nothing nature: a program either runs at line rate if it can fit within pipeline resources, or does not run at all. It is the compiler's responsibility to fit programs into pipeline resources. However, switch compilers, which use rewrite rules to generate switch machine code, often reject programs because the rules fail to transform programs into a form that can be mapped to a pipeline's limited resources---even if a mapping actually exists. Michael D. Wong, Divya Raghunathan, Aatish Kishan Varma, Pravein G. Kannan, Anirudh Sivaraman, Srinivas Narayana, Aarti Gupta |
SIGCOMM | 8 |
| 2019 | Autogenerating Fast Packet-Processing Code Using Program SynthesisabstractPacket-processing code should be fast. But, it is hard to write fast code for programmable substrates such as high-speed switches, multicore SoC SmarfNICs, FP-GAs, middleboxes, and the end-host stack. Today, expert developers with deep familiarity with the underlying hardware handcraft such code. Making things worse, building optimizing compilers for these substrates requires significant development effort, which may not be available for these new, niche, and evolving substrates. Aatish Kishan Varma, Anirudh Sivaraman, Srinivas Narayana |
HotNets | 5 |
| 2018 | Restructuring endpoint congestion controlabstractThis paper describes the implementation and evaluation of a system to implement complex congestion control functions by placing them in a separate agent outside the datapath. Each datapath---such as the Linux kernel TCP, UDP-based QUIC, or kernel-bypass transports like mTCP-on-DPDK---summarizes information about packet round-trip times, receptions, losses, and ECN via a well-defined interface to algorithms running in the off-datapath Congestion Control Plane (CCP). The algorithms use this information to control the datapath's congestion window or pacing rate. Algorithms written in CCP can run on multiple datapaths. CCP improves both the pace of development and ease of maintenance of congestion control algorithms by providing better, modular abstractions, and supports aggregation capabilities of the Congestion Manager, all with one-time changes to datapaths. CCP also enables new capabilities, such as Copa in Linux TCP, several algorithms running on QUIC and mTCP/DPDK, and the use of signal processing algorithms to detect whether cross-traffic is ACK-clocked. Experiments with our user-level Linux CCP implementation show that CCP algorithms behave similarly to kernel algorithms, and incur modest CPU overhead of a few percent. Akshay Narayan 0001, Frank Cangialosi, Deepti Raghavan, Prateesh Goyal, Srinivas Narayana, Radhika Mittal, Mohammad Alizadeh, Hari Balakrishnan |
SIGCOMM | 5 |
| 2017 | The Case for Moving Congestion Control Out of the DatapathabstractWith Moore's law ending, the gap between general-purpose processor speeds and network link rates is widening. This trend has led to new packet-processing "datapaths" in endpoints, including kernel bypass software and emerging SmartNIC hardware. In addition, several applications are rolling out their own protocols atop UDP (e.g., QUIC, WebRTC, Mosh, etc.), forming new datapaths different from the traditional kernel TCP stack. All these datapaths require congestion control, but they must implement it separately because it is not possible to reuse the kernel's TCP implementations. This paper proposes moving congestion control from the datapath into a separate agent. This agent, which we call the congestion control plane (CCP), must provide both an expressive congestion control API as well as a specification for datapath designers to implement and deploy CCP. We propose an API for congestion control, datapath primitives, and a user-space agent design that uses a batching method to communicate with the datapath. Our approach promises to preserve the behavior and performance of in-datapath implementations while making it significantly easier to implement and deploy new congestion control algorithms. Akshay Narayan 0001, Frank Cangialosi, Prateesh Goyal, Srinivas Narayana, Mohammad Alizadeh, Hari Balakrishnan |
HotNets | 4 |
| 2017 | Language-Directed Hardware Design for Network Performance MonitoringabstractNetwork performance monitoring today is restricted by existing switch support for measurement, forcing operators to rely heavily on endpoints with poor visibility into the network core. Switch vendors have added progressively more monitoring features to switches, but the current trajectory of adding specific features is unsustainable given the ever-changing demands of network operators. Instead, we ask what switch hardware primitives are required to support an expressive language of network performance questions. We believe that the resulting switch hardware design could address a wide variety of current and future performance monitoring needs. Srinivas Narayana, Anirudh Sivaraman, Vikram Nathan, Prateesh Goyal, Venkat Arun, Mohammad Alizadeh, Vimalkumar Jeyakumar, Changhoon Kim |
SIGCOMM | 1 |
| 2016 | Hardware-Software Co-Design for Network Performance MeasurementabstractDiagnosing performance problems in networks is important, for example to determine where packets experience high latency or loss. However, existing performance diagnoses are constrained by limited switch mechanisms for measurement. Alternatively, operators use endpoint information indirectly to infer root causes for problematic latency or drops. Srinivas Narayana, Anirudh Sivaraman, Vikram Nathan, Mohammad Alizadeh, David Walker 0001, Jennifer Rexford, Vimalkumar Jeyakumar, Changhoon Kim |
HotNets | 1 |
| 2016 | Compiling Path Queries
Srinivas Narayana, Mina Tahmasbi Arashloo, Jennifer Rexford, David Walker 0001 |
NSDI | 1 |
| 2013 | Abstractions for model checking SDN controllers
Divjyot Sethi, Srinivas Narayana, Sharad Malik |
FMCAD | 2 |
| 2012 | Distributed wide-area traffic management for cloud servicesabstractThe performance of interactive cloud services depends heavily on which data centers handle client requests, and which wide-area paths carry traffic. While making these decisions, cloud service providers also need to weigh operational considerations like electricity and bandwidth costs, and balancing server loads across replicas. We argue that selecting data centers and network routes independently, as is common in today's services, can lead to much lower performance or higher costs than a coordinated decision. However, fine-grained joint control of two large distributed systems---e.g., DNS-based replica-mapping and data center multi-homed routing---can be administratively challenging. In this paper, we introduce the design of a system that jointly optimizes replica-mapping and multi-homed routing, while retaining the functional separation that exists between them today. We show how to construct a provably optimal distributed solution implemented through local computations and message exchanges between the mapping and routing systems. Srinivas Narayana, Wenjie Jiang 0001, Jennifer Rexford, Mung Chiang |
SIGMETRICS | 1 |