Nandita Dukkipati

dblp:74/5250 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0003-2926-1482ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 21 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 CSIG: Congestion Signaling for Datacenter Transports
abstract
Optimizing burst-heavy datacenter workloads necessitates finegrained network control and visibility. We introduce CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header. The architecture captures μsgranularity switch metrics, such as available bandwidth, and signals them to end-hosts using in-band, line-rate operations. We propose Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production. Beyond transport-level performance, CSIG enables flow-aware observability by embedding μs-scale metrics into every packet, allowing individual application transfers to pinpoint their bottleneck location, such as the topology tier limiting their performance. CSIG thus transforms network telemetry from post-hoc correlation into a real time, context-aware capability. We demonstrate CSIG's broad deployability by validating it across five generations of commodity switch hardware (up to 102.4 Tbps), four NIC generations, and five transport stacks. Our design proves that a streamlined Layer 2 approach, focusing exclusively on the principal path bottleneck, provides transport-agnostic gains without requiring forklift hardware upgrades.
Abhiram Ravi, Nandita Dukkipati, Weiwu Pang, Neal Cardwell, Brad Karp, Mohammad Jafar Akhbarizadeh, Weida Huang, Konstantinos Prasopoulos, Kok-Kiong Yap, Amin Vahdat
SIGCOMM2
2025 Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski, Liqun Cheng, Amin Vahdat
NSDI7
2025 Learnings from Deploying Network QoS Alignment to Application Priorities for Storage Services
Matthew Buckley, Parsa Pazhooheshy, Z. Morley Mao, Nandita Dukkipati, Hamid Hajabdolali Bazzaz, Priyaranjan Jha, Yingjie Bi, Steve Middlekauff, Yashar Ganjali
NSDI4
2025 Firefly: Scalable, Ultra-Accurate Clock Synchronization for Datacenters
abstract
Cloud-based financial exchanges require sub-10ns device-to-device clock synchronization accuracy while adhering to Coordinated Universal Time (UTC). Existing clock sync techniques struggle to meet this demand at scale and are vulnerable to clock drift, jitter, and path asymmetries. Firefly, a software-driven datacenter clock sync system, scalably, cost-effectively, and reliably achieves very high clock sync accuracy. It employs a distributed consensus algorithm on a random overlay graph to rapidly converge to a common time while applying gradual adjustments to device hardware clocks. To realize consistent sync-to-UTC (external sync) across devices while maintaining a stable device-to-device internal sync, Firefly uses a novel technique, layered synchronization, that decouples internal and external syncs. In a 248-machine Clos network, Firefly achieves sub-10ns device-to-device and ≤1μs device-to-UTC sync, and is resilient to time server failure and unstable clocks.
Pooria Namyar, Nandita Dukkipati, KK Yap, Junzhi Gong, Peixuan Gao, Devdeep Ray, Gautam Kumar 0001, Ramesh Govindan, Amin Vahdat
SIGCOMM4
2025 Falcon: A Reliable, Low Latency Hardware Transport
abstract
Hardware transports such as RoCE deliver high performance with minimal host CPU, but are best suited to special-purpose deployments that limit their use, e.g., backend networks or Ethernet with Priority Flow Control (PFC). We introduce Falcon, the first hardware transport that supports multiple Upper Layer Protocols (ULPs) and heterogeneous application workloads in general-purpose Ethernet datacenter environments (with losses and without special switch support). Key design elements include: delay-based congestion control with multipath load balancing; a layered design with a simple request-response transaction interface for multi-ULP support; hardware-based retransmissions and error-handling for scalability; and a programmable engine for flexibility. The first Falcon hardware implementation delivers a peak performance of 200 Gbps, 120 Mops/sec, with near-optimal operation completion times that are up to 8× lower than CX-7 RoCE under network congestion, and up to 65% higher goodput under lossy conditions.
Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan M. G. Wassel, Naveen Kr. Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar 0003, Behnam Montazeri, Neelesh Bansod, Sarin Thomas, Inho Cho, Hyojeong Lee Seibert, Baijun Wu, Rui Yang 0034, Qianwen Yin, Srinivas Vaduvatha, Weihuang Wang, Masoud Moshref, David Wetherall, Amin Vahdat
SIGCOMM2
2023 Bolt: Sub-RTT Congestion Control for Ultra-Low Latency
Serhat Arslan, Gautam Kumar 0001, Nandita Dukkipati
NSDI4
2023 Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INT
Masoud Moshref, Gautam Kumar 0001, T. S. Eugene Ng, Neal Cardwell, Nandita Dukkipati
NSDI7
2022 Aequitas: admission control for performance-critical RPCs in datacenters
abstract
With the increasing popularity of disaggregated storage and microservice architectures, high fan-out and fan-in Remote Procedure Calls (RPCs) now generate most of the traffic in modern datacenters. While the network plays a crucial role in RPC performance, traditional traffic classification categories cannot sufficiently capture their importance due to wide variations in RPC characteristics. As a result, meeting service-level objectives (SLOs), especially for performance-critical (PC) RPCs, remains challenging.
Yiwen Zhang 0008, Gautam Kumar 0001, Nandita Dukkipati, Xian Wu 0001, Priyaranjan Jha, Mosharaf Chowdhury, Amin Vahdat
SIGCOMM3
2020 Sundial: Fault-tolerant Clock Synchronization for Datacenters
Gautam Kumar 0001, Hema Hariharan, Hassan M. G. Wassel, Peter Hochschild, Dave Platt, Simon L. Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, Amin Vahdat
OSDI9
2020 Swift: Delay is Simple and Effective for Congestion Control in the Datacenter
abstract
We report on experiences with Swift congestion control in Google datacenters. Swift targets an end-to-end delay by using AIMD control, with pacing under extreme congestion. With accurate RTT measurement and care in reasoning about delay targets, we find this design is a foundation for excellent performance when network distances are well-known. Importantly, its simplicity helps us to meet operational challenges. Delay is easy to decompose into fabric and host components to separate concerns, and effortless to deploy and maintain as a congestion signal while the datacenter evolves. In large-scale testbed experiments, Swift delivers a tail latency of <50μs for short RPCs, with near-zero packet drops, while sustaining ~100Gbps throughput per server. This is a tail of <3x the minimal latency at a load close to 100%. In production use in many different clusters, Swift achieves consistently low tail completion times for short RPCs, while providing high throughput for long RPCs. It has loss rates that are at least 10x lower than a DCTCP protocol, and handles O(10k) incasts that sharply degrade with DCTCP.
Gautam Kumar 0001, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel, Xian Wu 0001, Behnam Montazeri, Yaogong Wang, Kevin Springborn, Christopher Alfeld, Michael Ryan, David Wetherall, Amin Vahdat
SIGCOMM2
2019 Eiffel: Efficient and Flexible Software Packet Scheduling
Ahmed Saeed 0001, Yimeng Zhao, Nandita Dukkipati, Ellen Zegura, Mostafa H. Ammar, Khaled A. Harras, Amin Vahdat
NSDI3
2019 PicNIC: predictable virtualized NIC
abstract
Network virtualization stacks are the linchpins of public clouds. A key goal is to provide performance isolation so that workloads on one Virtual Machine (VM) do not adversely impact the network experience of another VM. Using data from a major public cloud provider, we systematically characterize how performance isolation can break in current virtualization stacks and find a fundamental tradeoff between isolation and resource multiplexing for efficiency. In order to provide predictable performance, we propose a new system called PicNIC that shares resources efficiently in the common case while rapidly reacting to ensure isolation. PicNIC builds on three constructs to quickly detect isolation breakdown and to enforce it when necessary: CPU-fair weighted fair queues at receivers, receiver-driven congestion control for backpressure, and sender-side admission control with shaping. Based on an extensive evaluation, we show that this combination ensures isolation for VMs at sub-millisecond timescales with negligible overhead.
Praveen Kumar 0003, Nandita Dukkipati, Nathan Lewis, Yaogong Wang, Chonggang Li, Valas Valancius, Jake Adriaens, Steve D. Gribble, Nate Foster, Amin Vahdat
SIGCOMM2
2019 Snap: a microkernel approach to host networking
abstract
This paper presents our design and experience with a microkernel-inspired approach to host networking called Snap. Snap is a userspace networking system that supports Google's rapidly evolving needs with flexible modules that implement a range of network functions, including edge packet switching, virtualization for our cloud platform, traffic shaping policy enforcement, and a high-performance reliable messaging and RDMA-like service. Snap has been running in production for over three years, supporting the extensible communication needs of several large and critical systems.
Michael Marty, Marc de Kruijf, Jacob Adriaens, Christopher Alfeld, Sean Bauer, Carlo Contavalli, Michael Dalton, Nandita Dukkipati, William C. Evans, Steve D. Gribble, Nicholas Kidd, Roman Kononov, Gautam Kumar 0001, Carl Mauer, Emily Musick, Lena E. Olson, Erik Rubow, Michael Ryan, Kevin Springborn, Valas Valancius, Amin Vahdat
SOSP8
2017 Carousel: Scalable Traffic Shaping at End Hosts
abstract
Traffic shaping, including pacing and rate limiting, is fundamental to the correct and efficient operation of both datacenter and wide area networks. Sample use cases include policy-based bandwidth allocation to flow aggregates, rate-based congestion control algorithms, and packet pacing to avoid bursty transmissions that can overwhelm router buffers. Driven by the need to scale to millions of flows and to apply complex policies, traffic shaping is moving from network switches into the end hosts, typically implemented in software in the kernel networking stack.
Ahmed Saeed 0001, Nandita Dukkipati, Vytautas Valancius, Vinh The Lam, Carlo Contavalli, Amin Vahdat
SIGCOMM2
2015 TIMELY: RTT-based Congestion Control for the Datacenter
abstract
Datacenter transports aim to deliver low latency messaging together with high throughput. We show that simple packet delay, measured as round-trip times at hosts, is an effective congestion signal without the need for switch feedback. First, we show that advances in NIC hardware have made RTT measurement possible with microsecond accuracy, and that these RTTs are sufficient to estimate switch queueing. Then we describe how TIMELY can adjust transmission rates using RTT gradients to keep packet latency low while delivering high bandwidth. We implement our design in host software running over NICs with OS-bypass capabilities. We show using experiments with up to hundreds of machines on a Clos network topology that it provides excellent performance: turning on TIMELY for OS-bypass messaging over a fabric with PFC lowers 99 percentile tail latency by 9X while maintaining near line-rate throughput. Our system also outperforms DCTCP running in an optimized kernel, reducing tail latency by $13$X. To the best of our knowledge, TIMELY is the first delay-based congestion control protocol for use in the datacenter, and it achieves its results despite having an order of magnitude fewer RTT signals (due to NIC offload) than earlier delay-based schemes such as Vegas.
Radhika Mittal, Vinh The Lam, Nandita Dukkipati, Emily R. Blem, Hassan M. G. Wassel, Manya Ghobadi, Amin Vahdat, Yaogong Wang, David Wetherall, David Zats
SIGCOMM3
2013 Reducing web latency: the virtue of gentle aggression
abstract
To serve users quickly, Web service providers build infrastructure closer to clients and use multi-stage transport connections. Although these changes reduce client-perceived round-trip times, TCP's current mechanisms fundamentally limit latency improvements. We performed a measurement study of a large Web service provider and found that, while connections with no loss complete close to the ideal latency of one round-trip time, TCP's timeout-driven recovery causes transfers with loss to take five times longer on average.
Tobias Flach, Nandita Dukkipati, Andreas Terzis, Barath Raghavan, Neal Cardwell, Yuchung Cheng, Shuai Hao 0002, Ethan Katz-Bassett, Ramesh Govindan
SIGCOMM2
2013 packetdrill: Scriptable Network Stack Testing, from Sockets to Packets
Neal Cardwell, Yuchung Cheng, Lawrence Brakmo, Matthew Mathis, Barath Raghavan, Nandita Dukkipati, Hsiao-Keng Jerry Chu, Andreas Terzis, Tom Herbert
USENIX ATC6
2011 Proportional rate reduction for TCP
abstract
Packet losses increase latency for Web users. Fast recovery is a key mechanism for TCP to recover from packet losses. In this paper, we explore some of the weaknesses of the standard algorithm described in RFC 3517 and the non-standard algorithms implemented in Linux. We find that these algorithms deviate from their intended behavior in the real world due to the combined effect of short flows, application stalls, burst losses, acknowledgment (ACK) loss and reordering, and stretch ACKs. Linux suffers from excessive congestion window reductions while RFC 3517 transmits large bursts under high losses, both of which harm the rest of the flow and increase Web latency.
Nandita Dukkipati, Matthew Mathis, Yuchung Cheng, Manya Ghobadi
Internet Measurement Conference1
2011 Layered Internet Video Adaptation (LIVA): Network-Assisted Bandwidth Sharing and Transient Loss Protection for Video Streaming
abstract
As video traffic increases in the Internet and competes for limited bandwidth resources, it is important to design bandwidth-sharing and loss-protection schemes that account for video characteristics, beyond the traditional paradigm of fair-rate allocation among data flows. Ideally, such a scheme should handle both persistent and transient congestion as video streaming applications demand low-latency transmissions and low packet-loss ratios. This paper presents a novel scheme, layered Internet video adaptation (LIVA), in which network nodes feed back virtual congestion levels to video senders to assist both media-aware bandwidth sharing and transient-loss protection. The video senders respond to such feedback by adapting the rates of encoded scalable bitstreams based on their respective video rate-distortion (R-D) characteristics. The same feedback is employed to calculate the amount of forward error correction (FEC) protection for combating transient losses. Simulation studies show that LIVA can minimize the total distortion of all participating video streams and hence maximize their overall quality. At steady state, video streams experience no queueing delays or packet losses. In the face of transient congestion, the network-assisted adaptive FEC promptly protects video packets from losses. Our Linux-based demonstration showcases how LIVA can be implemented in a simple manner in real systems. We also present a solution for LIVA streams to coexist with TCP flows based on explicit congestion notification signaling. Finally, our theoretical analysis guarantees system stability for an arbitrary number of streams with round-trip delays below a prescribed limit.
Mythili Suryanarayana Prabhu, Nandita Dukkipati, Vijay G. Subramanian, Flavio Bonomi
IEEE Trans. Multim.4
2010 Layered Internet Video Engineering (LIVE): Network-Assisted Bandwidth Sharing and Transient Loss Protection for Scalable Video Streaming
abstract
This paper presents a novel scheme, Layered Internet Video Engineering (LIVE), in which network nodes feedback virtual congestion levels to video senders to assist both media-aware bandwidth sharing and transient loss protection. The video senders respond to such feedback by adapting the rates of encoded H.264/SVC streams based on their respective video rate-distortion (R-D) characteristics. The same feedback is employed to calculate the amount of forward error correction (FEC) protection for combating transient losses. Simulation studies show that LIVE can minimize the total distortion of all participating video streams and hence maximize their overall quality. At steady state, video streams experience no queuing delays or packet losses. In face of transient congestion, the network-assisted adaptive FEC effectively protect video packets from losses while keeping a minimum overhead. Our theoretical analysis further guarantees system stability for arbitrary number of streams with arbitrary round trip delays below a prescribed limit. Finally, we show that LIVE streams can coexist with TCP flows within the existing explicit congestion notification (ECN) framework.
Nandita Dukkipati, Vijay G. Subramanian, Flavio Bonomi
INFOCOM3
2008 Making Large Scale Deployment of RCP Practical for Real Networks
abstract
We recently proposed the rate control protocol (RCP) as a way to minimize download times (or flow-completion times). Simulations suggest that if RCP were widely deployed, downloads would frequently finish an order of magnitude faster than with TCP. This is because RCP involves explicit feedback from the routers along the path, allowing a sender to pick a fast starting rate, and adapt quickly to network conditions. RCP is particularly appealing because it can be shown to be stable under broad operating conditions, and its performance is independent of the flow-size distribution and the RTT. Although it requires changes to the routers, the changes are small: The routers keep no per-flow state or per-flow queues, and the per-packet processing is minimal. However, the bar is high for a new congestion control mechanism - introducing a new scheme requires enormous change, and the argument needs to be compelling. And so, to enable incremental deployment of RCP, we have built and tested an open and public implementation of RCP, and proposed solutions for deployments that require no fork-lift network upgrades. In this paper we describe our end-host and router implementation of RCP in Linux, and solutions to how RCP can coexist in a network carrying predominantly non-RCP traffic, and coordinate with routers that don't implement RCP. We hope that these solutions will take us closer to having an impact in real networks, not just for RCP but also for many other explicit congestion control protocols proposed in literature.
Chia-Hui Tai, Nandita Dukkipati
INFOCOM3
2006 RCP-AC: Congestion Control to Make Flows Complete Quickly in Any Environment
abstract
We believe that a congestion control algorithm should make flows finish quickly - as quickly as possible, while staying stable and fair among flows. Recently, we proposed RCP (Rate Control Protocol) which enables typical Internet-sized flows to complete one to two orders of magnitude faster than the existing (TCP Reno) and the proposed (XCP) congestion control algorithm. Like XCP, RCP uses explicit feedback from routers, but doesn't require per-packet calculations. A router maintains just one rate that it gives to all flows, making it simple and inherently fair. Flows finish quickly because RCP aggressively gives excess bandwidth to flows, making it work well in the common case. However - and this is a design tradeoff - RCP will experience short-term transient overflows when network conditions change quickly (e.g. a route change or flash crowds). In this paper we extend RCP and propose RCP-AC (Rate Control Protocol with Acceleration Control) that allows the aggressiveness of RCP to be tuned, enabling fast completion of flows over a broad set of operating conditions.
Nandita Dukkipati, Nick McKeown, Alexander G. Fraser 0001
INFOCOM1
2005 Processor Sharing Flows in the Internet
Nandita Dukkipati, Masayoshi Kobayashi, Rui Zhang-Shen, Nick McKeown
IWQoS1
2004 Optimal call admission control on a single link with a GPS scheduler
abstract
The problem of call admission control (CAC) is considered for leaky bucket constrained sessions with deterministic service guarantees (zero loss and finite delay bound) served by a generalized processor sharing scheduler at a single node in the presence of best effort traffic. Based on an optimization process, a CAC algorithm capable of determining the (unique) optimal solution is derived. The derived algorithm is also applicable, under a slight modification, in a system where the best effort traffic is absent and is capable of guaranteeing that if it does not find a solution to the CAC problem, then a solution does not exist. The numerical results indicate that the CAC algorithm can achieve a significant improvement on bandwidth utilization as compared to a (deterministic) effective bandwidth-based CAC scheme.
Antonis Panagakis, Nandita Dukkipati, Ioannis Stavrakakis, Joy Kuri
IEEE/ACM Trans. Netw.2
2001 Optimal Call Admission Control in Generalized Processor Sharing (GPS) Schedulers
abstract
Generalized processor sharing (GPS) is an idealized fluid discipline with a number of desirable properties. Its packetized version PGPS is considered to be a good choice is a packet scheduling discipline to guarantee quality-of-service in IP and ATM networks. The existing connection admission control (CAC) frameworks for GPS result in a conservative resource allocation. We propose an optimal CAC algorithm for a GPS scheduler, for leaky-bucket constrained connections with deterministic delay guarantees. Our numerical results show that the optimal CAC results in a higher network utilization than the existing CACs.
Nandita Dukkipati, Joy Kuri, H. S. Jamadagni
INFOCOM1