Prateesh Goyal

dblp:153/2265 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-7945-1821ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 15 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Queueless and Dropless Rate Control
abstract
Modern Internet traffic is increasingly dominated by interactive video applications (e.g. cloud gaming, conferencing, etc) that require high throughput, extremely low delay, and near-zero drop rates. Meeting these requirements is inherently challenging for such applications because video encoders often need to overshoot the target bitrates, causing transient traffic bursts. These bursts lead to inevitable delays and drops on traditional Internet bottlenecks with fixed capacities and finite physical queues. However, we observe that modern Internet bottlenecks are often artificial, where ISPs intentionally limit user's traffic to subscribed rates. Our work exploits the flexibility afforded in these artificial bottlenecks to achieve zero queuing and drops for interactive applications. Specifically, we make a case for using traffic policers (implemented as token bucket filters) at the ISPs to enforce desired rate limits on an average. Policers can absorb traffic bursts without any queuing, but incur packet drops. Our system, QLDL, avoids packet drops by augmenting the policer and the endpoints with a light-weight explicit rate control mechanism. Our prototype evaluation using WebRTC applications shows how QLDL thus achieves zero queuing and drops, thereby improving interactive applications' QoE on all fronts: 2× higher bitrate, 2 – 3× lower frame delays, and 20× lower playback freezes, when compared to state-of-the-art baselines.
Ammar Tahir, Prateesh Goyal, Yongzhou Chen, Radhika Mittal
SIGCOMM2
2025 Cuttlefish: A Fair, Predictable Execution Environment for Cloud-Hosted Financial Exchanges
Liangcheng Yu, Prateesh Goyal, Ilias Marinos, Vincent Liu 0001
AFT2
2024 m3: Accurate Flow-Level Performance Estimation using Machine Learning
abstract
Data center network operators often need accurate estimates of aggregate network performance. Unfortunately, existing methods for estimating aggregate network statistics are either inaccurate or too slow to be practical at the data center scale.
Chenning Li, Arash Nasr-Esfahany, Kevin Zhao, Kimia Noorbakhsh, Prateesh Goyal, Mohammad Alizadeh, Thomas E. Anderson
SIGCOMM5
2024 Efficient Policy-Rich Rate Enforcement with Phantom Queues
abstract
ISPs routinely rate-limit user traffic. In addition to correctly enforcing the desired rates, rate-limiting mechanisms must be able to support rich rate-sharing policies within each traffic aggregate (e.g. per-flow fairness, weighted fairness, and prioritization). This must be done at scale to support the vast magnitude of users efficiently. There are two primary rate-limiting mechanisms - traffic shaping (that buffers packets in queues to enforce the desired rates and policies) and traffic policing (that filters packets as per the desired rates without buffering them). Policers are lightweight and scalable but don't support rich policy enforcement and often provide poor rate enforcement (being notoriously hard to configure). Shapers, on the other hand, achieve desired rates and policies, but at the cost of high system resource (memory and CPU) utilization impacting scalability. This paper explores whether we can get the best of both worlds. We present our system BC-PQP, which augments a policer with (i) multiple phantom queues that simulate buffer occupancy using counters and enable rich policy enforcement, and (ii) a novel burst-control mechanism that enables auto-configuration of the queues for correct rate enforcement. Our system achieves the rate and policy enforcement properties close to that of a shaper with 7× higher efficiency.
Ammar Tahir, Prateesh Goyal, Ilias Marinos, Mike Evans, Radhika Mittal
SIGCOMM2
2023 Scalable Tail Latency Estimation for Data Center Networks
Kevin Zhao, Prateesh Goyal, Mohammad Alizadeh, Thomas E. Anderson
NSDI2
2023 DBO: Fairness for Cloud-Hosted Financial Exchanges
abstract
We consider the problem of hosting financial exchanges in the cloud. Exchanges necessitate strong fairness guarantees for competing participants, particularly for use cases such as "high frequency trading". Today, exchanges achieve such guarantees by providing equal latency across all market participants in their on-premise deployments. However, ensuring equal latency for fairness is notably challenging in current multi-tenant cloud deployments, mainly due to factors such as network congestion and non-equidistant network paths.
Eashan Gupta, Prateesh Goyal, Ilias Marinos, Chenxingyu Zhao, Radhika Mittal, Ranveer Chandra
SIGCOMM2
2022 Rethinking cloud-hosted financial exchanges for response time fairness
abstract
In this paper, we consider the problem of supporting modern financial exchange services on the cloud premises. Important exchange services rely on predictable, equal latency from the servers to the participants for fair competition. Existing cloud networks, however, are unable to offer such property, as they were not originally designed for this purpose. We attempt to tackle the problem of unfairness that stems from the lack of determinism in cloud networks. We argue that predictable or bounded latency is not necessary to achieve fairness. Inspired by the use of logical clocks in distributed systems, we propose a new approach that instead corrects for differences in latency to the participants for fairness. We evaluate our approach in simulation and show that it is feasible to achieve fairness under highly variable network latency. Our approach is deployable in contemporary cloud environments; it avoids limitations of state-of-the-art and outperforms it.
Prateesh Goyal, Ilias Marinos, Eashan Gupta, Chaitanya Bandi, Alan Ross, Ranveer Chandra
HotNets1
2022 Backpressure Flow Control
Prateesh Goyal, Preey Shah, Kevin Zhao, Georgios Nikolaidis, Mohammad Alizadeh, Thomas E. Anderson
NSDI1
2022 Elasticity detection: a building block for internet congestion control
abstract
This paper introduces a new metric, "elasticity," which characterizes the nature of cross-traffic competing with a flow. Elasticity captures whether the cross traffic reacts to changes in available bandwidth. We show that it is possible to robustly detect the elasticity of cross traffic at a sender without router support, and that elasticity detection can reduce delays in the Internet by enabling delay-controlling congestion control protocols to be deployed without hurting flow throughput. Our results show that the proposed method achieves more than 85% accuracy under a variety of network conditions, and that congestion control using elasticity detection achieves throughput comparable to Cubic but with delays that are 50--70 ms lower when cross traffic is inelastic.
Prateesh Goyal, Akshay Narayan 0001, Frank Cangialosi, Srinivas Narayana, Mohammad Alizadeh, Hari Balakrishnan
SIGCOMM1
2021 Site-to-site internet traffic control
abstract
Queues allow network operators to control traffic: where queues build, they can enforce scheduling and shaping policies. In the Internet today, however, there is a mismatch between where queues build and where control is most effectively enforced; queues build at bottleneck links that are often not under the control of the data sender. To resolve this mismatch, we propose a new kind of middlebox, called Bundler. Bundler uses a novel inner control loop between a sendbox (in the sender's site) and a receivebox (in the receiver's site) to determine the aggregate rate for the bundle, leaving the end-to-end connections and their control loops intact. Enforcing this sending rate ensures that bottleneck queues that would have built up from the bundle's packets now shift from the bottleneck to the sendbox. This enables the sendbox to exercise control over its traffic by scheduling packets according to any policy necessary to achieve the network operator's higher-level objectives. We have implemented Bundler in Linux and evaluated it with real-world and emulation experiments. We find that Bundler allows the sender-chosen policy to be effective: when configured to implement Stochastic Fairness Queueing (SFQ), it improves median flow completion time (FCT) by between 28% and 97% across various scenarios.
Frank Cangialosi, Akshay Narayan 0001, Prateesh Goyal, Radhika Mittal, Mohammad Alizadeh, Hari Balakrishnan
EuroSys3
2020 ABC: A Simple Explicit Congestion Controller for Wireless Networks
Prateesh Goyal, Anup Agarwal, Ravi Netravali, Mohammad Alizadeh, Hari Balakrishnan
NSDI1
2020 Annulus: A Dual Congestion Control Loop for Datacenter and WAN Traffic Aggregates
abstract
Cloud services are deployed in datacenters connected though high-bandwidth Wide Area Networks (WANs). We find that WAN traffic negatively impacts the performance of datacenter traffic, increasing tail latency by 2.5x, despite its small bandwidth demand. This behavior is caused by the long round-trip time (RTT) for WAN traffic, combined with limited buffering in datacenter switches. The long WAN RTT forces datacenter traffic to take the full burden of reacting to congestion. Furthermore, datacenter traffic changes on a faster time-scale than the WAN RTT, making it difficult for WAN congestion control to estimate available bandwidth accurately.
Ahmed Saeed 0001, Prateesh Goyal, Milad Sharif, Mostafa H. Ammar, Ellen Zegura, Keon Jang, Mohammad Alizadeh, Abdul Kabbani, Amin Vahdat
SIGCOMM3
2019 End-to-end transport for video QoE fairness
abstract
The growth of video traffic makes it increasingly likely that multiple clients share a bottleneck link, giving video content providers an opportunity to optimize the experience of multiple users jointly. But today's transport protocols are oblivious to video streaming applications and provide only connection-level fairness. We design and build Minerva, the first end-to-end transport protocol for multi-user video streaming. Minerva uses information about the player state and video characteristics to adjust its congestion control behavior to optimize for QoE fairness. Minerva clients receive no explicit information about other video clients, yet when multiple of them share a bottleneck link, their rates converge to a bandwidth allocation that maximizes QoE fairness. At the same time, Minerva videos occupy only their fair share of the bottleneck link bandwidth, competing fairly with existing TCP traffic. We implement Minerva on an industry standard video player and server and show that, compared to Cubic and BBR, 15-32%of the videos using Minerva experience an improvement in viewing experience equivalent to a jump in resolution from 720p to 1080p. Additionally, in a scenario with dynamic video arrivals and departures, Minerva reduces rebuffering time by an average of 47%.
Vikram Nathan, Vibhaalakshmi Sivaraman, Ravichandra Addanki, Mehrdad Khani Shirkoohi, Prateesh Goyal, Mohammad Alizadeh
SIGCOMM5
2019 Faster Algorithms for Dynamic Algebraic Queries in Basic RSMs with Constant Treewidth
abstract
Interprocedural analysis is at the heart of numerous applications in programming languages, such as alias analysis, constant propagation, and so on. Recursive state machines (RSMs) are standard models for interprocedural analysis. We consider a general framework with RSMs where the transitions are labeled from a semiring and path properties are algebraic with semiring operations. RSMs with algebraic path properties can model interprocedural dataflow analysis problems, the shortest path problem, the most probable path problem, and so on. The traditional algorithms for interprocedural analysis focus on path properties where the starting point is fixed as the entry point of a specific method. In this work, we consider possible multiple queries as required in many applications such as in alias analysis. The study of multiple queries allows us to bring in an important algorithmic distinction between the resource usage of the one-time preprocessing vs for each individual query. The second aspect we consider is that the control flow graphs for most programs have constant treewidth. Our main contributions are simple and implementable algorithms that support multiple queries for algebraic path properties for RSMs that have constant treewidth. Our theoretical results show that our algorithms have small additional one-time preprocessing but can answer subsequent queries significantly faster as compared to the current algorithmic solutions for interprocedural dataflow analysis. We have also implemented our algorithms and evaluated their performance for performing on-demand interprocedural dataflow analysis on various domains, such as for live variable analysis and reaching definitions, on a standard benchmark set. Our experimental results align with our theoretical statements and show that after a lightweight preprocessing, on-demand queries are answered much faster than the standard existing algorithmic approaches.
Krishnendu Chatterjee, Amir Kafshdar Goharshady, Prateesh Goyal, Rasmus Ibsen-Jensen, Andreas Pavlogiannis
ACM Trans. Program. Lang. Syst.3
2018 Restructuring endpoint congestion control
abstract
This paper describes the implementation and evaluation of a system to implement complex congestion control functions by placing them in a separate agent outside the datapath. Each datapath---such as the Linux kernel TCP, UDP-based QUIC, or kernel-bypass transports like mTCP-on-DPDK---summarizes information about packet round-trip times, receptions, losses, and ECN via a well-defined interface to algorithms running in the off-datapath Congestion Control Plane (CCP). The algorithms use this information to control the datapath's congestion window or pacing rate. Algorithms written in CCP can run on multiple datapaths. CCP improves both the pace of development and ease of maintenance of congestion control algorithms by providing better, modular abstractions, and supports aggregation capabilities of the Congestion Manager, all with one-time changes to datapaths. CCP also enables new capabilities, such as Copa in Linux TCP, several algorithms running on QUIC and mTCP/DPDK, and the use of signal processing algorithms to detect whether cross-traffic is ACK-clocked. Experiments with our user-level Linux CCP implementation show that CCP algorithms behave similarly to kernel algorithms, and incur modest CPU overhead of a few percent.
Akshay Narayan 0001, Frank Cangialosi, Deepti Raghavan, Prateesh Goyal, Srinivas Narayana, Radhika Mittal, Mohammad Alizadeh, Hari Balakrishnan
SIGCOMM4
2017 Rethinking Congestion Control for Cellular Networks
abstract
We propose Accel-Brake Control (ABC), a protocol that integrates a simple and deployable signaling scheme at cellular base stations with an endpoint mechanism to respond to these signals. The key idea is for the base station to enable each sender to achieve a computed target rate by marking each packet with an "accelerate" or "brake" notification, which causes the sender to either slightly increase or slightly reduce its congestion window. ABC is designed to rapidly acquire any capacity that opens up, a common occurrence in cellular networks, while responding promptly to congestion. It is also incrementally deployable using existing ECN infrastructure and can co-exist with legacy ECN routers. Preliminary results obtained over cellular network traces show that ABC outperforms prior approaches significantly.
Prateesh Goyal, Mohammad Alizadeh, Hari Balakrishnan
HotNets1
2017 The Case for Moving Congestion Control Out of the Datapath
abstract
With Moore's law ending, the gap between general-purpose processor speeds and network link rates is widening. This trend has led to new packet-processing "datapaths" in endpoints, including kernel bypass software and emerging SmartNIC hardware. In addition, several applications are rolling out their own protocols atop UDP (e.g., QUIC, WebRTC, Mosh, etc.), forming new datapaths different from the traditional kernel TCP stack. All these datapaths require congestion control, but they must implement it separately because it is not possible to reuse the kernel's TCP implementations. This paper proposes moving congestion control from the datapath into a separate agent. This agent, which we call the congestion control plane (CCP), must provide both an expressive congestion control API as well as a specification for datapath designers to implement and deploy CCP. We propose an API for congestion control, datapath primitives, and a user-space agent design that uses a batching method to communicate with the datapath. Our approach promises to preserve the behavior and performance of in-datapath implementations while making it significantly easier to implement and deploy new congestion control algorithms.
Akshay Narayan 0001, Frank Cangialosi, Prateesh Goyal, Srinivas Narayana, Mohammad Alizadeh, Hari Balakrishnan
HotNets3
2017 Language-Directed Hardware Design for Network Performance Monitoring
abstract
Network performance monitoring today is restricted by existing switch support for measurement, forcing operators to rely heavily on endpoints with poor visibility into the network core. Switch vendors have added progressively more monitoring features to switches, but the current trajectory of adding specific features is unsustainable given the ever-changing demands of network operators. Instead, we ask what switch hardware primitives are required to support an expressive language of network performance questions. We believe that the resulting switch hardware design could address a wide variety of current and future performance monitoring needs.
Srinivas Narayana, Anirudh Sivaraman, Vikram Nathan, Prateesh Goyal, Venkat Arun, Mohammad Alizadeh, Vimalkumar Jeyakumar, Changhoon Kim
SIGCOMM4
2015 Faster Algorithms for Algebraic Path Properties in Recursive State Machines with Constant Treewidth
abstract
Interprocedural analysis is at the heart of numerous applications in programming languages, such as alias analysis, constant propagation, etc. Recursive state machines (RSMs) are standard models for interprocedural analysis. We consider a general framework with RSMs where the transitions are labeled from a semiring, and path properties are algebraic with semiring operations. RSMs with algebraic path properties can model interprocedural dataflow analysis problems, the shortest path problem, the most probable path problem, etc. The traditional algorithms for interprocedural analysis focus on path properties where the starting point is fixed as the entry point of a specific method. In this work, we consider possible multiple queries as required in many applications such as in alias analysis. The study of multiple queries allows us to bring in a very important algorithmic distinction between the resource usage of the one-time preprocessing vs for each individual query. The second aspect that we consider is that the control flow graphs for most programs have constant treewidth. Our main contributions are simple and implementable algorithms that support multiple queries for algebraic path properties for RSMs that have constant treewidth. Our theoretical results show that our algorithms have small additional one-time preprocessing, but can answer subsequent queries significantly faster as compared to the current best-known solutions for several important problems, such as interprocedural reachability and shortest path. We provide a prototype implementation for interprocedural reachability and intraprocedural shortest path that gives a significant speed-up on several benchmarks.
Krishnendu Chatterjee, Rasmus Ibsen-Jensen, Andreas Pavlogiannis, Prateesh Goyal
POPL4