EDBT 2026 Demo / reviewers in the wild / expert
Manya Ghobadi
dblp:37/1704 · also Monia Ghobadi
· DBLP profile ↗
46ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-4095-1519ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 39 · 7 first-author · 16 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient Replication
Ankit Bhardwaj 0002, Weiyang Wang, Jeremy Carin, Adam Belay, Manya Ghobadi |
NSDI | 5 |
| 2026 | MonkeyTree: Near-Minimal Congestion for Multi-tenant Training via MigrationabstractWe present MonkeyTree, the first system to mitigate network congestion in multi-tenant GPU clusters through job-migration based defragmentation rather than network-layer techniques. As cloud operators co-locate ML training jobs on shared, oversubscribed networks, congestion degrades training throughput for over a third of jobs. Prior approaches either rely on routing and flow scheduling—which we show have fundamental limits when traffic exceeds capacity, or require costly full-bisection bandwidth topologies with packet spraying. Anton A. Zabreyko, Weiyang Wang, Manya Ghobadi |
SIGCOMM | 3 |
| 2025 | FORESIGHT: Joint Time and Space Scheduling for Efficient Distributed ML Training
Farid Zandi Shafagh, Manya Ghobadi, Yashar Ganjali |
Networking | 2 |
| 2024 | Rail-only: A Low-Cost High-Performance Network for Training LLMs with Trillion ParametersabstractThis paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a novel datacenter network design tailored to LLM's unique communication pattern. We show that LLM training generates sparse communication patterns in the network and, therefore, does not require any-to-any full-bisection network to complete efficiently. As a result, our design eliminates the spine layer in traditional GPU clusters. We name this design a Rail-only network and demonstrate that it achieves the same training performance while reducing the network cost by 38% to 77% and network power consumption by 37% to 75% compared to a conventional GPU datacenter. Our architecture also supports Mixture-of-Expert (MoE) models with all-to-all communication through forwarding, with only 4.1% to 5.6% completion time overhead for all-to-all traffic. We study the failure robustness of Rail-only networks and provide insights into the performance impact of different network and training parameters. Weiyang Wang, Manya Ghobadi, Kayvon Shakeri, Ying Zhang 0022, Naader Hasani |
HOTI | 2 |
| 2024 | The Case for Decentralized Fallback NetworksabstractThis paper argues that network and application delivery infrastructures have become highly centralized and are more vulnerable to attacks and disasters than is desirable. It proposes a research agenda for decentralized fallback networks and focuses on a key component---a city-scale decentralized network using existing Wi-Fi access points, which are deployed across almost all buildings in cities. It proposes a routing system that uses information about buildings from geospatial maps instead of traditional routing mechanisms to scale well to millions of Wi-Fi nodes. James C. Lynch, Chenning Li, Manya Ghobadi, Hari Balakrishnan |
HotNets | 4 |
| 2024 | MLTCP: A Distributed Technique to Approximate Centralized Flow Scheduling For Machine LearningabstractThis paper argues that congestion control protocols in machine learning datacenters sit at a sweet spot between centralized and distributed flow scheduling solutions. We present MLTCP, a technique to augment today's congestion control algorithms to approximate an interleaved centralized flow schedule. At the heart of MLTCP lies a straight-forward principle based on a key conceptual insight: by scaling the congestion window size (or sending rate) based on the number of bytes sent at each iteration, MLTCP flows eventually converge into a schedule that reduces network contention. We demonstrate that MLTCP uses a gradient descent trend with a step taken at every training (or fine-tuning) iteration towards reducing network congestion among competing jobs. Sudarsanan Rajasekaran, Sanjoli Narang, Anton A. Zabreyko, Manya Ghobadi |
HotNets | 4 |
| 2024 | CASSINI: Network-Aware Job Scheduling in Machine Learning Clusters
Sudarsanan Rajasekaran, Manya Ghobadi, Aditya Akella |
NSDI | 2 |
| 2024 | NetBlocks: Staging Layouts for High-Performance Custom Host Network StacksabstractModern network applications and environments, ranging from data centers and IoT devices to AR/VR headsets and underwater robotics, present diverse requirements that cannot be satisfied by the all-or-nothing approach of TCP and UDP protocols. Network researchers and engineers need to create highly tailored protocols targeting individual problem domains. Existing library-based approaches either fall short on the flexibility in features or offer them at a significant performance overhead. To address this challenge, we present NetBlocks, a domain-specific language, and compiler for designing ad-hoc protocols and generating their highly optimized host network stack implementations. NetBlocks DSL input allows users to configure protocols by selecting and customizing features. Unlike other DSL compilers, NetBlocks also allows network researchers to extend the system and add more features easily without any prior compiler knowledge. Our design and implementation employ a high-performance Aspect-Oriented Programming framework written with the staging framework BuildIt. We also introduce a novel Layout Customization Layer that allows "staging packet layouts" alongside the implementation, which is critical for getting the best performance out of the protocol when possible, while allowing the practitioners to maintain compatibility with existing protocol layers where needed. Our evaluations on three applications ranging across deployments in data centers and underwater acoustic networks demonstrate a trade-off between performance (both latency and throughput) and selected features allowing the user to only pay-for-what-they-use. Ajay Brahmakshatriya, Christopher Rinard, Manya Ghobadi, Saman P. Amarasinghe |
Proc. ACM Program. Lang. | 3 |
| 2023 | On-Fiber Photonic ComputingabstractIn the 1800s, Charles Babbage envisioned computers as analog devices. However, it was not until 150 years later that a Mechanical Analog Computer was constructed for the US Navy to solve differential equations. With the end of Moore's Law, photonic computing is revitalizing the promise of analog computing by leveraging photons' speed, bandwidth, and energy efficiency for faster, more efficient, and scalable analog computing systems. This paper argues that the networking community should augment pluggable transponders with photonic computing capabilities to enable a backward-compatible solution for in-network computing. We propose on-fiber photonic computing to perform computing operations inside network transponders while the data is in the optical domain. We discuss the components required to enable the seamless integration of computation into the very fabric of optical communication links. We then discuss several use cases of on-fiber photonic computing, including machine learning inference, video encoding, load balancing, and intrusion detection. Mingran Yang, Zhizhen Zhong, Manya Ghobadi |
HotNets | 3 |
| 2023 | InfoShape: Task-Based Neural Data Shaping via Mutual InformationabstractThe use of mutual information as a tool in private data sharing has remained an open challenge due to the difficulty of its estimation in practice. In this paper, we propose InfoShape, a task-based encoder that aims to remove unnecessary sensitive information from training data while maintaining enough relevant information for a particular ML training task. We achieve this goal by utilizing mutual information estimators that are based on neural networks, in order to measure two performance metrics, privacy and utility. Using these together in a Lagrangian optimization, we train a separate neural network as a lossy encoder. We empirically show that InfoShape is capable of shaping the encoded samples to be informative for a specific downstream task while eliminating unnecessary sensitive information. Moreover, we demonstrate that the classification accuracy of downstream models has a meaningful connection with our utility and privacy measures. Homa Esfahanizadeh, William Wu, Manya Ghobadi, Regina Barzilay, Muriel Médard |
ICASSP | 3 |
| 2023 | TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Dheevatsa Mudigere, Ying Zhang 0022, Anthony Kewitsch |
NSDI | 4 |
| 2023 | Demo: First Demonstration of Real-Time Photonic-Electronic DNN Acceleration on SmartNICsabstractWe demonstrate Lightning, a reconfigurable photonic-electronic deep learning smartNIC that serves real-time inference requests at 4.055 GHz compute frequency. To do so, Lightning uses a novel datapath to feed traffic from the NIC into its photonic computing cores without incurring digital data movement bottlenecks. Lightning achieves this by employing a reconfigurable count-action abstraction, which decouples the compute control plane from the data plane. The count-action abstraction counts the number of operations for each computation task in the Directed Acyclic Graph (DAG). It then triggers the execution of the next task(s) as soon as the previous task is finished without interrupting the dataflow. Our prototype shows that Lightning achieves 99.25% photonic MAC accuracy. When serving real-time inference requests, Lightning accelerates the end-to-end inference latency of the LeNet DNN by 9.4× and 6.6× compared to Nvidia P4 and A100 GPUs, respectively. Zhizhen Zhong, Mingran Yang, Jay Lang, Dirk R. Englund, Manya Ghobadi |
SIGCOMM | 5 |
| 2023 | Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient InferenceabstractThe massive growth of machine learning-based applications and the end of Moore's law have created a pressing need to redesign computing platforms. We propose Lightning, the first reconfigurable photonic-electronic smartNIC to serve real-time deep neural network inference requests. Lightning uses a fast datapath to feed traffic from the NIC into the photonic domain without creating digital packet processing and data movement bottlenecks. To do so, Lightning leverages a novel reconfigurable count-action abstraction that keeps track of the required computation operations of each inference packet. Our count-action abstraction decouples the compute control plane from the data plane by counting the number of operations in each task and triggers the execution of the next task(s) without interrupting the dataflow. We evaluate Lightning's performance using four platforms: a prototype, chip synthesis, emulations, and simulations. Our prototype demonstrates the feasibility of performing 8-bit photonic multiply-accumulate operations with 99.25% accuracy. To the best of our knowledge, our prototype is the highest-frequency photonic computing system, capable of serving real-time inference queries at 4.055 GHz end-to-end. Our simulations with large DNN models show that compared to Nvidia A100 GPU, A100X DPU, and Brainwave smartNIC, Lightning accelerates the average inference serve time by 337×, 329×, and 42×, while consuming 352×, 419×, and 54× less energy, respectively. Zhizhen Zhong, Mingran Yang, Jay Lang, Christian Williams, Liam Kronman, Alex Sludds, Homa Esfahanizadeh, Dirk R. Englund, Manya Ghobadi |
SIGCOMM | 9 |
| 2022 | Congestion control in machine learning clustersabstractThis paper argues that fair-sharing, the holy grail of congestion control algorithms for decades, is not necessarily a desirable property in Machine Learning (ML) training clusters. We demonstrate that for a specific combination of jobs, introducing unfairness improves the training time for all competing jobs. We call this specific combination of jobs compatible and define the compatibility criterion using a novel geometric abstraction. Our abstraction rolls time around a circle and rotates the communication phases of jobs to identify fully compatible jobs. Using this abstraction, we demonstrate up to 1.3× improvement in the average training iteration time of popular ML models. We advocate that resource management algorithms should take job compatibility on network links into account. We then propose three directions to ameliorate the impact of network congestion in ML training clusters: (i) an adaptively unfair congestion control scheme, (ii) priority queues on switches, and (iii) precise flow scheduling. Sudarsanan Rajasekaran, Manya Ghobadi, Gautam Kumar 0001, Aditya Akella |
HotNets | 2 |
| 2022 | ABM: active buffer management in datacentersabstractToday's network devices share buffer across queues to avoid drops during transient congestion and absorb bursts. As the buffer-per-bandwidth-unit in datacenter decreases, the need for optimal buffer utilization becomes more pressing. Typical devices use a hierarchical packet admission control scheme: First, a Buffer Management (BM) scheme decides the maximum length per queue at the device level and then an Active Queue Management (AQM) scheme decides which packets will be admitted at the queue level. Unfortunately, the lack of cooperation between the two control schemes leads to (i) harmful interference across queues, due to the lack of isolation; (ii) increased queueing delay, due to the obliviousness to the per-queue drain time; and (iii) thus unpredictable burst tolerance. To overcome these limitations, we propose ABM, Active Buffer Management which incorporates insights from both BM and AQM. Concretely, ABM accounts for both total buffer occupancy (typically used by BM) and queue drain time (typically used by AQM). We analytically prove that ABM provides isolation, bounded buffer drain time and achieves predictable burst tolerance without sacrificing throughput. We empirically find that ABM improves the 99th percentile FCT for short flows by up to 94% compared to the state-of-the-art buffer management. We further show that ABM improves the performance of advanced datacenter transport protocols in terms of FCT by up to 76% compared to DCTCP, TIMELY and PowerTCP under bursty workloads even at moderate load conditions. Vamsi Addanki, Maria Apostolaki, Manya Ghobadi, Stefan Schmid 0001, Laurent Vanbever |
SIGCOMM | 3 |
| 2022 | Using trio: juniper networks' programmable chipset - for emerging in-network applicationsabstractThis paper describes Trio, a programmable chipset used in Juniper Networks' MX-series routers and switches. Trio's architecture is based on a multi-threaded programmable packet processing engine and a hierarchy of high-capacity memory systems, making it fundamentally different from pipeline-based architectures. Trio gracefully handles non-homogeneous packet processing rates for a wide range of networking use cases and protocols, making it an ideal platform for emerging in-network applications. We begin by describing the Trio chipset's fundamental building blocks, including its multi-threaded Packet Forwarding and Packet Processing Engines. We then discuss Trio's programming language, called Microcode. To showcase Trio's flexible Microcode-based programming environment, we describe two use cases. First, we demonstrate Trio's ability to perform in-network aggregation for distributed machine learning. Second, we propose and design an in-network straggler mitigation technique using Trio's timer threads. We prototype both use cases on a testbed using three real DNN models (ResNet50, DenseNet161, and VGG11) to demonstrate Trio's ability to mitigate stragglers while performing in-network aggregation. Our evaluations show that when stragglers occur in the cluster, Trio outperforms today's pipeline-based solutions by up to 1.8x. Mingran Yang, Alex Baban, Valery Kugel, Jeff Libby, Scott Mackie, Swamy Sadashivaiah Renu Kananda, Chang-Hong Wu, Manya Ghobadi |
SIGCOMM | 8 |
| 2021 | A Social Network Under Social Distancing: Risk-Driven Backbone Management During COVID-19 and Beyond
Yiting Xia, Ying Zhang 0022, Zhizhen Zhong, Guanqing Yan, Chiunlin Lim, Satyajeet Ahuja, Soshant Bali, Alexander Nikolaidis, Kimia Ghobadi, Manya Ghobadi |
NSDI | 10 |
| 2021 | SiP-ML: high-bandwidth optical network interconnects for machine learning trainingabstractThis paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3--9.1x. Mehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, Eiman Ebrahimi |
SIGCOMM | 2 |
| 2021 | ARROW: restoration-aware traffic engineeringabstractFiber cut events reduce the capacity of wide-area networks (WANs) by several Tbps. In this paper, we revive the lost capacity by reconfiguring the wavelengths from cut fibers into healthy fibers. We highlight two challenges that made prior solutions impractical and propose a system called Arrow to address them. First, our measurements show that contrary to common belief, in most cases, the lost capacity is only partially restorable. This poses a cross-layer challenge from the Traffic Engineering (TE) perspective that has not been considered before: “Which IP links should be restored and by how much to best match the TE objective?” To address this challenge, Arrow's restoration-aware TE system takes a set of partial restoration candidates (that we call LotteryTickets) as input and proactively finds the best restoration plan. Second, prior work has not considered the reconfiguration latency of amplifiers. However, in practical settings, amplifiers add tens of minutes of reconfiguration delay. To enable fast and practical restoration, Arrow leverages optical noise loading and bypasses amplifier reconfiguration altogether. We evaluate Arrow using large-scale simulations and a testbed. Our testbed demonstrates Arrow's end-to-end restoration latency is eight seconds. Our large-scale simulations compare Arrow to the state-of-the-art TE schemes and show it can support 2.0x--2.4x more demand without compromising 99.99% availability. Zhizhen Zhong, Manya Ghobadi, Alaa Khaddaj, Jonathan Leach, Yiting Xia, Ying Zhang 0022 |
SIGCOMM | 2 |
| 2020 | Challenging the Stateless Quo of Programmable SwitchesabstractProgrammable switches based on the Protocol Independent Switch Architecture (PISA) have greatly enhanced the flexibility of today's networks by allowing new packet protocols to be deployed without any hardware changes. They have also been instrumental in enabling a new computing paradigm in which parts of an application's logic run within the network core (in-network computing). Nadeen Gebara, Alberto Lerner, Mingran Yang, Minlan Yu, Paolo Costa, Manya Ghobadi |
HotNets | 6 |
| 2020 | Enabling Programmable Transport Protocols in High-Speed NICs
Mina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford, David Walker 0001, David Wentzlaff |
NSDI | 3 |
| 2019 | Bandwidth steering in HPC using silicon nanophotonicsabstractAs bytes-per-FLOP ratios continue to decline, communication is becoming a bottleneck for performance scaling. This paper describes bandwidth steering in HPC using emerging reconfigurable silicon photonic switches. We demonstrate that placing photonics in the lower layers of a hierarchical topology efficiently changes the connectivity and consequently allows operators to recover from system fragmentation that is otherwise hard to mitigate using common task placement strategies. Bandwidth steering enables efficient utilization of the higher layers of the topology and reduces cost with no performance penalties. In our simulations with a few thousand network endpoints, bandwidth steering reduces static power consumption per unit throughput by 36% and dynamic power consumption by 14% compared to a reference fat tree topology. Such improvements magnify as we taper the bandwidth of the upper network layer. In our hardware testbed, bandwidth steering improves total application execution time by 69%, unaffected by bandwidth tapering. George Michelogiannakis, Yiwen Shen 0002, Min Yee Teh, Xiang Meng 0003, Benjamin Aivazi, Taylor L. Groves, John Shalf, Madeleine Glick, Manya Ghobadi, Larry Dennison, Keren Bergman |
SC | 9 |
| 2019 | TEAVAR: striking the right utilization-availability balance in WAN traffic engineeringabstractTo keep up with the continuous growth in demand, cloud providers spend millions of dollars augmenting the capacity of their wide-area backbones and devote significant effort to efficiently utilizing WAN capacity. A key challenge is striking a good balance between network utilization and availability, as these are inherently at odds; a highly utilized network might not be able to withstand unexpected traffic shifts resulting from link/node failures. We advocate a novel approach to this challenge that draws inspiration from financial risk theory: leverage empirical data to generate a probabilistic model of network failures and maximize bandwidth allocation to network users subject to an operator-specified availability target. Our approach enables network operators to strike the utilization-availability balance that best suits their goals and operational reality. We present TEAVAR (Traffic Engineering Applying Value at Risk), a system that realizes this risk management approach to traffic engineering (TE). We compare TEAVAR to state-of-the-art TE solutions through extensive simulations across many network topologies, failure scenarios, and traffic patterns, including benchmarks extrapolated from Microsoft's WAN. Our results show that with TEAVAR, operators can support up to twice as much throughput as state-of-the-art TE schemes, at the same level of availability. Jeremy Bogle, Nikhil Bhatia, Manya Ghobadi, Ishai Menache, Nikolaj S. Bjørner, Asaf Valadarsky, Michael Schapira |
SIGCOMM | 3 |
| 2018 | Characterizing the algorithmic complexity of reconfigurable data center architecturesabstractEmerging data center architectures are becoming reconfigurable. While prior work has shown the practical benefits of reconfigurable topologies, the underlying algorithmic complexity is not yet well understood. In particular, most reconfigurable topologies are hybrid, where parts of the network are reconfigurable (consisting of optical or wireless devices) while other parts are static (consisting of electrical switches). Current proposals enforce a routing policy that routes flows on either part "exclusively" by labeling flows as mice or elephant. We show that such artificial segregation in routing policy results in non-optimal paths and argue for algorithms that route packets across the network seamlessly. In doing so, we present the first algorithmic study of reconfigurable network architectures and provide optimality and hardness proofs in terms of topology and routing policy. Our results show that classical matching algorithms, as used in prior work, are optimal only when the topology consists of one reconfigurable switch, and the routing policy is enforced to be segregated. In other words, if there is an option of routing flows seamlessly along reconfigurable and non-reconfigurable parts of the network, matching algorithms are not optimal. In fact, when the hybrid network is seen from a joint perspective, optimal routing is an NP-hard problem. We further show that optimally routing even two flows in a network with multiple reconfigurable switches is an NP-hard problem as well. Klaus-Tycho Förster, Manya Ghobadi, Stefan Schmid 0001 |
ANCS | 2 |
| 2018 | Beyond SmartNICs: Towards a Fully Programmable Cloud: Invited PaperabstractFPGA-based SmartNICs and programmable switches have been recently introduced to leverage hardware acceleration and custom pipelines inside the cloud infrastructure. These devices are capable of handling the per-packet processing needs at line rate, including load balancing, encapsulation, congestion management, and security. We argue, however, that the benefits provided by these new devices could extend beyond software-defined networking use cases and they prompt a shift towards a fully programmable cloud, which would enable hardware-software co-design across all layers, ranging from application to hardware and networks. In this paper, we focus on the potential of FPGA-based SmartNICs and programmable switches to realize this vision and illustrate some of the research challenges that need to be addressed to fully unleash its benefits. Adrian M. Caulfield, Paolo Costa, Manya Ghobadi |
HPSR | 3 |
| 2018 | RADWAN: rate adaptive wide area networkabstractFiber optic cables connecting data centers are an expensive but important resource for large organizations. Their importance has driven a conservative deployment approach, with redundancy and reliability baked in at multiple layers. In this work, we take a more aggressive approach and argue for adapting the capacity of fiber optic links based on their signal-to-noise ratio (SNR). We investigate this idea by analyzing the SNR of over 8,000 links in an optical backbone for a period of three years. We show that the capacity of 64% of 100 Gbps IP links can be augmented by at least 75 Gbps, leading to an overall capacity gain of over 134 Tbps. Moreover, adapting link capacity to a lower rate can prevent up to 25% of link failures. Our analysis shows that using the same links, we get higher capacity, better availability, and 32% lower cost per gigabit per second. To accomplish this, we propose RADWAN, a traffic engineering system that allows optical links to adapt their rate based on the observed SNR to achieve higher throughput and availability while minimizing the churn during capacity reconfigurations. We evaluate RADWAN using a testbed consisting of 1,540 km fiber with 16 amplifiers and attenuators. We then simulate the throughput gains of RADWAN at scale and compare them to the gains of state-of-the-art traffic engineering systems. Our data-driven simulations show that RADWAN improves the overall network throughput by 40% while also improving the average link availability. Rachee Singh, Manya Ghobadi, Klaus-Tycho Förster, Mark Filer, Phillipa Gill |
SIGCOMM | 2 |
| 2017 | HotCocoa: Hardware Congestion Control AbstractionsabstractCongestion control in multi-tenant data centers is an active area of research because of its significant impact on customer experience, and, consequently, on revenue. Therefore, new algorithms and protocols are expected to emerge as the Cloud evolves. Deploying new congestion control algorithms in the end host's hypervisor allows frequent updates, but processing packets at high rates in the hypervisor and implementing the elements of a congestion control algorithm, such as traffic shapers and timestamps, in software have well-studied inaccuracies and CPU inefficiencies. In this paper, we argue for implementing the entire congestion control algorithm in programmable NICs. To do so, we identify the absence of hardware-aware programming abstractions as the most immediate challenge and solve it using a simple high-level domain specific language called HotCocoa. HotCocoa lies at a sweet spot between the ability to express a broad set of congestion control algorithms and efficient hardware implementation. It offers a set of hardware-aware COngestion COntrol Abstractions that enable operators to specify their algorithm without having to worry about low-level hardware primitives. To evaluate HotCocoa, we implement four congestion control algorithms (Reno, DCTCP, PCC, and TIMELY) and use simulations to show that HotCocoa's implementation of Reno perfectly tracks the behavior of a native implementation in C++. Mina Tahmasbi Arashloo, Manya Ghobadi, Jennifer Rexford, David Walker 0001 |
HotNets | 2 |
| 2017 | Run, Walk, Crawl: Towards Dynamic Link CapacitiesabstractFiber optic cables are the workhorses of today's Internet services. Operators spend millions of dollars to purchase, lease and maintain their optical backbone, making the efficiency of fiber essential to their business. In this work, we make a case for adapting the capacity of optical links based on their signal-to-noise ratio (SNR). We show two immediate benefits of this by analyzing the SNR of over 2000 links in an optical backbone over a period of 2.5 years. First, the capacity of 80% of IP links can be augmented by 75% or more, leading to an overall capacity gain of 145 Tbps in a large optical backbone in North America. Second, at least 25% of link failures are caused by SNR degradation, not complete loss-of-light, highlighting the opportunity to replace link failures by link flaps wherein the capacity is adjusted according to the new SNR. Given these benefits, we identify the disconnect between current optical and networking infrastructure which hinders the deployment of dynamic capacity links in wide area networks (WANs). To bridge this gap, we propose a graph abstraction that enables existing traffic engineering algorithms to benefit from dynamic link capacities. We evaluate the feasibility of dynamic link capacities using a small testbed and simulate the throughput gains from deploying our approach. Rachee Singh, Manya Ghobadi, Klaus-Tycho Förster, Mark Filer, Phillipa Gill |
HotNets | 2 |
| 2017 | RAIL: A Case for Redundant Arrays of Inexpensive Links in Data Center Networks
Danyang Zhuo, Manya Ghobadi, Ratul Mahajan, Amar Phanishayee, Xuan Kelvin Zou, Hang Guan, Arvind Krishnamurthy, Thomas E. Anderson |
NSDI | 2 |
| 2017 | Understanding and Mitigating Packet Corruption in Data Center NetworksabstractWe take a comprehensive look at packet corruption in data center networks, which leads to packet losses and application performance degradation. By studying 350K links across 15 production data centers, we find that the extent of corruption losses is significant and that its characteristics differ markedly from congestion losses. Corruption impacts fewer links than congestion, but imposes a heavier loss rate; and unlike congestion, corruption rate on a link is stable over time and is not correlated with its utilization. Danyang Zhuo, Manya Ghobadi, Ratul Mahajan, Klaus-Tycho Förster, Arvind Krishnamurthy, Thomas E. Anderson |
SIGCOMM | 2 |
| 2016 | ECN or Delay: Lessons Learnt from Analysis of DCQCN and TIMELYabstractData center networks, and especially drop-free RoCEv2 networks require efficient congestion control protocols. DCQCN (ECN-based) and TIMELY (delay-based) are two recent proposals for this purpose. In this paper, we analyze DCQCN and TIMELY using fluid models and simulations, for stability, convergence, fairness and flow completion time. We uncover several surprising behaviors of these protocols. For example, we show that DCQCN exhibits non-monotonic stability behavior, and that TIMELY can converge to stable regime with arbitrary unfairness. We propose simple fixes and tuning for ensuring that both protocols converge to and are stable at the fair share point. Finally, using lessons learnt from the analysis, we address the broader question: are there fundamental reasons to prefer either ECN or delay for end-to-end congestion control in data center networks? We argue that ECN is a better congestion signal, due to the way modern switches mark packets, and due to a fundamental limitation of end-to-end delay-based protocols, that we derive. Yibo Zhu 0001, Manya Ghobadi, Vishal Misra, Jitendra Padhye |
CoNEXT | 2 |
| 2016 | Optical Layer Failures in a Large Backbone
Manya Ghobadi, Ratul Mahajan |
Internet Measurement Conference | 1 |
| 2016 | ProjecToR: Agile Reconfigurable Data Center InterconnectabstractWe explore a novel, free-space optics based approach for building data center interconnects. It uses a digital micromirror device (DMD) and mirror assembly combination as a transmitter and a photodetector on top of the rack as a receiver (Figure 1). Our approach enables all pairs of racks to establish direct links, and we can reconfigure such links (i.e., connect different rack pairs) within 12 us. To carry traffic from a source to a destination rack, transmitters and receivers in our interconnect can be dynamically linked in millions of ways. We develop topology construction and routing methods to exploit this flexibility, including a flow scheduling algorithm that is a constant factor approximation to the offline optimal solution. Experiments with a small prototype point to the feasibility of our approach. Simulations using realistic data center workloads show that, compared to the conventional folded-Clos interconnect, our approach can improve mean flow completion time by 30-95% and reduce cost by 25-40%. Manya Ghobadi, Ratul Mahajan, Amar Phanishayee, Nikhil R. Devanur, Janardhan Kulkarni, Gireeja Ranade, Pierre-Alexandre Blanche, Houman Rastegarfar, Madeleine Glick, Daniel C. Kilper |
SIGCOMM | 1 |
| 2015 | Efficient traffic splitting on commodity switchesabstractTraffic often needs to be split over multiple equivalent backend servers, links, paths, or middleboxes. For example, in a load-balancing system, switches distribute requests of online services to backend servers. Hash-based approaches like Equal-Cost Multi-Path (ECMP) have low accuracy due to hash collision and incur significant churn during update. In a Software-Defined Network (SDN) the accuracy of traffic splits can be improved by crafting a set of wildcard rules for switches that better match the actual traffic distribution. The drawback of existing SDN-based traffic-splitting solutions is poor scalability as they generate too many rules for small rule-tables on switches. In this paper, we propose Niagara, an SDN-based traffic-splitting scheme that achieves accurate traffic splits while being extremely efficient in the use of rule-table space available on commodity switches. Niagara uses an incremental update strategy to minimize the traffic churn given an update. Experiments demonstrate that Niagara (1) achieves nearly optimal accuracy using only 1.2%--37% of the rule space of the current state-of-art, (2) scales to tens of thousands of services with the constrained rule-table capacity and (3) offers nearly minimum churn. Nanxi Kang, Manya Ghobadi, John Reumann, Alexander Shraer, Jennifer Rexford |
CoNEXT | 2 |
| 2015 | TIMELY: RTT-based Congestion Control for the DatacenterabstractDatacenter transports aim to deliver low latency messaging together with high throughput. We show that simple packet delay, measured as round-trip times at hosts, is an effective congestion signal without the need for switch feedback. First, we show that advances in NIC hardware have made RTT measurement possible with microsecond accuracy, and that these RTTs are sufficient to estimate switch queueing. Then we describe how TIMELY can adjust transmission rates using RTT gradients to keep packet latency low while delivering high bandwidth. We implement our design in host software running over NICs with OS-bypass capabilities. We show using experiments with up to hundreds of machines on a Clos network topology that it provides excellent performance: turning on TIMELY for OS-bypass messaging over a fabric with PFC lowers 99 percentile tail latency by 9X while maintaining near line-rate throughput. Our system also outperforms DCTCP running in an optimized kernel, reducing tail latency by $13$X. To the best of our knowledge, TIMELY is the first delay-based congestion control protocol for use in the datacenter, and it achieves its results despite having an order of magnitude fewer RTT signals (due to NIC offload) than earlier delay-based schemes such as Vegas. Radhika Mittal, Vinh The Lam, Nandita Dukkipati, Emily R. Blem, Hassan M. G. Wassel, Manya Ghobadi, Amin Vahdat, Yaogong Wang, David Wetherall, David Zats |
SIGCOMM | 6 |
| 2012 | Rethinking end-to-end congestion control in software-defined networksabstractTCP is designed to operate in a wide range of networks. Without any knowledge of the underlying network and traffic characteristics, TCP is doomed to continuously increase and decrease its congestion window size to embrace changes in network or traffic. In light of emerging popularity of centrally controlled Software-Defined Networks (SDNs), one might wonder whether we can take advantage of the global network view available at the controller to make faster and more accurate congestion control decisions. In this paper, we identify the need and the underlying requirements for a congestion control adaptation mechanism. To this end, we propose OpenTCP as a TCP adaptation framework that works in SDNs. OpenTCP allows network operators to define rules for tuning TCP as a function of network and traffic conditions. We also present a preliminary implementation of OpenTCP in a ~4000 node data center. Manya Ghobadi, Soheil Hassas Yeganeh, Yashar Ganjali |
HotNets | 1 |
| 2012 | Trickle: Rate Limiting YouTube Video Streaming
Manya Ghobadi, Yuchung Cheng, Matthew Mathis |
USENIX ATC | 1 |
| 2011 | Proportional rate reduction for TCPabstractPacket losses increase latency for Web users. Fast recovery is a key mechanism for TCP to recover from packet losses. In this paper, we explore some of the weaknesses of the standard algorithm described in RFC 3517 and the non-standard algorithms implemented in Linux. We find that these algorithms deviate from their intended behavior in the real world due to the combined effect of short flows, application stalls, burst losses, acknowledgment (ACK) loss and reordering, and stretch ACKs. Linux suffers from excessive congestion window reductions while RFC 3517 transmits large bursts under high losses, both of which harm the rest of the flow and increase Web latency. Nandita Dukkipati, Matthew Mathis, Yuchung Cheng, Manya Ghobadi |
Internet Measurement Conference | 4 |
| 2010 | OpenTM: Traffic Matrix Estimator for OpenFlow Networks
Amin Tootoonchian, Manya Ghobadi, Yashar Ganjali |
PAM | 2 |
| 2010 | Caliper: a tool to generate precise and closed-loop trafficabstractGenerating realistic and responsive traffic that reflects different network conditions is a challenging problem associated with performing valid experiments in network testbeds. In this work, we preset Caliper, a highly precise traffic generation tool, built on NetThreads, a flexible platform that we have created for developing packet processing applications on FPGA-based devices and the NetFPGA in particular. We will demonstrate the effect of ad-hoc inter-departure times on a commodity NIC compared to precisely timed inter-departures with Caliper. Both NetThreads and Caliper are available as free software to download. Manya Ghobadi, Martin Labrecque, Geoffrey Salmon, Kaveh Aasaraai, Soheil Hassas Yeganeh, Yashar Ganjali, J. Gregory Steffan |
SIGCOMM | 1 |
| 2009 | Emulation of Optical PIFO BuffersabstractWith recent advances in optical technology, we are closer to building all-optical routers than ever before. A major problem in this area, however, is the lack of all-optical memories similar to what we have in electronics. To overcome this problem, recently, there have been several proposals that show how we can emulate First-In First-Out (FIFO) queues using a combination of fiber delay lines and switches. Unfortunately, FIFO queues cannot be used for implementing many link scheduling policies including weighted fair queuing, weighted round-robin, or strict priority, which are essential components of any modern router today. In this paper, we introduce an architecture based on fiber delay lines and optical switches that can be used for emulating Push-In First-Out (PIFO) queues. In a PIFO queue, an incoming packet can be pushed anywhere in the queue, and therefore it can be used for the implementation of various link scheduling policies. We describe a scheduling algorithm for this architecture and show that with a small speedup, we can build a PIFO queue of size N - 1 using only O(log2N) 3 × 3 optical switches. The resulting system has a minimum reliability of 99.5%, and even for the small portion of departure requests that cannot be fulfilled immediately, the requested packet is ready to depart within approximately five time slots from the request time. Houman Rastegarfar, Manya Ghobadi, Yashar Ganjali |
GLOBECOM | 2 |
| 2008 | Performing time-sensitive network experimentsabstractIt is commonly believed that the Internet has deficiencies that need to be fixed. However, making changes to the current Internet infrastructure is not easy, if possible at all. Any new protocol or design to be implemented on a global scale requires extensive experimental testing in sufficiently realistic settings; simulations alone are not enough. On the other hand, performing network experiments is intrinsically difficult for several reasons: i) Creating a network with multiple routers and a topology that is representative of a real backbone network requires significant resources, ii) Network components have proprietary architectures, which makes it almost impossible to figure out all of their internal details, iii) Making changes to network components is not always possible, iv) We cannot always use real network traces and generating high volumes of artificial traffic which closely resemble operational traffic is not trivial, and v) We need a measurement infrastructure which collects traces and measures various metrics throughout the network. These problems become even more pronounced in the context of time-sensitive network experiments. These are experiments that need very high-precision timings for packet injections into the network, or require packet-level traffic measurements with accurate timing. Experimenting with new congestion control algorithms, buffer sizing in Internet routers, and denial of service attacks which use low-rate packet injections are all examples of time-sensitive experiments, where a subtle variation in packet injection times can change the results significantly. In this work we study the challenges of conducting time-sensitive network experiments in a testbed. We provide a set of guidelines that aim at eliminating sources of inaccuracy in a time-sensitive network experiment. We should note that these guidelines are not meant to be comprehensive. For the sake of space, we only focus on issues that are most likely to be overlooked, and thus unknowingly distort the results of a time-sensitive network experiment. Neda Beheshti, Yashar Ganjali, Manya Ghobadi, Nick McKeown, Jad Naous, Geoffrey Salmon |
ANCS | 3 |
| 2008 | Experimental study of router buffer sizingabstractDuring the past four years, several papers have proposed rules for sizing buffers in Internet core routers. Appenzeller et al. suggest that a link needs a buffer of size O(C/√N), where C is the capacity of the link, and N is the number of flows sharing the link. If correct, buffers could be reduced by 99% in a typical backbone router today without loss in throughput. Enachecsu et al., and Raina et al. suggest that buffers can be reduced even further to 20-50 packets if we are willing to sacrifice a fraction of link capacities, and if there is a large ratio between the speed of core and access links. If correct, this is a five orders of magnitude reduction in buffer sizes. Each proposal is based on theoretical analysis and validated using simulations. Given the potential benefits (and the risk of getting it wrong!) it is worth asking if these results hold in real operational networks. In this paper, we report buffer-sizing experiments performed on real networks - either laboratory networks with commercial routers as well as customized switching and monitoring equipment (UW Madison, Sprint ATL, and University of Toronto), or operational backbone networks (Level 3 Communications backbone network, Internet2, and Stanford). The good news: Subject to the limited scenarios we can create, the buffer sizing results appear to hold. While we are confident that the O(C/√N) will hold quite generally for backbone routers, the 20-50 packet rule should be applied with extra caution to ensure that network components satisfy the underlying assumptions. Neda Beheshti, Yashar Ganjali, Manya Ghobadi, Nick McKeown, Geoffrey Salmon |
Internet Measurement Conference | 3 |
| 2008 | Resource optimization algorithms for virtual private networks using the hose model
Manya Ghobadi, Sudhakar Ganti, Gholamali C. Shoja |
Comput. Networks | 1 |
| 2007 | Hierarchical Provisioning Algorithm for Virtual Private Networks Using the Hose ModelabstractVirtual Private Networks (VPN) provide a secure and reliable communication between customer sites over a shared network. Two models were proposed for the service provisioning in VPNs. The "hose model" for VPNs alleviates the scalability problem of the "pipe model" by reserving bandwidths for aggregate ingress and egress requirements instead of between every pair of VPN endpoints. In this work, VPN endpoints are connected using a tree structure and our algorithm optimizes the total bandwidth reserved on edges of the VPN tree. We introduce a fast and efficient algorithm in finding the shared VPN tree to reduce the total provisioning cost. Our simulation results indicate that the VPN trees constructed by our proposed algorithm reduce bandwidth requirements as compared to previously proposed algorithms while having a much smaller execution time. Manya Ghobadi, Sudhakar Ganti, Gholamali C. Shoja |
GLOBECOM | 1 |
| 2007 | Resource Optimization to Provision a Virtual Private Network Using the Hose ModelabstractVirtual private networks (VPN) provide a secure and reliable communication between customer sites over a shared network. With increase in number and size of VPNs, providers need efficient provisioning techniques that adapt to customer demands. The recently proposed hose model for VPN alleviates the scalability problem of the pipe model by reserving for its aggregate ingress and egress bandwidth instead of between every pair of VPN endpoints. Existing studies on quality of service guarantees in the hose model either deal only with bandwidth requirements or regard the delay requirement as the main objective ignoring the bandwidth cost. In this work we propose a new approach to enhance the hose model to guarantee delay requirements between endpoints while optimizing the provisioning bandwidth cost. We connect VPN endpoints using a tree structure and our algorithm attempts to optimize the total bandwidth reserved on edges of the VPN tree. Our proposed approach takes into account the user preferences in meeting the delay requirements and provisioning cost to find the optimal solution of resource allocation problem. Our experimental results indicate that the VPN trees constructed by our proposed algorithm meet minimum delay requirements while reducing the bandwidth requirements as compared to previously proposed algorithms. Manya Ghobadi, Sudhakar Ganti, Gholamali C. Shoja |
ICC | 1 |