Soudeh Ghorbani

dblp:121/4255 · also Soudeh Ghorbani Khaledi · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-1331-1372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 3 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Orderly Management of Packets in RDMA by Eunomia
Sana Mahmood, Jinqi Lu, Soudeh Ghorbani
APNet3
2025 One to Many: Closing the Bandwidth Gap in AI Datacenters with Scalable Multicast
abstract
AI training now floods datacenter fabrics with thousands of simultaneous collectives, yet most frameworks still move data the hard way: O(N) unicasts for a group of N processors. Classic multicast could slash those bytes but has long been deemed unscalable: computing an optimal tree in an asymmetric Clos is NP-hard and group-specific rules quickly exhaust switch TCAM.
Sepehr Abdous, Jinqi Lu, Jiacheng Wan, Erfan Sharafzadeh, Ying Zhang 0022, Soudeh Ghorbani
HotNets6
2025 Congestion Patterns in a Large-scale RDMA Datacenter
abstract
RDMA datacenters are proliferating to meet the demand of emerging workloads such as AI training and inference as well as distributed storage. This trend has opened up a critical knowledge gap: the traffic characteristics of congestion in these networks remain unknown. We do not know, for example, which layers of the network are the most congested, if the network is load balanced effectively, how long congestion events last, and how accurate existing telemetry systems are in capturing congestion. This paper bridges this gap by investigating congestion in a large-scale RDMA datacenter dedicated to distributed AI training. We provide insights into three specific congestion patterns: (a) location and distribution in the network, (b) burstiness, e.g., the duration and synchrony of bursts, and (c) observability using existing telemetry methods. We show, for instance, that the deployment of Priority Flow Control (PFC) in RDMA networks has shifted the location of congestion one level up: from the edge-host in legacy TCP/IP datacenters to the network core in RDMA datacenters. At the same time, we show that the same protocol enables us to observe and understand congestion better, even bursty events. The findings of this research reveal open challenges for measuring, characterizing, and managing congestion in RDMA networks, paving the way for future research.
Soudeh Ghorbani, Yimeng Zhao, Srikanth Sundaresan, Ying Zhang 0022, Yijing Zeng, Abhigyan Sharma, Prashanth Kannan, Cristian Lumezanu
IMC1
2025 Self-Clocked Round-Robin Packet Scheduling
Erfan Sharafzadeh, Raymond Matson, Jean Tourrilhes, Puneet Sharma 0001, Soudeh Ghorbani
NSDI5
2025 Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity
abstract
Cloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter traffic within datacenters plus the disparity between the RTTs of intra- and inter-datacenter networks complicates congestion management and traffic routing. Particularly, faster congestion responses of intra-datacenter traffic causes rate unfairness when competing with slower inter-datacenter flows. Additionally, inter-datacenter messages suffer from slow loss recovery and, thus, require reliability. Existing solutions overlook these challenges and handle inter- and intra-datacenter congestion with separate control loops or at different granularities. We propose Uno, a unified system for both inter- and intra-DC environments that integrates a transport protocol for rapid congestion reaction and fair rate control with a load balancing scheme that combines erasure coding and adaptive routing. Our findings show that Uno significantly improves the completion times of both inter- and intra-DC flows compared to state-of-the-art methods such as Gemini.
Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini, Nadeen Gebara, Terry Lam, Anup Agarwal, Tiancheng Chen, Zhuolong Yu, Konstantin Taranov, Mahmoud Elhaddad, Daniele De Sensi, Soudeh Ghorbani, Torsten Hoefler
SC13
2025 Intent-Driven Network Management with Multi-Agent LLMs: The Confucius Framework
abstract
Advancements in Large Language Models (LLMs) are significantly transforming network management practices. In this paper, we present our experience developing Confucius, a multi-agent framework for network management at Meta. We model network management workflows as directed acyclic graphs (DAGs) to aid planning. Our framework integrates LLMs with existing management tools to achieve seamless operational integration, employs retrieval-augmented generation (RAG) to improve long-term memory, and establishes a set of primitives to systematically support human/model interaction. To ensure the accuracy of critical network operations, Confucius closely integrates with existing network validation methods and incorporates its own validation framework to prevent regressions. Remarkably, Confucius is a production-ready LLM development framework that has been operational for two years, with over 60 applications onboarded. To our knowledge, this is the first report on employing multi-agent LLMs for hyper-scale networks.
Zhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani, Minlan Yu, Jiawei Zhou 0012, Nathan Hu, Lopa Baruah, Sam Peters, Srikanth Kamath, Jerry Yang, Ying Zhang 0022
SIGCOMM4
2024 A large-scale deployment of DCTCP
Abhishek Dhamija, Balasubramanian Madhavan, Hechao Li, Shrikrishna Khare, Madhavi Rao, Lawrence Brakmo, Neil Spring, Prashanth Kannan, Srikanth Sundaresan, Soudeh Ghorbani
NSDI11
2023 Understanding the impact of host networking elements on traffic bursts
Erfan Sharafzadeh, Sepehr Abdous, Soudeh Ghorbani
NSDI3
2021 Burst-tolerant datacenter networks with Vertigo
abstract
Microsecond-scale congestion events, known as microbursts, are a main cause of packet loss and poor application performance in today's datacenters. Given the low network utilization in datacenters, one would expect packet deflection, in-situ re-routing of packets that arrive at a full buffer to a different port, to effectively prevent packet loss. However, if deployed naively, deflection leads to excessive packet re-ordering, exacerbated congestion, and head-of-the-line blocking in switch buffers. In this study, we resolve the above challenges by selectively deflecting the packets that cause persistent congestion in the network. To enable this, we augment the end-host network stacks with a transport-independent extension that tracks and marks flows with their remaining bytes. Our in-network deflection component uses the flow size information to re-route packets from flows with more data to send. Finally, an extension to the receive-side of end-host stacks retrieves the correct ordering of packets before passing them to transport and higherlevel protocols. We evaluate our design, Vertigo, under diverse datacenter workloads and show that it is effective in managing microbursts under light and heavy loads and when combined with various congestion control algorithms. For example, in a leaf-spine network under 85% load, Vertigo reduces the mean incast query completion times by 3.5x, 3.3x, 5x compared to ECMP, DRILL, and DIBS when using TCP, 3x, 3.5x, 4.5x alongside DCTCP, and 43x, 33x, 16x when using Swift, respectively.
Sepehr Abdous, Erfan Sharafzadeh, Soudeh Ghorbani
CoNEXT3
2021 A high-resolution study of data center traffic at its origin
Erfan Sharafzadeh, Soudeh Ghorbani
CoNEXT2
2020 Liveness Verification of Stateful Network Functions
Farnaz Yousefi, Anubhavnidhi Abhashkumar, Kausik Subramanian, Kartik Hans, Soudeh Ghorbani, Aditya Akella
NSDI5
2017 COCONUT: Seamless Scale-out of Network Elements
abstract
A key use of software-defined networking is to enable scale-out of network data plane elements. Naively scaling networking elements, however, can cause incorrect behavior. For example, we show that an IDS system which operates correctly as a single network element can erroneously and permanently block hosts when it is replicated.
Soudeh Ghorbani, Brighten Godfrey
EuroSys1
2017 DRILL: Micro Load Balancing for Low-latency Data Center Networks
abstract
The trend towards simple datacenter network fabric strips most network functionality, including load balancing, out of the network core and pushes it to the edge. This slows reaction to microbursts, the main culprit of packet loss in datacenters. We investigate the opposite direction: could slightly smarter fabric significantly improve load balancing? This paper presents DRILL, a datacenter fabric for Clos networks which performs micro load balancing to distribute load as evenly as possible on microsecond timescales. DRILL employs per-packet decisions at each switch based on local queue occupancies and randomized algorithms to distribute load. Our design addresses the resulting key challenges of packet reordering and topological asymmetry. In simulations with a detailed switch hardware model and realistic workloads, DRILL outperforms recent edge-based load balancers, particularly under heavy load. Under 80% load, for example, it achieves 1.3-1.4x lower mean flow completion time than recent proposals, primarily due to shorter upstream queues. To test hardware feasibility, we implement DRILL in Verilog and estimate its area overhead to be less than 1%. Finally, we analyze DRILL's stability and throughput-efficiency.
Soudeh Ghorbani, Zibin Yang, Brighten Godfrey, Yashar Ganjali, Amin Firoozshahian
SIGCOMM1
2015 Micro Load Balancing in Data Centers with DRILL
abstract
The trend towards simple data center network fabric strips most network functionality, including load balancing capabilities, out of the network core and pushes them to the edge. We investigate a different direction of incorporating minimal load balancing intelligence into the network fabric and show that this slightly smarter fabric significantly enhances performance. We provide a very simple in-network load balancing scheduling algorithm called DRILL which is purely local to each switch. DRILL leverages local load sensing and randomization concepts to distribute load among multiple paths. Through simulation, we show that this simple approach outperforms CONGA, a recent global edge-based load balancing scheme for data centers. We also formally prove the switch-level stability and throughput-efficiency of DRILL's scheduling algorithm.
Soudeh Ghorbani, Brighten Godfrey, Yashar Ganjali, Amin Firoozshahian
HotNets1
2014 Transparent, Live Migration of a Software-Defined Network
abstract
Increasingly, datacenters are virtualized and software-defined. Live virtual machine (VM) migration is becoming an indispensable management tool in such environments. However, VMs often have a tight coupling with the underlying network. Hence, cloud providers are beginning to offer tenants more control over their virtual networks. Seamless migration of all (or part) of a virtual network greatly simplifies management tasks like planned maintenance, optimizing resource usage, and cloud bursting. Our LIME architecture efficiently migrates an ensemble, a collection of virtual machines and virtual switches, for any arbitrary controller and end-host applications. To minimize performance disruptions, during the migration, LIME temporarily runs all or part of a virtual switch on multiple physical switches. Running a virtual switch on multiple physical switches must be done carefully to avoid compromising application correctness. To that end, LIME merges events, combines traffic statistics, and preserves consistency among multiple physical switches even across changes to the packet-handling rules. Using a formal model, we prove that migration under LIME is transparent to applications, i.e., any execution of the controller and end-host applications during migration is a completely valid execution that could have taken place in a migration-free setting. Experiments with our prototype, built on the Floodlight controller, show that ensemble migration can be an efficient tool for network management.
Soudeh Ghorbani, Cole Schlesinger, Matthew Monaco, Eric Keller, Matthew Caesar 0001, Jennifer Rexford, David Walker 0001
SoCC1
2012 Live migration of an entire network (and its hosts)
abstract
Live virtual machine (VM) migration can move applications from one location to another without a disruption in service. However, applications often consist of multiple VMs and rely on the state of the underlying network for basic reachability, access control, and QoS functionality. Rather than migrating an individual VM, we show how to migrate an ensemble---the VMs, the network, and the management system---to a different set of physical resources. Our LIME (LIve Migration of Ensembles) design leverages recent advances in Software Defined Networking (SDN) for a clear separation between the controller and the data-plane state in the switches. Transparent to the application running on the controller, LIME clones the data-plane state to a new set of switches, and then incrementally migrates the traffic sources (e.g., the VMs). During this transition, both networks deliver traffic and LIME maintains synchronized state. Experiments with our initial prototype, built on the Floodlight OpenFlow controller, suggest that network migration does not have to be a disruptive, middle-of-the-night maintenance event, but can become an integral network management mechanism completely transparent to applications.
Eric Keller, Soudeh Ghorbani, Matthew Caesar 0001, Jennifer Rexford
HotNets2