Abdul Kabbani

dblp:91/4608 · also Abdul Kader Kabbani · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-7106-3824ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 3 first-author · 3 since 2021Systems, architecture and hardware · 8 · 4 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure Mitigation
abstract
Next-generation datacenters require highly efficient network load balancing to manage the growing scale of artificial intelligence (AI) training and general datacenter traffic. However, existing Ethernet-based solutions, such as Equal Cost Multi-Path (ECMP) and oblivious packet spraying (OPS), struggle to maintain high network utilization due to both increasing traffic demands and the expanding scale of datacenter topologies, which also exacerbate network failures. To address these limitations, we propose REPS, a lightweight decentralized per-packet adaptive load balancing algorithm designed to optimize network utilization while ensuring rapid recovery from link failures. REPS adapts to network conditions by caching good-performing paths. In case of a network failure, REPS re-routes traffic away from it in less than 100 microseconds. REPS is designed to be deployed with next-generation out-of-order transports, such as Ultra Ethernet, and uses less than 25 bytes of per-connection state regardless of the topology size. We extensively evaluate REPS in large-scale simulations and FPGA-based NICs.
Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Michael Papamichael, Mohammad Dohadwala, Lukas Gianinazzi, Mikhail Khalilov, Elias Achermann, Daniele De Sensi, Torsten Hoefler
EuroSys2
2026 Spritz: Path-Aware Load Balancing in Low-Diameter Networks
Tommaso Bonato, Ales Kubicek, Abdul Kabbani, Ahmad Ghalayini, Maciej Besta, Torsten Hoefler
IPDPS3
2025 Mitigating Inter-datacenter Incast with a Proxy: The shortest path is not necessarily the fastest
abstract
Many-to-one communication (i.e., incast) is a long-standing challenge in networking with a wide range of proposed solutions. However, as incast-inducing applications today (e.g., storage, ML training) scale beyond a single datacenter, they introduce new challenges that current solutions do not handle. In particular, inter-datacenter links have orders of magnitude higher latency than intra-datacenter paths, lengthening the feedback loop that senders rely on to adjust their sending rates and drastically increasing incast completion times.
Anchengcheng Zhou, Carter Costic, Hongyu Hè, Ahmad Ghalayini, Abdul Kabbani, Maria Apostolaki
HotNets5
2025 Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity
abstract
Cloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter traffic within datacenters plus the disparity between the RTTs of intra- and inter-datacenter networks complicates congestion management and traffic routing. Particularly, faster congestion responses of intra-datacenter traffic causes rate unfairness when competing with slower inter-datacenter flows. Additionally, inter-datacenter messages suffer from slow loss recovery and, thus, require reliability. Existing solutions overlook these challenges and handle inter- and intra-datacenter congestion with separate control loops or at different granularities. We propose Uno, a unified system for both inter- and intra-DC environments that integrates a transport protocol for rapid congestion reaction and fair rate control with a load balancing scheme that combines erasure coding and adaptive routing. Our findings show that Uno significantly improves the completion times of both inter- and intra-DC flows compared to state-of-the-art methods such as Gemini.
Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini, Nadeen Gebara, Terry Lam, Anup Agarwal, Tiancheng Chen, Zhuolong Yu, Konstantin Taranov, Mahmoud Elhaddad, Daniele De Sensi, Soudeh Ghorbani, Torsten Hoefler
SC3
2025 SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
abstract
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training.
Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler
SC12
2023 Improving Network Availability with Protective ReRoute
abstract
We present PRR (Protective ReRoute), a transport technique for shortening user-visible outages that complements routing repair. It can be added to any transport to provide benefits in multipath networks. PRR responds to flow connectivity failure signals, e.g., retransmission timeouts, by changing the FlowLabel on packets of the flow, which causes switches and hosts to choose a different network path that may avoid the outage. To enable it, we shifted our IPv6 network architecture to use the FlowLabel, so that hosts can change the paths of their flows without application involvement. PRR is deployed fleetwide at Google for TCP and Pony Express, where it has been protecting all production traffic for several years. It is also available to our Cloud customers. We find it highly effective for real outages. In a measurement study on our network backbones, adding PRR reduced the cumulative region-pair outage time for RPC traffic by 63--84%. This is the equivalent of adding 0.4--0.8 "nines" of availability.
David Wetherall, Abdul Kabbani, Van Jacobson, Jim Winget, Yuchung Cheng, Charles B. Morrey III, Uma Parthavi Moravapalle, Phillipa Gill, Steven Knight, Amin Vahdat
SIGCOMM2
2022 PLB: congestion signals are simple and effective for network load balancing
abstract
We present a new, host-based design for link load balancing and report the first experiences of link imbalance in datacenters. Our design, PLB (Protective Load Balancing), builds on transport protocols and ECMP/WCMP to reduce network hotspots. PLB randomly changes the paths of connections that experience congestion, preferring to repath after idle periods to minimize packet reordering. It repaths a connection by changing the IPv6 Flow Label on its packets, which switches include as part of ECMP/WCMP. Across hosts, this action drives down hotspots in the network, and lowers the latency of RPCs.
Mubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu, Gautam Kumar 0001, Masoud Moshref, Junhua Yan, Van Jacobson, David Wetherall, Abdul Kabbani
SIGCOMM10
2020 Annulus: A Dual Congestion Control Loop for Datacenter and WAN Traffic Aggregates
abstract
Cloud services are deployed in datacenters connected though high-bandwidth Wide Area Networks (WANs). We find that WAN traffic negatively impacts the performance of datacenter traffic, increasing tail latency by 2.5x, despite its small bandwidth demand. This behavior is caused by the long round-trip time (RTT) for WAN traffic, combined with limited buffering in datacenter switches. The long WAN RTT forces datacenter traffic to take the full burden of reacting to congestion. Furthermore, datacenter traffic changes on a faster time-scale than the WAN RTT, making it difficult for WAN congestion control to estimate available bandwidth accurately.
Ahmed Saeed 0001, Prateesh Goyal, Milad Sharif, Mostafa H. Ammar, Ellen Zegura, Keon Jang, Mohammad Alizadeh, Abdul Kabbani, Amin Vahdat
SIGCOMM10
2017 Flier: Flow-level congestion-aware routing for direct-connect data centers
abstract
Various topologies have been proposed in the context of high-performance computing and data center networking. Direct-connect topologies generally offer large capacity with high path diversity and are highly cost effective for general data center traffic patterns. However, the lack of simple yet efficient load balancing techniques for direct-connect fabrics has hindered these networks from gaining traction in data centers. This paper presents the design, implementation, and evaluation of Flicr, a light-weight host-based load balancing mechanism for direct-connect data centers. Flicr dynamically reroutes traffic through minimal and non-minimal routes to avoid congesting the fabric. This enables Flicr to efficiently minimize networking resource consumption while exploiting high path diversity in direct-connect fabrics to balance the network and gracefully handle link failures. Flicr requires only a simple kernel modification and is readily deployable in commodity data centers today. Our evaluations show that Flicr consistently outperforms other state-of-the-art load balancing designs, achieving 25-60% lower average flow completion time compared to adaptive routing. Flicr is also more robust against link failures and has 5-8 χ better performance relative to other schemes in the presence of link failures.
Abdul Kabbani, Milad Sharif
INFOCOM1
2016 Juggler: a practical reordering resilient network stack for datacenters
abstract
We present Juggler, a practical reordering resilient network stack for datacenters that enables any packet to be sent on any path at any level of priority. Juggler adds functionality to the Generic Receive Offload layer at the entry of the network stack to put packets in order in a best-effort fashion. Juggler's design exploits the small packet delays in datacenter networks and the inherent burstiness of traffic to eliminate the negative effects of packet reordering almost entirely while keeping state for only a small number of flows at any given time. Extensive testbed experiments at 10Gb/s and 40Gb/s speeds show that Juggler is effective and lightweight: it prevents performance loss even with severe packet reordering while imposing low CPU overhead. We demonstrate the use of Juggler for per-packet multi-path load balancing and a novel system that provides bandwidth guarantees by dynamically prioritizing packets.
Yilong Geng, Vimalkumar Jeyakumar, Abdul Kabbani, Mohammad Alizadeh
EuroSys3
2014 FlowBender: Flow-level Adaptive Routing for Improved Latency and Throughput in Datacenter Networks
abstract
Datacenter networks provide high path diversity for traffic between machines. Load balancing traffic across these paths is important for both, latency- and throughput-sensitive applications. The standard load balancing techniques used today obliviously hash a flow to a random path. When long flows collide on the same path, this might lead to long lasting congestion while other paths could be underutilized, degrading performance of other flows as well. Recent proposals to address this shortcoming incur significant implementation complexity at the host that would actually slow down short flows (MPTCP), depend on relatively slow centralized controllers for rerouting large congesting flows (Hedera), or require custom switch hardware, hindering near-term deployment (DeTail).
Abdul Kabbani, Balajee Vamanan, Jahangir Hasan, Fabien Duchene 0001
CoNEXT1
2014 WCMP: weighted cost multipathing for improved fairness in data centers
abstract
Data Center topologies employ multiple paths among servers to deliver scalable, cost-effective network capacity. The simplest and the most widely deployed approach for load balancing among these paths, Equal Cost Multipath (ECMP), hashes flows among the shortest paths toward a destination. ECMP leverages uniform hashing of balanced flow sizes to achieve fairness and good load balancing in data centers. However, we show that ECMP further assumes a balanced, regular, and fault-free topology, which are invalid assumptions in practice that can lead to substantial performance degradation and, worse, variation in flow bandwidths even for same size flows.
Junlan Zhou, Malveeka Tewari, Abdul Kabbani, Leonid B. Poutievski, Amin Vahdat
EuroSys4
2014 SENIC: Scalable NIC for End-Host Rate Limiting
Sivasankar Radhakrishnan, Yilong Geng, Vimalkumar Jeyakumar, Abdul Kabbani, George Porter, Amin Vahdat
NSDI4
2012 Less Is More: Trading a Little Bandwidth for Ultra-Low Latency in the Data Center
Mohammad Alizadeh, Abdul Kabbani, Tom Edsall, Balaji Prabhakar, Amin Vahdat, Masato Yasuda
NSDI2
2011 Stability analysis of QCN: the averaging principle
abstract
Data Center Networks have recently caused much excitement in the industry and in the research community. They represent the convergence of networking, storage, computing and virtualization. This paper is concerned with the Quantized Congestion Notification (QCN) algorithm, developed for Layer 2 congestion management. QCN has recently been standardized as the IEEE 802.1Qau Ethernet Congestion Notification standard.
Mohammad Alizadeh, Abdul Kabbani, Berk Atikoglu, Balaji Prabhakar
SIGMETRICS2
2008 Counter braids: a novel counter architecture for per-flow measurement
abstract
Fine-grained network measurement requires routers and switches to update large arrays of counters at very high link speed (e.g. 40 Gbps). A naive algorithm needs an infeasible amount of SRAM to store both the counters and a flow-to-counter association rule, so that arriving packets can update corresponding counters at link speed. This has made accurate per-flow measurement complex and expensive, and motivated approximate methods that detect and measure only the large flows.This paper revisits the problem of accurate per-flow measurement. We present a counter architecture, called Counter Braids, inspired by sparse random graph codes. In a nutshell, Counter Braids compresses while counting. It solves the central problems (counter space and flow-to-counter association) of per-flow measurement by braiding a hierarchy of counters with random graphs. Braiding results in drastic space reduction by sharing counters among flows; and using random graphs generated on-the-fly with hash functions avoids the storage of flow-to-counter association.The Counter Braids architecture is optimal (albeit with a complex decoder) as it achieves the maximum compression rate asymptotically. For implementation, we present a low-complexity message passing decoding algorithm, which can recover flow sizes with essentially zero error. Evaluation on Internet traces demonstrates that almost all flow sizes are recovered exactly with only a few bits of counter space per flow.
Yi Lu 0001, Andrea Montanari, Balaji Prabhakar, Sarang Dharmapurikar, Abdul Kabbani
SIGMETRICS5
2007 Distributed Low-Complexity Maximum-Throughput Scheduling for Wireless Backhaul Networks
abstract
We introduce a low-complexity distributed slotted MAC protocol that can support all feasible arrival rates in a wireless backhaul network (WBN). For arbitrary wireless networks, such a maximum throughput protocol has been notoriously hard to realize because even if global topology information is available, the problem of computing the optimal link transmission set at each slot is NP-complete. For the logical tree structures induced by WBN traffic matrices, we first introduce a centralized algorithm that solves the optimal scheduling problem in a number of steps at most linear in the number of nodes in the network. This is achieved by discovering and exploiting a novel set of graph-theoretical properties of WBN contention graph. Guided by the centralized algorithm, we design a distributed protocol where, at the beginning of each slot, nodes coordinate and incrementally compute the optimal link transmission set. We then introduce an algorithm to compute the minimum number of steps to complete this computation, thus minimizing the per-slot overhead. Using both analysis and simulations, we show that in practice our protocol yields low overhead when implemented over existing wireless technologies and significantly outperforms existing suboptimal distributed slotted scheduling mechanisms.
Abdul Kabbani, Theodoros Salonidis, Edward W. Knightly
INFOCOM1