VLDB 2026 Research / reviewers in the wild / expert
Mahmoud Elhaddad
dblp:05/7032
· DBLP profile ↗
8ranked-venue papers
4as first author
4since 2021 · last 2025
0009-0000-0335-7898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
3 papers |
Datacenter networks · 61% Network optimization and economics · 22% Network performance modeling · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 39% Storage systems · 36% Distributed systems · 14% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Datacenter networks
RDMA |
1.5 | 2 | 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025 Understanding RDMA Microarchitecture Resources for Performance Isolation · NSDI 2023 |
Datacenter networks
load balancing |
0.9 | 1 | 2025 | Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity · SC 2025 |
Network optimization and economics › resource allocation › fair resource allocation
rate fairness |
0.9 | 1 | 2025 | Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity · SC 2025 |
Network performance modeling
performance isolation |
0.7 | 1 | 2023 | Understanding RDMA Microarchitecture Resources for Performance Isolation · NSDI 2023 |
Storage systems › networked storage › storage networking
RDMA storage |
0.7 | 1 | 2023 | Empowering Azure Storage with RDMA · NSDI 2023 |
Cloud and datacenter computing › datacenter network
datacenter communication |
0.3 | 1 | 2025 | Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity · SC 2025 |
Distributed systems › distributed machine learning
distributed training |
0.3 | 1 | 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025 |
Cloud and datacenter computing › cloud networking
inter-datacenter network |
0.3 | 1 | 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025 |
Cloud and datacenter computing
cloud storage |
0.2 | 1 | 2023 | Empowering Azure Storage with RDMA · NSDI 2023 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2023 | Understanding RDMA Microarchitecture Resources for Performance Isolation · NSDI 2023 |
Methods — techniques the papers use, named apart from their topics
erasure coding · 3.5selective repeat · 1.7adaptive routing · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable ConnectivityabstractCloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter traffic within datacenters plus the disparity between the RTTs of intra- and inter-datacenter networks complicates congestion management and traffic routing. Particularly, faster congestion responses of intra-datacenter traffic causes rate unfairness when competing with slower inter-datacenter flows. Additionally, inter-datacenter messages suffer from slow loss recovery and, thus, require reliability. Existing solutions overlook these challenges and handle inter- and intra-datacenter congestion with separate control loops or at different granularities. We propose Uno, a unified system for both inter- and intra-DC environments that integrates a transport protocol for rapid congestion reaction and fair rate control with a load balancing scheme that combines erasure coding and adaptive routing. Our findings show that Uno significantly improves the completion times of both inter- and intra-DC flows compared to state-of-the-art methods such as Gemini. Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini, Nadeen Gebara, Terry Lam, Anup Agarwal, Tiancheng Chen, Zhuolong Yu, Konstantin Taranov, Mahmoud Elhaddad, Daniele De Sensi, Soudeh Ghorbani, Torsten Hoefler |
SC | 11 |
| 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA CommunicationabstractRDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training. Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler |
SC | 18 |
| 2023 | Empowering Azure Storage with RDMA
Wei Bai 0001, Shanim Sainul Abdeen, Ankit Agrawal 0013, Krishan Kumar Attre, Paramvir Bahl, Ameya Bhagat, Gowri Bhaskara, Tanya Brokhman, Ahmad Cheema, Rebecca Chow, Jeff Cohen, Mahmoud Elhaddad, Vivek Ette, Igal Figlin, Daniel Firestone, Mathew George, Ilya German, Lakhmeet Ghai, Eric Green, Albert G. Greenberg, Randy Haagens, Matthew Hendel, Ridwan Howlader, Neetha John, Julia Johnstone, Tom Jolly, Greg Kramer, David Kruse, Erica Lan, Avi Levy, Marina Lipshteyn, Guohan Lu, Yuemin Lu, Xiakun Lu, Vadim Makhervaks, Ulad Malashanka, David A. Maltz, Ilias Marinos, Rohan Mehta, Sharda Murthi, Anup Namdhari, Aaron Ogus, Jitendra Padhye, Madhav Pandya, Douglas Phillips, Adrian Power, Suraj Puri, Shachar Raindel, Jordan Rhee, Anthony Russo, Maneesh Sah, Ali Sheriff, Chris Sparacino, Ashutosh Srivastava, Weixiang Sun, Nick Swanson, Fuhou Tian, Lukasz Tomczyk, Vamsi Vadlamuri, Alec Wolman, Joyce Yom, Yanzhao Zhang, Brian Zill |
NSDI | 13 |
| 2023 | Understanding RDMA Microarchitecture Resources for Performance Isolation
Xinhao Kong, Jingrong Chen 0002, Wei Bai 0001, Yechen Xu, Mahmoud Elhaddad, Shachar Raindel, Jitendra Padhye, Alvin R. Lebeck, Danyang Zhuo |
NSDI | 5 |
| 2008 | On the Emulation of Finite-Buffered Output Queued Switches Using Combined Input-Output Queuing
Mahmoud Elhaddad, Rami G. Melhem |
DISC | 1 |
| 2007 | Scheduling to Minimize theWorst-Case Loss RateabstractWe study link scheduling in networks with small router buffers, with the goal of minimizing the guaranteed packet loss rate bound for each ingress-egress traffic aggregate (connection). Given a link scheduling algorithm (a service discipline and a packet drop policy), the guaranteed loss rate for a connection is the loss rate under worst-case routing and bandwidth allocations for competing traffic. Under simplifying assumptions, we show that a local min-max fairness property with respect to apportioning loss events among the connections sharing each link, and a condition on the correlation of scheduling decisions at different links are two necessary and (together) sufficient conditions for optimality in the minimization problem. Based on these conditions, we introduce a randomized link-scheduling algorithm called rolling priority where packet scheduling at each link relies exclusively on local information. We show that RP satisfies both conditions and is therefore optimal. Mahmoud Elhaddad, Hammad Iqbal, Taieb Znati, Rami G. Melhem |
ICDCS | 1 |
| 2006 | Supporting Loss Guarantees in Buffer-Limited NetworksabstractWe consider the problem of packet scheduling in a network with small router buffers. The objective is to provide a statistical bound on the worst-case packet loss rate for a traffic aggregate (connection) routed along any network path, given maximum permissible link utilization (load). This problem is argued to be of interest in networks providing statistical loss-rate guarantees to ingress-egress connections with fixed bandwidth demands. We introduce a scheduling algorithm for networks using per packet transmission reservation. Reservations allow loss guarantees at the aggregate level to hold for individual flows within the aggregate. The algorithm employs randomization and traffic regulation at the ingress, and batch local scheduling at the links. It ensures that a large fraction of packets from each connection are consistently subject to small loss probability at every link. These packets are therefore likely to survive long paths. To obtain the desired loss-rate bound, we analyze the performance of the algorithm under global routing and bandwidth allocation scenarios that maximize the loss rate of a connection routed along an arbitrary network path. We compare the bound to that obtained using the scheduling algorithm that combines the FCFS service discipline and the drop-tail policy. We find that the proposed algorithm significantly improves the constraints on link utilization and path length necessary to achieve strong loss-rate guarantees Mahmoud Elhaddad, Rami G. Melhem, Taieb Znati |
IWQoS | 1 |
| 2004 | Decoupling Packet Loss from Blocking in Proactive Reservation-Based SwitchingabstractWe consider the maximization of network throughput in buffer-constrained optical networks using aggregate bandwidth allocation and reservation-based transmission control. Assuming that all flows are subject to loss-based TCP congestion control, we quantify the effects of buffer capacity constraints on bandwidth utilization efficiency through contention-induced packet loss. The analysis shows that the ability of TCP flows to efficiently utilize successful reservations is highly sensitive to the available buffer capacity. Maximizing the bandwidth utilization efficiency under buffer capacity constraints thus requires decoupling packet loss from contention-induced blocking of transmission requests. We describe a confirmed (two-way) reservation scheme that eliminates contention-induced loss, so that no packets are dropped at the network's core, and loss can be incurred only at the adequately buffer-provisioned ingress routers, where it is exclusively congestion-induced. For the confirmed signaling scheme, analytical and simulation results indicate that TCP aggregates are able to efficiently utilize the successful reservations independently of buffer constraints. Mahmoud Elhaddad, Rami G. Melhem, Taieb Znati |
BROADNETS | 1 |