EDBT 2026 Demo / reviewers in the wild / expert
Tommaso Bonato
dblp:329/3681
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-2345-473XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure MitigationabstractNext-generation datacenters require highly efficient network load balancing to manage the growing scale of artificial intelligence (AI) training and general datacenter traffic. However, existing Ethernet-based solutions, such as Equal Cost Multi-Path (ECMP) and oblivious packet spraying (OPS), struggle to maintain high network utilization due to both increasing traffic demands and the expanding scale of datacenter topologies, which also exacerbate network failures. To address these limitations, we propose REPS, a lightweight decentralized per-packet adaptive load balancing algorithm designed to optimize network utilization while ensuring rapid recovery from link failures. REPS adapts to network conditions by caching good-performing paths. In case of a network failure, REPS re-routes traffic away from it in less than 100 microseconds. REPS is designed to be deployed with next-generation out-of-order transports, such as Ultra Ethernet, and uses less than 25 bytes of per-connection state regardless of the topology size. We extensively evaluate REPS in large-scale simulations and FPGA-based NICs. Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Michael Papamichael, Mohammad Dohadwala, Lukas Gianinazzi, Mikhail Khalilov, Elias Achermann, Daniele De Sensi, Torsten Hoefler |
EuroSys | 1 |
| 2026 | Spritz: Path-Aware Load Balancing in Low-Diameter Networks
Tommaso Bonato, Ales Kubicek, Abdul Kabbani, Ahmad Ghalayini, Maciej Besta, Torsten Hoefler |
IPDPS | 1 |
| 2026 | Ampel: Scheduling at the Network Cut in ML TrainingabstractThe size and communication patterns of modern ML training workloads place significant strain on datacenter fabrics. When bandwidth demand exceeds capacity, flows experience slowdowns and iteration time grows larger. Fine-grained load-balancing such as packet spraying cannot fully resolve this issue, yet it can shift the bottleneck from individual links to groups of links partitioning the network. In this work, we present Ampel: a system to schedule ML training flows at the network cuts. It leverages packet-spraying's ability to spread traffic evenly across all available paths to simplify the view of the topology into one only containing potential bottlenecks, and then bridges this new simplified model with past work on coflow scheduling. Our simulated experiments show that Ampel can reduce average training iteration time by up to 18% compared to state-of-the-art ML schedulers. Valerio Torsiello, Ayush Mishra, Sushovan Das, Lukas Röllin, Tommaso Bonato, Torsten Hoefler, Laurent Vanbever |
SIGCOMM | 5 |
| 2026 | Flowcut Switching: High-Performance Adaptive Routing With In-Order Delivery GuaranteesabstractNetwork latency severely impacts the performance of applications running on supercomputers. Adaptive routing algorithms route packets over different available paths to reduce latency and improve network utilization. However, if a switch routes packets belonging to the same network flow on different paths, they might arrive at the destination out-of-order due to differences in the latency of these paths. For some transport protocols like TCP, QUIC, and RoCE, out-of-order (OOO) packets might cause large performance drops or significantly increase CPU utilization. In this work, we proposeFlowcut switching, a new adaptive routing algorithm that provides high-performance in-order packet delivery. Differently from existing solutions likeFlowlet switching, which are based on the assumption of bursty traffic and that might still reorder packets,Flowcut switchingguarantees in-order delivery under any network conditions, and is effective also for non-bursty traffic, as it is often the case for RDMA. On top of this, Flowcut can be implemented either at the switch or NIC level providing flexibility and different tradeoffs. Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Abdulla Bataineh, David Hewson, Duncan Roweth, Torsten Hoefler |
IEEE Trans. Netw. | 1 |
| 2025 | Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and AlgorithmsabstractThe NVIDIA Collective Communication Library (NCCL) is a critical software layer enabling high-performance collectives on large-scale GPU clusters. Despite being open source with a documented API, its internal design remains largely opaque. The orchestration of communication channels, selection of protocols, and handling of memory movement across devices and nodes are not clearly understood, making it difficult to analyze performance or identify bottlenecks. This paper presents a comprehensive analysis of NCCL, focusing on its communication protocol variants (Simple, LL, and LL128), the mechanisms governing intra-node and inter-node data movement, and ring-and tree-based collective communication algorithms. The insights obtained from this study serve as the foundation for ATLAHS, an application-trace-driven network simulation toolchain capable of accurately reproducing NCCL communication patterns in large-scale AI training workloads. By demystifying NCCL's internal architecture, this work provides guidance for system researchers and performance engineers working to optimize or simulate collective communication at scale. Zhiyi Hu, Tommaso Bonato, Sylvain Jeaugey, Cedell Alexander, Eric Spada, James Dinan, Jeff R. Hammond, Torsten Hoefler |
HOTI | 3 |
| 2025 | Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable ConnectivityabstractCloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter traffic within datacenters plus the disparity between the RTTs of intra- and inter-datacenter networks complicates congestion management and traffic routing. Particularly, faster congestion responses of intra-datacenter traffic causes rate unfairness when competing with slower inter-datacenter flows. Additionally, inter-datacenter messages suffer from slow loss recovery and, thus, require reliability. Existing solutions overlook these challenges and handle inter- and intra-datacenter congestion with separate control loops or at different granularities. We propose Uno, a unified system for both inter- and intra-DC environments that integrates a transport protocol for rapid congestion reaction and fair rate control with a load balancing scheme that combines erasure coding and adaptive routing. Our findings show that Uno significantly improves the completion times of both inter- and intra-DC flows compared to state-of-the-art methods such as Gemini. Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini, Nadeen Gebara, Terry Lam, Anup Agarwal, Tiancheng Chen, Zhuolong Yu, Konstantin Taranov, Mahmoud Elhaddad, Daniele De Sensi, Soudeh Ghorbani, Torsten Hoefler |
SC | 1 |
| 2025 | Bine Trees: Enhancing Collective Operations by Optimizing Communication LocalityabstractCommunication locality plays a key role in the performance of collective operations on large HPC systems, especially on oversubscribed networks where groups of nodes are fully connected internally but sparsely linked through global connections. We present Bine (binomial negabinary) trees, a family of collective algorithms that improve communication locality. Bine trees maintain the generality of binomial trees and butterflies while cutting global-link traffic by up to \(33\%\). We implement eight Bine-based collectives and evaluate them on four large-scale supercomputers with Dragonfly, Dragonfly+, oversubscribed fat-tree, and torus topologies, achieving up to 5 × speedups and consistent reductions in global-link traffic across different vector sizes and node counts. Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato, Seydou Ba, Matteo Turisini, Jens Domke, Torsten Hoefler |
SC | 4 |
| 2025 | ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed StorageabstractNetwork simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the Group Operation Assembly Language (GOAL) format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of both simulation runtime and trace size efficiency. We further illustrate ATLAHS’s utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes. Tommaso Bonato, Zhiyi Hu, Pasquale Jordan, Tiancheng Chen, Torsten Hoefler |
SC | 2 |
| 2024 | Swing: Short-cutting Rings for Higher Bandwidth Allreduce
Daniele De Sensi, Tommaso Bonato, David Saam, Torsten Hoefler |
NSDI | 2 |
| 2022 | HammingMesh: A Network Topology for Large-Scale Deep LearningabstractNumerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the AI revolution. With the exhaustion of such optimizations, the growth of modern AI is now gated by the performance of training systems, especially their data movement. Instead of focusing on single accelerators, we investigate data-movement characteristics of large-scale training at full system scale. Based on our workload analysis, we design HammingMesh, a novel network topology that provides high bandwidth at low cost with high job scheduling flexibility. Specifically, HammingMesh can support full bandwidth and isolation to deep learning training jobs with two dimensions of parallelism. Furthermore, it also supports high global bandwidth for generic traffic. Thus, HammingMesh will power future large-scale deep learning systems with extreme bandwidth requirements. Torsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Shigang Li 0002, Marco Heddes, Jon Belk, Deepak Goel, Miguel Castro 0001, Steve Scott |
SC | 2 |