Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rami Nudelman

dblp:386/2133 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0002-3601-2786ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 30% Hardware accelerators and domain-specific architectures · 30% Distributed systems · 29%
Computer networks
1 paper
Datacenter networks · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Datacenter networks
RDMA
0.912025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025
High-performance computing
collective communication
0.812024
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI · SC 2024
Hardware accelerators and domain-specific architectures › network accelerator
SmartNIC
0.812024
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI · SC 2024
Distributed systems › distributed machine learning
distributed training
0.522025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI · SC 2024
Cloud and datacenter computing › cloud networking
inter-datacenter network
0.312025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025
Distributed systems › communication optimization
communication-computation overlap
0.212024
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI · SC 2024

Methods — techniques the papers use, named apart from their topics

selective repeat · 1.7erasure coding · 1.7hardware multicast · 0.8SmartNIC offloading · 0.8
YearPublicationVenuePosition
2025 SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
abstract
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training.
Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler
SC9
2024 Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
abstract
In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstanding operations such as Allgather and Reduce-Scatter can compete for the injection bandwidth and create pipeline bubbles. To address this problem, we propose a novel bandwidth-optimal Allgather collective algorithm that leverages hardware multicast. We use multicast to build a constant-time reliable Broadcast protocol, a building block for constructing an optimal Allgather schedule. Our Allgather algorithm achieves $2 \times$ traffic reduction on a 188 -node testbed. To free the host side from running the protocol, we employ SmartNIC offloading. We extract the parallelism in our Allgather algorithm and map it to a SmartNIC specialized for hiding the cost of data movement. We show that our SmartNIC-offloaded collective progress engine can scale to the next generation of 1.6 Tbit/s links.
Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman, Gil Bloch, Torsten Hoefler
SC4