VLDB 2026 Research / reviewers in the wild / expert
Scott Moe
dblp:246/4286
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0007-8500-5810ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
1 paper |
Datacenter networks · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 50% Distributed systems · 50% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Datacenter networks
RDMA |
0.9 | 1 | 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025 |
Distributed systems › distributed machine learning
distributed training |
0.3 | 1 | 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025 |
Cloud and datacenter computing › cloud networking
inter-datacenter network |
0.3 | 1 | 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025 |
Methods — techniques the papers use, named apart from their topics
selective repeat · 1.7erasure coding · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA CommunicationabstractRDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training. Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler |
SC | 17 |