Peter-Jan Gootzen

dblp:309/6238 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0008-8809-254XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
1 paper
Datacenter networks · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 50% Distributed systems · 50%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Datacenter networks
RDMA
0.912025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025
Distributed systems › distributed machine learning
distributed training
0.312025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025
Cloud and datacenter computing › cloud networking
inter-datacenter network
0.312025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication · SC 2025

Methods — techniques the papers use, named apart from their topics

selective repeat · 1.7erasure coding · 1.7
YearPublicationVenuePosition
2026 Kernel-Initiated One-Sided Networking for GPU-Accelerated AI Workloads
Khaled Hamidouche, John Bachan, Pak Markthub, Peter-Jan Gootzen, Elena Agostini, Sylvain Jeaugey, Aamir Shafi, Georgios Theodorakis, Manjunath Gorentla Venkata
CCGrid4
2025 SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
abstract
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training.
Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler
SC7
2023 DPFS: DPU-Powered File System Virtualization
abstract
As we move towards hyper-converged cloud solutions, the efficiency and overheads of distributed file systems at the cloud tenant side (i.e., client) become of paramount importance. Often, the clientside driver of a cloud file system is complex and CPU intensive, deeply coupled with the backend implementation, and requires optimizing multiple intrusive knobs. In this work, we propose to decouple the file system client from its backend implementation by virtualizing it with an off-the-shelf DPU using the Linux virtio-fs software stack. The decoupling allows us to offload the file system client execution to a DPU, which is managed and optimized by the cloud provider, while freeing the host CPU cycles. DPFS, our proposed framework, is 4.4× more host CPU efficient per I/O, delivers comparable performance to a tenant with zero-configuration and without modification of their host software stack, while allowing workload and hardware specific backend optimizations. The DPFS framework and its artifacts are publically available at https://github.com/IBM/DPFS.
Peter-Jan Gootzen, Jonas Pfefferle, Radu Stoica, Animesh Trivedi
SYSTOR1