Gil Bloch

dblp:99/2609 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0004-6224-9802ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2025 SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
abstract
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-haul link characteristics, such as drop rate, distance and bandwidth, the widely used Selective Repeat algorithm can be inefficient, warranting alternatives like Erasure Coding. To enable such alternatives on existing hardware, we propose SDR-RDMA, a software-defined reliability stack for RDMA. Its core is a lightweight SDR SDK that extends standard point-to-point RDMA semantics — fundamental to AI networking stacks — with a receive buffer bitmap. SDR bitmap enables partial message completion to let applications implement custom reliability schemes tailored to specific deployments, while preserving zero-copy RDMA benefits. By offloading the SDR backend to NVIDIA’s Data Path Accelerator (DPA), we achieve line-rate performance, enabling efficient inter-datacenter communication and advancing reliability innovation for inter-datacenter training.
Mikhail Khalilov, Marcin Chrapek, Tiancheng Chen, Kenji Nakano, Nicola Mazzoletti, Peter-Jan Gootzen, Salvatore Di Girolamo, Rami Nudelman, Gil Bloch, Abdul Kabbani, Sreevatsa Anantharamu, Konstantin Taranov, Zhuolong Yu, Scott Moe, Mahmoud Elhaddad, Torsten Hoefler
SC10
2024 Unified Collective Communication (UCC): An Unified Library for CPU, GPU, and DPU Collectives
abstract
Unified Collective Communication (UCC) is an API and library implementation of collective communication operations. The goal of UCC is to provide a unified API and library serving the collective communication needs of various workloads running on a wide variety of system architectures. Particularly, we aim to unify the collective communication interfaces and semantics and provide a common implementation framework for (i) parallel programming models including HPC (message passing and Partioned Global Address Space (PGAS)), Deep Learning (DL) and 1/0, (ii) collectives moving data in CPU main memory and device memory (GPU, and Data Processing Unit (DPU)), and (iii) collective operations using software point-to-point transports, and hardware transports. In this paper, we present an overview of UCC's design, interfaces, semantics, and an implementation. We demonstrate UCC's capabilities through evaluations with microbenchmarks representing diverse workloads and applications across various programming models including MPI, PGAS (OpenSHMEM), and PyTorch, on multiple hardware architectures (CPU, GPU and DPU). Our results show that with UCC, microbenchmarks perform up to 3.4X and 5.4X better on latency and bandwidth, respectively compared to the Tuned collective component in Open MPI. In addition, we show up to 12% performance improvement in applications without any modifications to the application itself and an up to 4X improvement when leveraging the flexibility UCC provides. These improvements highlight UCC's success in unifying collective communication needs. UCC has been adopted by multiple institutions, and it is currently integrated with Open MPI for Message Passing Interface (MPI) and OpenSHMEM programming models, and PyTorch for DL workloads.
Manjunath Gorentla Venkata, Valentine Petrov, Sergey Lebedev, Devendar Bureddy, Ferrol Aderholdt, Joshua Ladd, Gil Bloch, Mike Dubman, Gilad Shainer
HOTI7
2024 Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
abstract
In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstanding operations such as Allgather and Reduce-Scatter can compete for the injection bandwidth and create pipeline bubbles. To address this problem, we propose a novel bandwidth-optimal Allgather collective algorithm that leverages hardware multicast. We use multicast to build a constant-time reliable Broadcast protocol, a building block for constructing an optimal Allgather schedule. Our Allgather algorithm achieves $2 \times$ traffic reduction on a 188 -node testbed. To free the host side from running the protocol, we employ SmartNIC offloading. We extract the parallelism in our Allgather algorithm and map it to a SmartNIC specialized for hiding the cost of data movement. We show that our SmartNIC-offloaded collective progress engine can scale to the next generation of 1.6 Tbit/s links.
Mikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman, Gil Bloch, Torsten Hoefler
SC5
2019 Accelerating OpenSHMEM Collectives Using In-Network Computing Approach
abstract
OpenSHMEM is one of the key programming models for High Performance Computing (HPC) applications with irregular communication patterns. Particularly, it is useful for problems that cannot be decomposed easily such as graph partitioning. The programming model supports Remote Memory Access (RMA), atomics, and collective operations. In this paper, we explore and evaluate the In-network Computing approach for accelerating the OpenSHMEM collective operations, particularly barrier, broadcast, and reduction operations. To achieve acceleration, In-network Computing leverages hardware engines on the networking elements and effective software that can efficiently use these capabilities. We explore the value of this approach for collective operations on the InfiniBand Host Channel Adapters (HCAs) and switches. Particularly, we focus on the recently introduced collective offload feature provided by the Mellanox Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) TM capability, which accelerates the barriers and reduction operations; the multicast capability accelerates the broadcast collective operation. To leverage the hardware capabilities, we complement it with an effective software stack that includes Hierarchical Collectives (HCOLL) library, and SHARP layer. Our evaluation on Oak Ridge National Laboratory (ORNL)'s Summit system, which is the fastest supercomputer on the June 2019 Top 500 list, show that the hardware and software acceleration in the In-network Computing approach is key for achieving the performance and scalability required for collectives and applications. For a 5120 process OpenSHMEM job, our results show that the barrier operation is 710% faster, broadcast is 370% faster, and reduction operation is 10 times faster when compared with the implementation of collective operations with no acceleration. Further, experiments with a 2D-Heat kernel show that the In-network Computing approach is very effective for realworld applications.
Manjunath Gorentla Venkata, Gil Bloch, Gilad Shainer, Richard L. Graham
SBAC-PAD2
2010 ConnectX-2 InfiniBand Management Queues: First Investigation of the New Support for Network Offloaded Collective Operations
abstract
This paper introduces the newly developed Infini-Band (IB) Management Queue capability, used by the Host Channel Adapter (HCA) to manage network task data flow dependancies, and progress the communications associated with such flows. These tasks include sends, receives, and the newly supported wait task, and are scheduled by the HCA based on a data dependency description provided by the user. This functionality is supported by the ConnectX-2 HCA, and provides the means for delegating collective communication management and progress to the HCA, also known as collective communication offload. This provides a means for overlapping collective communications managed by the HCA and computation on the Central Processing Unit (CPU), thus making it possible to reduce the impact of system noise on parallel applications using collective operations. This paper further describes how this new capability can be used to implement scalable Message Passing Interface (MPI) collective operations, describing the high level details of how this new capability is used to implement the MPI_Barrier collective operation, focusing on the latency sensitive performance aspects of this new capability. This paper concludes with small scale benchmark experiments comparing implementations of the barrier collective operation, using the new network offload capabilities, with established point-to-point based implementations of these same algorithms, which manage the data flow using the central processing unit. These early results demonstrate the promise this new capability provides to improve the scalability of high-performance applications using collective communications. The latency of the HCA based implementation of the barrier is similar to that of the best performing point-to-point based implementation managed by the central processing unit, starting to outperform these as the number of processes involved in the collective operation increases.
Richard L. Graham, Stephen W. Poole, Pavel Shamis, Gil Bloch, Noam Bloch, Hillel Chapman, Michael Kagan, Ariel Shahar, Ishai Rabinovitz, Gilad Shainer
CCGRID4