Goutham Kalikrishna Reddy Kuncham

dblp:308/0119 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-2112-4769ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 One Memory-Many Paths: Early Experiences with Allocation and Data Copy Strategies on MI300A
Goutham Kalikrishna Reddy Kuncham, Siyuan Zhang 0001, Dhabaleswar K. Panda 0001
IPDPS1
2025 Towards Dynamic Message Passing Protocols for Stencil-Based Communication Patterns
abstract
Halo-exchange communication patterns occur in many stencil-based HPC applications such as MiniAMR, MiniGhost, and MILC. In this pattern, each process performs a mix of inter-node and intra-node transfers. Depending on the input and processor grid size, the amount of time spent in inter-node or intra-node could dominate the total communication time. Therefore, in this work, we propose a dynamic protocol for intranode and inter-node transfers that optimizes the communication time. With the proposed designs, we show up to 48% improvements over state-of-the-art libraries in 3D stencil communication benchmarks and 28% in the MiniAMR application at a scale of 2304 processes.
Kaushik Kandadi Suresh, Bharath Ramesh 0005, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda 0001
CLUSTER3
2024 Design and Implementation of Kernel-based MPI Reduction Operations for Intel GPU s
abstract
The demand for computing power in high- performance computing and deep learning applications is steadily increasing, leading to a noticeable inclination toward equipping modern exascale clusters with accelerators. In particular, dis-tributed Deep Learning training necessitates high-performance G PU -aware MPI operations, with reduction operations being widely employed. Unlike data movement-based MPI runtimes, reduction operations encompass both communication and computation, making them inherently more intricate to design and optimize for data transmission between GPU buffers. Acknowl-edging the success of NVIDIA and AMD GPUs in HPC, Intel has actively participated in the development of GPU products, while also fostering their associated ecosystems in recent years. However, existing MPI libraries supporting Intel GPUs rely on naive staging approaches, resulting in elevated latencies and subpar performance. In this paper, we propose a kernel- based reduction collective MPI library designed specifically for Intel G PU s. Our approach leverages IPC techniques to minimize data movement overhead during communication while harnessing highly efficient GPU kernels for the computational aspects of reduction operations. We assess the advantages of our designs through benchmark and application-level evaluations, conducted on ACES and Stampede3 systems. In benchmark- level evaluations, our Allreduce implementations demonstrate an 13.3x performance enhancement compared to Intel MPI at 1GB with 8 GPUs. Moreover, with 32 GPUs, we achieve a 42% performance enhancement. In application-level evaluations, our proposed designs exhibit up to a 22 % enhancement for the Deep Learning application TensorFlow with Horovod and a 28% improvement for PyTorch with Horovod on 32 GPUs compared to Intel MPI.
Chen-Chun Chen, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC2
2024 OHIO: Improving RDMA Network Scalability in MPI_Alltoall Through Optimized Hierarchical and Intra/Inter-Node Communication Overlap Design
abstract
The presence of exascale computers has pushed a new boundary in computing capability, which poses performance challenges in parallel programming models on how to exploit such systems efficiently. A dominant programming model for running parallel programs is the Message Passing Interface. Among primitives provided by MPI, Alltoall is a communication-intensive operation, which is utilized by many applications and is well-known for being difficult to optimize. Alltoall algorithms can be mainly classified into flat and hierarchical. The hierarchical designs avoid the slowdown of intra-node communication by inter-node communication by decoupling them. The hierarchical designs also reduce network congestion by reducing concurrently injected messages into the network. This work demonstrates an additional benefit of hierarchical designs to improve connection scalability in RDMA networks. This is attributed to the cache thrashing happening inside network adapters. All of these advantages of hierarchical schemes collectively contribute to the network scalability of Alltoall. This motivates us to propose a further optimized hierarchical design to enhance performance and network scalability. The design is network-agnostic and evaluated on clusters with InfiniBand and Omni-Path network adapters. The proposed design achieves average latency improvements of 61.13%, 56.40%, 37.49%, and 51.90% over Open MPI + UCX, HPC-X, Intel MPI, and MVAPICH2-X at micro-benchmark level with up to 7168 cores, respectively. In addition, the evaluation at application-level with Car-Parrinello Molecular Dynamics code shows 24.98 %, 40.44 % and 50.48 % improvement in the simulation time, compared to MVAPICH2-X, Open MPI + UCX, and Intel MPI, respectively.
Tu Tran, Goutham Kalikrishna Reddy Kuncham, Bharath Ramesh 0005, Shulei Xu, Hari Subramoni, Mustafa Abdul Jabbar, Dhabaleswar K. Panda 0001
HOTI2
2023 Implementing and Optimizing a GPU-aware MPI Library for Intel GPUs: Early Experiences
abstract
As the demand for computing power from High-Performance Computing (HPC) and Deep Learning (DL) applications increase, there is a growing trend of equipping modern exascale clusters with accelerators, such as NVIDIA and AMD GPUs. GPU-aware MPI libraries allow the applications to communicate between GPUs in a parallel environment with high productivity and performance. Although NVIDIA and AMD GPUs have dominated the accelerator market for top supercomputers over the past several years, Intel has recently developed and released its GPUs and associated software stack, and provided a unified programming model to program their GPUs, referred to as oneAPI. The emergence of Intel GPUs drives the need for initial MPI-level GPU-aware support that utilizes the underlying software stack specific to these GPUs and a thorough evaluation of communication. In this paper, we propose a GPU-aware MPI library for Intel GPUs using oneAPI and an SYCL backend. We delve into our experiments using Intel GPUs and the challenges to consider at the MPI layer when adding GPU-aware support using the software stack provided by Intel for their GPUs. We explore different memory allocation approaches and benchmark the memory copy performance with Intel GPUs. We propose implementations based on our experiments on Intel GPUs to support point-to-point GPU-aware MPI operations and show the high adaptability of our approach by extending the implementations to MPI collective operations, such as MPI_Bcast and MPI_Reduce. We evaluate the benefits of our implementations at the benchmark level by extending support for Intel GPU buffers over OSU Micro-Benchmarks. Our implementations provide up to 1.8x and 2.2x speedups on point-to-point latency using device buffers at small messages compared to Intel MPI and a naive benchmark, respectively; and have up to 1.3x and 1.5x speedups at large message sizes. At collective MPI operations, our implementations show 8x and 5x speedups for MPI_Allreduce and MPI_Allgather at large messages. At the application-level evaluation, our implementations provide up to 40% improvement for 3DStencil compared to Intel MPI.
Chen-Chun Chen, Kawthar Shafie Khorassani, Goutham Kalikrishna Reddy Kuncham, Rahul Vaidya, Mustafa Abdul Jabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
CCGrid3
2023 Optimized All-to-All Connection Establishment for High-Performance MPI Libraries Over InfiniBand
abstract
In modern multi-/many-core HPC systems, the increasing number of processor cores presents new challenges in managing parallel compute workloads across multiple nodes. One crucial aspect that significantly impacts the startup phase of parallel MPI jobs is the methodology used for connection establishment. In this paper, we investigate the limitations of existing all-to-all connection establishment designs in-depth, identify the primary sources of performance overhead, and propose an optimized all-to-all connection establishment design. This is done through an enforced ordering rank-by-rank scheme that significantly cuts down on the data exchange overhead via PMI. To address the increasing overheads associated with queue pair creation and synchronization, we explore multi-thread parallelism and incorporate CPU affinity awareness into our design. We implement our proposed design in the state-of-the-art MVAPICH2 MPI library and conduct extensive experiments on two emerging architectures. Through a comprehensive performance evaluation of these architectures, we demonstrate the efficacy of our optimized all-to-all connection establishment design. Our microbenchmark results reveal up to 20 times faster MPI_Init time, while evaluations with application kernels exhibit a 31% improvement in throughput.
Shulei Xu, Goutham Kalikrishna Reddy Kuncham, Mustafa Abdul Jabbar, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC2
2023 Designing In-network Computing Aware Reduction Collectives in MPI
abstract
The Message-Passing Interface (MPI) provides convenient abstractions such as MPI_Allreduce for inter-process collective reduction operations. With the advent of deep learning and large-scale HPC systems, it is ever so important to optimize the latency of the MPI_Allreduce operation for large messages. Due to the amount of compute and communication involved in MPI_Allreduce, it is beneficial to offload collective computation/communication to the network to allow the CPU to work on other important operations and provide maximal overlap/scalability. NVIDIA’s HDR InfiniBand switches provide in-network computing features using the Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) for this purpose with two protocols targeted at different message ranges: 1) Local Latency Tree (LLT) for small messages, and 2) Streaming aggregation tree (SAT) for large messages. In this paper, we first analyze the overheads involved in using SHARP-based reductions with SAT in an MPI library using micro-benchmarks. Next, we propose designs for large message MPI_Allreduce by fully utilizing the capabilities provided by the SHARP runtime while overcoming various bottlenecks. The efficacy of our proposed designs is demonstrated using micro-benchmark results. We observe up to 89% improvements over MVAPICH2-X and HPC-X for large message reductions.
Bharath Ramesh 0005, Goutham Kalikrishna Reddy Kuncham, Kaushik Kandadi Suresh, Rahul Vaidya, Nawras Alnaasan, Mustafa Abdul Jabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
HOTI2