EDBT 2026 Demo / reviewers in the wild / expert
Chen-Chun Chen
dblp:185/8604
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design and Implementation of Casting Compression for GPU-Aware MPI Collectives
Chen-Chun Chen, Nicholas Contini, Jacob Queiser, Hari Subramoni, Dhabaleswar K. Panda 0001 |
IPDPS | 1 |
| 2025 | Design and Optimization of GPU-Aware MPI Allreduce Using Direct Sendrecv CommunicationabstractModern GPU-accelerated high-performance computing (HPC) and deep learning (DL) applications rely heavily on collective communication, particularly the Allreduce operation. As systems scale to hundreds or thousands of GPUs, conventional algorithms such as Ring or vendor libraries like NCCL struggle to sustain performance for large scale due to communication bottlenecks and algorithmic dependencies. In this work, we propose a novel GPU-aware Allreduce design using a Direct Sendrecv algorithm with throttling to improve scalability and bandwidth utilization across various scales and interconnects. To further reduce overhead, we introduce computation-communication overlap and kernel fusion techniques. The design also extends to CPU-staging scenarios for small messages. Evaluations on large-scale GPU systems demonstrate that our designs outperform baseline NCCL implementations by up to 40% at medium message sizes. In application-level evaluations, the proposed design achieves up to 7% improvement in nanoGPT training and 27% improvement in the Amber HPC simulation. Chen-Chun Chen, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda 0001 |
ICPP | 1 |
| 2025 | Unified Designs of Multi-Rail-Aware MPI Allreduce and Alltoall Operations Across Diverse GPU and Interconnect SystemsabstractThe growing demand for computing power in highperformance computing is driving the adoption of diverse accelerators and interconnect networks in modern exascale clusters. Within a node, device interconnects like NVLink, Infinity Fabric, and$\mathbf{X}^{e}$Link, along with IPC techniques, provide high throughput in dense GPU environments. Recently, multi-rail interconnects, such as InfiniBand, Slingshot, and Omni-Path, have enabled highbandwidth communication between nodes. Additionally, highperformance computing applications impose significant demands on collective operations such as Allreduce and Alltoall. Therefore, designing efficient and scalable MPI runtimes for diverse system architectures at large scales is essential. In this paper, we propose unified designs to optimize MPI Allreduce and Alltoall operations using multi-rail-aware, two-level algorithms. The designs support a variety of GPU and interconnect combinations, including NVIDIA, AMD, and Intel GPUs, across InfiniBand, Slingshot, and Omni-Path networks, in the modern dense GPU systems. We optimized the Allreduce operation using a persistent device buffer for device-side reduction and employed an early-triggered, pipelined approach to overlap computation with communication. Additionally, we leveraged the device buffer with IPC techniques as a shared buffer to enhance the two-level Alltoall algorithm, designing both PUSH and PULL variants. We evaluate the advantages of our designs through benchmark and applicationlevel tests on the IsambardAI, Frontier, Cardinal, and Stampede3 systems. In benchmark evaluations, the proposed Allreduce design shows a$2.8 \mathrm{x}, 2.9 \mathrm{x}$, and 2.9 x performance improvement at 1 GB with 32 NVIDIA, 64 AMD, and 64 Intel GPUs, respectively. Additionally, the proposed Alltoall design demonstrates a 1.3 x and 1.05 x improvement at 4 MB with 32 NVIDIA and 64 AMD GPUs, respectively. In application-level evaluations, the proposed Allreduce design demonstrates a$2 x$performance improvement in Amber, while the Alltoall design shows a 1.4x performance gain in heFFTe, both tested on 32 H100 and GH200 GPUs with Infiniband and Slingshot-11 interconnects, respectively. Chen-Chun Chen, Jinghan Yao, Hari Subramoni, Dhabaleswar K. Panda 0001 |
IPDPS | 1 |
| 2024 | Design and Implementation of Kernel-based MPI Reduction Operations for Intel GPU sabstractThe demand for computing power in high- performance computing and deep learning applications is steadily increasing, leading to a noticeable inclination toward equipping modern exascale clusters with accelerators. In particular, dis-tributed Deep Learning training necessitates high-performance G PU -aware MPI operations, with reduction operations being widely employed. Unlike data movement-based MPI runtimes, reduction operations encompass both communication and computation, making them inherently more intricate to design and optimize for data transmission between GPU buffers. Acknowl-edging the success of NVIDIA and AMD GPUs in HPC, Intel has actively participated in the development of GPU products, while also fostering their associated ecosystems in recent years. However, existing MPI libraries supporting Intel GPUs rely on naive staging approaches, resulting in elevated latencies and subpar performance. In this paper, we propose a kernel- based reduction collective MPI library designed specifically for Intel G PU s. Our approach leverages IPC techniques to minimize data movement overhead during communication while harnessing highly efficient GPU kernels for the computational aspects of reduction operations. We assess the advantages of our designs through benchmark and application-level evaluations, conducted on ACES and Stampede3 systems. In benchmark- level evaluations, our Allreduce implementations demonstrate an 13.3x performance enhancement compared to Intel MPI at 1GB with 8 GPUs. Moreover, with 32 GPUs, we achieve a 42% performance enhancement. In application-level evaluations, our proposed designs exhibit up to a 22 % enhancement for the Deep Learning application TensorFlow with Horovod and a 28% improvement for PyTorch with Horovod on 32 GPUs compared to Intel MPI. Chen-Chun Chen, Goutham Kalikrishna Reddy Kuncham, Hari Subramoni, Dhabaleswar K. Panda 0001 |
HiPC | 1 |
| 2023 | Implementing and Optimizing a GPU-aware MPI Library for Intel GPUs: Early ExperiencesabstractAs the demand for computing power from High-Performance Computing (HPC) and Deep Learning (DL) applications increase, there is a growing trend of equipping modern exascale clusters with accelerators, such as NVIDIA and AMD GPUs. GPU-aware MPI libraries allow the applications to communicate between GPUs in a parallel environment with high productivity and performance. Although NVIDIA and AMD GPUs have dominated the accelerator market for top supercomputers over the past several years, Intel has recently developed and released its GPUs and associated software stack, and provided a unified programming model to program their GPUs, referred to as oneAPI. The emergence of Intel GPUs drives the need for initial MPI-level GPU-aware support that utilizes the underlying software stack specific to these GPUs and a thorough evaluation of communication. In this paper, we propose a GPU-aware MPI library for Intel GPUs using oneAPI and an SYCL backend. We delve into our experiments using Intel GPUs and the challenges to consider at the MPI layer when adding GPU-aware support using the software stack provided by Intel for their GPUs. We explore different memory allocation approaches and benchmark the memory copy performance with Intel GPUs. We propose implementations based on our experiments on Intel GPUs to support point-to-point GPU-aware MPI operations and show the high adaptability of our approach by extending the implementations to MPI collective operations, such as MPI_Bcast and MPI_Reduce. We evaluate the benefits of our implementations at the benchmark level by extending support for Intel GPU buffers over OSU Micro-Benchmarks. Our implementations provide up to 1.8x and 2.2x speedups on point-to-point latency using device buffers at small messages compared to Intel MPI and a naive benchmark, respectively; and have up to 1.3x and 1.5x speedups at large message sizes. At collective MPI operations, our implementations show 8x and 5x speedups for MPI_Allreduce and MPI_Allgather at large messages. At the application-level evaluation, our implementations provide up to 40% improvement for 3DStencil compared to Intel MPI. Chen-Chun Chen, Kawthar Shafie Khorassani, Goutham Kalikrishna Reddy Kuncham, Rahul Vaidya, Mustafa Abdul Jabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001 |
CCGrid | 1 |
| 2023 | Designing and Optimizing GPU-aware Nonblocking MPI Neighborhood Collective Communication for PETSc*abstractMPI Neighborhood collectives are used for non-traditional collective operations involving uneven distribution of communication amongst processes such as sparse communication patterns. They provide flexibility to define the communication pattern involved when a neighborhood relationship can be defined. PETSc, the Portable, Extensible Toolkit for Scientific Computation, used extensively with scientific applications to provide scalable solutions through routines modeled by partial differential equations, utilizes neighborhood communication patterns to define various structures and routines.We propose GPU-aware MPI Neighborhood collective operations with support for AMD and NVIDIA GPU backends and propose optimized designs to provide scalable performance for various communication routines. We evaluate our designs using PETSc structures for scattering from a parallel vector to a parallel vector, scattering from a sequential vector to a parallel vector, and scattering from a parallel vector to a sequential vector using a star forest graph representation implemented with nonblocking MPI neighborhood alltoallv collective operations. We evaluate our neighborhood designs on 64 NVIDIA GPUs on the Lassen system with Infiniband networking, demonstrating30.90% improvement against a GPU implementation utilizing CPU-staging techniques, and 8.25% improvement against GPU-aware point-to-point implementations of the communication pattern. We also evaluate on 64 AMD GPUs on the Spock system with slingshot networking and present 39.52% improvement against the CPU-staging implementation of a neighborhood GPU vector type in PETSc, and 33.25% improvement against GPU-aware point-to-point implementation of the routine. Kawthar Shafie Khorassani, Chen-Chun Chen, Hari Subramoni, Dhabaleswar K. Panda 0001 |
IPDPS | 2 |
| 2023 | High Performance MPI over the Slingshot Interconnect
Kawthar Shafie Khorassani, Chen-Chun Chen, Bharath Ramesh 0005, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | PHY: A performance-driven hybrid communication compression method for distributed training
Chen-Chun Chen, Yu-Min Chou, Jerry Chou 0001 |
J. Parallel Distributed Comput. | 1 |
| 2022 | Network Assisted Non-Contiguous Transfers for GPU-Aware MPI LibrariesabstractThe importance of GPUs in accelerating HPC applications is evident by the fact that a large number of super-computing clusters are GPU-enabled. Many of these HPC applications use MPI as their programming model. These MPI applications oftentimes exchange data that is non-contiguous in GPU memory. MPI provides Derived Datatypes(DDTs) to represent such data. In the past, researchers have proposed solutions to optimize these MPI DDT based inter-node GPU exchanges. All of these solutions are aimed at optimizing the overheads associated with pack-unpack kernels that facilitate the non-contiguous exchanges. Modern HCAs are capable of gathering/scattering data from/to non-contiguous GPU memory regions. In this work, we analyze the challenges in using HCA's scatter/gather mechanism for GPU-based HPC workloads. We propose a low-overhead HCA-assisted scheme to improve the performance of GPU-based non-contiguous exchanges. We show that the proposed scheme provides up to 2X benefits compared to existing pack-based schemes at the benchmark level. Fur-thermore, on the layouts used by MILC, NASMG, Specfem3D applications, we show that the proposed scheme outperforms the state-of-the MPI libraries such as MVAPICH2-GDR and OpenMPI+UCX. Kaushik Kandadi Suresh, Kawthar Shafie Khorassani, Chen-Chun Chen, Bharath Ramesh 0005, Mustafa Abdul Jabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001 |
HOTI | 3 |
| 2022 | ALBERT: An automatic learning based execution and resource management system for optimizing Hadoop workload in clouds
Chen-Chun Chen, Kai-Siang Wang, Yu-Tung Hsiao, Jerry Chou 0001 |
J. Parallel Distributed Comput. | 1 |
| 2021 | Layout-aware Hardware-assisted Designs for Derived Data Types in MPIabstractModern MPI-based scientific applications frequently use derived datatypes (DDT) for inter-process communication. Designing scalable solutions capable of dynamically adapting themselves to the complex communication requirements posed by DDT-based applications bring forth several new challenges. In this work, we address these challenges and propose solutions to efficiently improve the performance of hardware-assisted datatype transfers. Further, we design a layout-aware DDT scheme that dynamically adapts the datatype processing to the communication requirements of the datatype layouts used by the application. The proposed layout-aware adaptive scheme is able to dynamically switch between different host-based and the proposed hardware-assisted schemes to deliver the best performance and scalability while hiding the communication overheads. The experimental evaluations on multiple HPC systems including Frontera at TACC and Expanse at SDSC show that our proposed designs achieve up to 22 % improvement in performance over state-of-the-art MPI libraries at the micro-benchmark level. We also evaluate our designs with various scientific application kernels such as MILC, WRF, and applications such as miniGhost and demonstrate up to 9 % improvement in performance at 128 nodes for the miniGhost application. Kaushik Kandadi Suresh, Bharath Ramesh 0005, Chen-Chun Chen, Seyedeh Mahdieh Ghazimirsaeed, Mohammadreza Bayatpour, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001 |
HiPC | 3 |