Shulei Xu

dblp:264/1318 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0007-5041-9853ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Enhanced MPI Intra-Node Communication Framework: A Hybrid Approach with Cooperative DMA Channel-Based Data Transfer
abstract
In modern multi-/many-core HPC systems, the increasing demands for data transfers and memory I/O workloads have emerged as a predominant challenge for parallel programming models like MPI. Specifically, memory copy performance within the node has emerged as a significant bottleneck of the existing state-of-the-art MPI intra-node communication scheme. This paper introduces a hybrid approach to offload memory copy tasks partly from CPU, utilizing the Linux DMA Engine API with support for various CPU architectures, including I/OAT and PTDMA. We characterize the performance of DMA channel-based copy design for further optimizations. Subsequently, we integrate the DMA memory copy design into the existing MPI intra-node communication framework, thereby empowering the MPI intra-node point-to-point communication to simultaneously leverage DMA-based schemes and existing CPU-driven schemes in a cooperative manner. To demonstrate the effectiveness of our proposed design, we conducted systematic experiments on two emerging architectures. Our experimental results demonstrate up to 26 % lower communication latency in microbenchmarks, up to 23 % faster execution time in the 3D-stencil application.
Shulei Xu, Tu Tran, Dhabaleswar K. Panda 0001
HiPC1
2024 OHIO: Improving RDMA Network Scalability in MPI_Alltoall Through Optimized Hierarchical and Intra/Inter-Node Communication Overlap Design
abstract
The presence of exascale computers has pushed a new boundary in computing capability, which poses performance challenges in parallel programming models on how to exploit such systems efficiently. A dominant programming model for running parallel programs is the Message Passing Interface. Among primitives provided by MPI, Alltoall is a communication-intensive operation, which is utilized by many applications and is well-known for being difficult to optimize. Alltoall algorithms can be mainly classified into flat and hierarchical. The hierarchical designs avoid the slowdown of intra-node communication by inter-node communication by decoupling them. The hierarchical designs also reduce network congestion by reducing concurrently injected messages into the network. This work demonstrates an additional benefit of hierarchical designs to improve connection scalability in RDMA networks. This is attributed to the cache thrashing happening inside network adapters. All of these advantages of hierarchical schemes collectively contribute to the network scalability of Alltoall. This motivates us to propose a further optimized hierarchical design to enhance performance and network scalability. The design is network-agnostic and evaluated on clusters with InfiniBand and Omni-Path network adapters. The proposed design achieves average latency improvements of 61.13%, 56.40%, 37.49%, and 51.90% over Open MPI + UCX, HPC-X, Intel MPI, and MVAPICH2-X at micro-benchmark level with up to 7168 cores, respectively. In addition, the evaluation at application-level with Car-Parrinello Molecular Dynamics code shows 24.98 %, 40.44 % and 50.48 % improvement in the simulation time, compared to MVAPICH2-X, Open MPI + UCX, and Intel MPI, respectively.
Tu Tran, Goutham Kalikrishna Reddy Kuncham, Bharath Ramesh 0005, Shulei Xu, Hari Subramoni, Mustafa Abdul Jabbar, Dhabaleswar K. Panda 0001
HOTI4
2023 Optimized All-to-All Connection Establishment for High-Performance MPI Libraries Over InfiniBand
abstract
In modern multi-/many-core HPC systems, the increasing number of processor cores presents new challenges in managing parallel compute workloads across multiple nodes. One crucial aspect that significantly impacts the startup phase of parallel MPI jobs is the methodology used for connection establishment. In this paper, we investigate the limitations of existing all-to-all connection establishment designs in-depth, identify the primary sources of performance overhead, and propose an optimized all-to-all connection establishment design. This is done through an enforced ordering rank-by-rank scheme that significantly cuts down on the data exchange overhead via PMI. To address the increasing overheads associated with queue pair creation and synchronization, we explore multi-thread parallelism and incorporate CPU affinity awareness into our design. We implement our proposed design in the state-of-the-art MVAPICH2 MPI library and conduct extensive experiments on two emerging architectures. Through a comprehensive performance evaluation of these architectures, we demonstrate the efficacy of our optimized all-to-all connection establishment design. Our microbenchmark results reveal up to 20 times faster MPI_Init time, while evaluations with application kernels exhibit a 31% improvement in throughput.
Shulei Xu, Goutham Kalikrishna Reddy Kuncham, Mustafa Abdul Jabbar, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC1
2023 A Novel Framework for Efficient Offloading of Communication Operations to Bluefield SmartNICs
abstract
Smart Network Interface Cards (SmartNICs) such as NVIDIA’s BlueField Data Processing Units (DPUs) provide advanced networking capabilities and processor cores, enabling the offload of complex operations away from the host. In the context of MPI, prior work has explored the use of DPUs to offload non-blocking collective operations. The limitations of current state-of-the-art approaches are twofold: They only work for a pre-defined set of algorithms/communication patterns and have degraded communication latency due to staging data between the DPU and the host. In this paper, we propose a framework that supports the offload of any communication pattern to the DPU while achieving low communication latency with perfect overlap. To achieve this, we first study the limitations of higher-level programming models such as MPI in expressing the offload of complex communication patterns to the DPU. We present a new set of APIs to alleviate these shortcomings and support any generic communication pattern. Then, we analyze the bottlenecks involved in offloading communication operations to the DPU and propose efficient designs for a few candidate communication patterns. To the best of our knowledge, this is the first framework providing both efficient and generic communication offload to the DPU. Our proposed framework outperforms state-of-the-art staging-based offload solutions by 47% in Alltoall micro-benchmarks, and at the application level, we see improvements up to 60% in P3DFFT and 15% in HPL on 512 processes.
Kaushik Kandadi Suresh, Benjamin Michalowicz, Bharath Ramesh 0005, Nicholas Contini, Jinghan Yao, Shulei Xu, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001
IPDPS6
2021 Towards Architecture-aware Hierarchical Communication Trees on Modern HPC Systems
abstract
Modern HPC systems built with emerging multi-/many -core architectures have high core-counts and deep memory hierarchies. It is challenging to design communication libraries on these systems with the conventional wisdom of using OS processes as the basic building block to build communication algorithms. Instead, the next generation of communication libraries should treat hardware as the “first-class citizen” and utilize the underlying topology as the basic building block. Driven by this overarching principle, we present a framework for Optimized Shared Memory Processing (OSMP) and communication for these platforms. An abstract representation of the underlying hardware topology is maintained by OSMP in the form of a topology tree, which is later exploited by runtime libraries to execute communication operations in a topology-aware manner. This can be done by simply traversing the topology tree with an existing communication primitive as the base-case. OSMP does not mandate any changes to the original communication algorithm. We focus on collective operations such as barrier, reduction, and broadcast as candidate communication patterns. We demonstrate the efficacy of OSMP by decoupling the implementation of collective algorithms and system topology and evaluate it on four state-of-the-art multi-tmany-core architectures: Intel Cascade Lake, AMD Rome, ARM A64fx and IBM POWER9. Results show that even the basic algorithms can be made topology-aware by exploiting OSMP. This provides significant benefits over state-of-the-art algorithm implementations for intra-node communication. Using various micro-benchmarks and applications, we demonstrate that our proposed designs can achieve up to 7.8× improvements at the micro-benchmark level, and 15% for applications over state-of-the-art intra-node collective communication designs employed by production MPI libraries.
Bharath Ramesh 0005, Jahanzeb Maqbool Hashmi, Shulei Xu, Aamir Shafi, Seyedeh Mahdieh Ghazimirsaeed, Mohammadreza Bayatpour, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC3
2020 Design and Characterization of InfiniBand Hardware Tag Matching in MPI
abstract
Message Passing Interface (MPI) standard uses (source rank, tag, and communicator id) to properly place the incoming data into the application receive buffer. The act of searching through the receive queues and finding the appropriate match is called Tag Matching (TM). In the state-of-the-art MPI libraries, this operation is either being performed by the main thread or a separate communication progress thread. Either way leads to underutilization of the resources and major synchronization overheads leading to less optimal performance. Mellanox ConnectX-5 network architecture has introduced a feature to offload the Tag Matching and communication progress from host to InfiniBand network card. This paper proposes a Hardware Tag Matching aware MPI library and discusses various aspects and challenges of leveraging this feature in MPI library. Moreover, it characterizes hardware Tag Matching using different benchmarks and provides guidelines for the application developers to develop Hardware Tag Matching-aware applications to maximize their usage of this feature. Our proposed designs are able to improve the performance of non-blocking collectives up to 42% on 512 nodes and improve the performance of 3Dstencil application kernel on 7168 processes and Nekbone on 512 processes by a factor 40% and 3.5%, respectively.
Mohammadreza Bayatpour, Seyedeh Mahdieh Ghazimirsaeed, Shulei Xu, Hari Subramoni, Dhabaleswar K. Panda 0001
CCGRID3
2020 Machine-agnostic and Communication-aware Designs for MPI on Emerging Architectures
abstract
Modern multi-/many-cores offer higher core-density, hardware multi-threading, deeper memory hierarchies, and diverse architectural capabilities. While emerging cloud-based HPC systems are able to deliver near-native performance, they bring more diversity to the architectures. The Message Passing Interface (MPI) offers the flexibility to arbitrarily bind application processes to CPU cores, however the static nature of these binding policies typically does not take applications' communication patterns and underlying machine architecture into consideration. This lack of association between the dynamic nature of applications and architectural diversity offered by modern processors makes it difficult for the application developers and MPI designers to exploit modern multi-/many-core systems to their full potential. In this paper, we propose a set of low-level benchmarking based approaches and MPI-level designs to infer vendor-specific machine characteristics e.g., physical to virtual machine topologies, and dynamic communication patterns of the applications. By utilizing this information, we propose two novel algorithms to construct efficient MPI mappings for any given architecture and application communication pattern. The proposed designs are implemented in the MVAPICH2 MPI library and are evaluated on three different architectures using various micro-benchmarks and application kernels. We demonstrate up to 2X performance improvement for MPI collectives, and up to 3.5X and 26% improvement for NAS-CG and miniAMR application kernels, respectively.
Jahanzeb Maqbool Hashmi, Shulei Xu, Bharath Ramesh 0005, Mohammadreza Bayatpour, Hari Subramoni, Dhabaleswar K. Panda 0001
IPDPS2