Tu Tran

dblp:42/5678 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-0040-8404ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enhanced MPI Intra-Node Communication Framework: A Hybrid Approach with Cooperative DMA Channel-Based Data Transfer
abstract
In modern multi-/many-core HPC systems, the increasing demands for data transfers and memory I/O workloads have emerged as a predominant challenge for parallel programming models like MPI. Specifically, memory copy performance within the node has emerged as a significant bottleneck of the existing state-of-the-art MPI intra-node communication scheme. This paper introduces a hybrid approach to offload memory copy tasks partly from CPU, utilizing the Linux DMA Engine API with support for various CPU architectures, including I/OAT and PTDMA. We characterize the performance of DMA channel-based copy design for further optimizations. Subsequently, we integrate the DMA memory copy design into the existing MPI intra-node communication framework, thereby empowering the MPI intra-node point-to-point communication to simultaneously leverage DMA-based schemes and existing CPU-driven schemes in a cooperative manner. To demonstrate the effectiveness of our proposed design, we conducted systematic experiments on two emerging architectures. Our experimental results demonstrate up to 26 % lower communication latency in microbenchmarks, up to 23 % faster execution time in the 3D-stencil application.
Shulei Xu, Tu Tran, Dhabaleswar K. Panda 0001
HiPC2
2024 MPI Allgather Utilizing CXL Shared Memory Pool in Multi-Node Computing Systems
abstract
In Artificial Intelligence (AI) and high-performance computing (HPC), growing data and model sizes require distributed processing across multiple nodes due to single-node limitations, increasing inter-node communication. To address these challenges, we propose a novel MPI allgather method leveraging CXL technology, which supports composable architectures and dynamic resource allocation in data centers and HPC systems. Notably, CXL 3.1 facilitates cache coherence among nodes. The proposed allgather method uses the CXL shared memory pool as a communication buffer, outperforming existing algorithms for two reasons: First, CXL provides lower latency than Ethernet and IB, and second, by using the CXL shared memory pool as a shared communication buffer across multiple nodes, it significantly reduces the number of communications. To the best of our knowledge, this work is the first to explore combining MPI collective communication with CXL technology to optimize MPI allgather. Our proposed allgather method significantly reduces communication latency compared to traditional allgather methods by up to 42.14x, with a minimum improvement of 2.91x, as measured using the OSU Micro-Benchmark (OMB), a standard MPI benchmarking suite.
Hooyoung Ahn, Seonyoung Kim, Yoo-Mi Park, Woojong Han, Shin-Young Ahn, Tu Tran, Bharath Ramesh 0005, Hari Subramoni, Dhabaleswar K. Panda 0001
IEEE Big Data6
2024 OHIO: Improving RDMA Network Scalability in MPI_Alltoall Through Optimized Hierarchical and Intra/Inter-Node Communication Overlap Design
abstract
The presence of exascale computers has pushed a new boundary in computing capability, which poses performance challenges in parallel programming models on how to exploit such systems efficiently. A dominant programming model for running parallel programs is the Message Passing Interface. Among primitives provided by MPI, Alltoall is a communication-intensive operation, which is utilized by many applications and is well-known for being difficult to optimize. Alltoall algorithms can be mainly classified into flat and hierarchical. The hierarchical designs avoid the slowdown of intra-node communication by inter-node communication by decoupling them. The hierarchical designs also reduce network congestion by reducing concurrently injected messages into the network. This work demonstrates an additional benefit of hierarchical designs to improve connection scalability in RDMA networks. This is attributed to the cache thrashing happening inside network adapters. All of these advantages of hierarchical schemes collectively contribute to the network scalability of Alltoall. This motivates us to propose a further optimized hierarchical design to enhance performance and network scalability. The design is network-agnostic and evaluated on clusters with InfiniBand and Omni-Path network adapters. The proposed design achieves average latency improvements of 61.13%, 56.40%, 37.49%, and 51.90% over Open MPI + UCX, HPC-X, Intel MPI, and MVAPICH2-X at micro-benchmark level with up to 7168 cores, respectively. In addition, the evaluation at application-level with Car-Parrinello Molecular Dynamics code shows 24.98 %, 40.44 % and 50.48 % improvement in the simulation time, compared to MVAPICH2-X, Open MPI + UCX, and Intel MPI, respectively.
Tu Tran, Goutham Kalikrishna Reddy Kuncham, Bharath Ramesh 0005, Shulei Xu, Hari Subramoni, Mustafa Abdul Jabbar, Dhabaleswar K. Panda 0001
HOTI1
2024 Accelerating communication with multi-HCA aware collectives in MPI
abstract
Summary To accelerate the communication between nodes, supercomputers are now equipped with multiple network adapters per node, also referred to as HCAs (Host Channel Adapters), resulting in a “multi‐rail”/“multi‐HCA” network. For example, the ThetaGPU system at Argonne National Laboratory (ANL) has eight adapters per node; with this many networking resources available, utilizing all of them becomes non‐trivial. The Message Passing Interface (MPI) is a dominant model for high‐performance computing clusters. Not all MPI collectives utilize all resources, and this becomes more apparent with advances in bandwidth and adapter count in a given cluster. In this work, we provide a thorough performance analysis of existing multirail solutions and their implications on collectives and present the necessity for further enhancement. Specifically, we propose novel designs for hierarchical, multi‐HCA‐aware Allgather. The proposed designs fully utilize all the available network adapters within a node and provide high overlap between inter‐node and intra‐node communication. At the micro‐benchmark level, we see large inter‐node improvements up to 62% and 61% better than HPC‐X and MVAPICH2‐X for 1024 processes. Because Allgather is used in Ring‐Allreduce, our designs also improve its performance by 56% and 44% compared to HPC‐X and MVAPICH2‐X, respectively. At the application level, our enhanced Allgather shows and improvement in a matrix‐vector multiplication kernel when compared to HPC‐X and MVAPICH2‐X, and Allreduce performs up to 7.83% better in deep learning training against MVAPICH2‐X.
Tu Tran, Bharath Ramesh 0005, Benjamin Michalowicz, Mustafa Abdul Jabbar, Hari Subramoni, Aamir Shafi, Dhabaleswar K. Panda 0001
Concurr. Comput. Pract. Exp.1
2023 Enabling Reconfigurable HPC through MPI-based Inter-FPGA Communication
abstract
Modern HPC faces new challenges with the slowing of Moore's Law and the end of Dennard Scaling. Traditional computing architectures can no longer be expected to drive today's HPC loads, as shown by the adoption of heterogeneous system design leveraging accelerators such as GPUs and TPUs. Recently, FPGAs have become viable candidates as HPC accelerators. These devices can accelerate workloads by replicating implemented compute units to enable task parallelism, overlapping computation between and within kernels to enable pipeline parallelism, and increasing data locality by sending data directly between compute units. While many solutions for inter-FPGA communication have been presented, these proposed designs generally rely on inter-FPGA networks, unique system setups, and/or the consumption of soft logic resources on the chip. In this paper, we propose an FPGA-aware MPI runtime that avoids such shortcomings. Our MPI implementation does not use any special system setup other than plugging FPGA accelerators into PCIe slots. All communication is orchestrated by the host, utilizing the PCIe interconnect and inter-host network to implement message passing. We propose advanced designs that address data movement challenges and reduce the need for explicit data movement between the device and host (staging) in FPGA applications. We achieve up to 50% reduction in latency for point-to-point transfers compared to application-level staging.
Nicholas Contini, Bharath Ramesh 0005, Kaushik Kandadi Suresh, Tu Tran, Benjamin Michalowicz, Mustafa Abdul Jabbar, Hari Subramoni, Dhabaleswar K. Panda 0001
ICS4
2021 Large-Message Nonblocking MPI_Iallgather and MPI Ibcast Offload via BlueField-2 DPU
abstract
Since the introduction of nonblocking collectives in the MPI-3 standard, communication has been progressed by several mechanisms. One such mechanism includes modifying the application code to periodically call MPI_ Test to enter the MPI library. Another launches an extra thread per core to progress communication asynchronously. Communication progression can also be offloaded to the Host Channel Adapter (HCA) using the latest hardware. In this paper, we explore this last option by using the Data Processing Unit (DPU) shipped with the BlueField-2 SmartNIC adapter to offload progression of non-blocking MPI_Ibcast and MPI_Iallgather collectives. For both collectives, we present several designs which take advantage of the DPU. We demonstrate the efficacy of our proposed designs through microbenchmark evaluations. At the microbenchmark level, total execution time of the osu_ibcast microbenchmark can be reduced by up to 54% using our DPU-based Ibcast designs. Total execution time of the osu_iallgather microbenchmark can be reduced by up to 43 %. To the best of our knowledge, this is the first work to optimize nonblocking broadcast and allgather collectives on emerging BlueField DPUs.
Nick Sarkauskas, Mohammadreza Bayatpour, Tu Tran, Bharath Ramesh 0005, Hari Subramoni, Dhabaleswar K. Panda 0001
HiPC3
2005 Is your web page accessible?: a comparative study of methods for assessing web page accessibility for the blind
abstract
Web access for users with disabilities is an important goal and challenging problem for web content developers and designers. This paper presents a comparison of different methods for finding accessibility problems affecting users who are blind. Our comparison focuses on techniques that might be of use to Web developers without accessibility experience, a large and important group that represents a major source of inaccessible pages. We compare a laboratory study with blind users to an automated tool, expert review by web designers with and without a screen reader, and remote testing by blind users. Multiple developers, using a screen reader, were most consistently successful at finding most classes of problems, and tended to find about 50% of known problems. Surprisingly, a remote study with blind users was one of the least effective methods. All of the techniques, however, had different, complementary strengths and weaknesses.
Jennifer Mankoff, Holly Fait, Tu Tran
CHI3