Thomas B. Rolinger

dblp:194/3555 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0001-8383-4737ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 JITSPMM: Just-in-Time Instruction Generation for Accelerated Sparse Matrix-Matrix Multiplication
abstract
Achieving high performance for Sparse Matrix-Matrix Multiplication (SpMM) has received increasing research attention, especially on multi-core CPUs, due to the large input data size in applications such as graph neural networks (GNNs). Most existing solutions for SpMM computation follow the ahead-of-time (AOT) compilation approach, which compiles a program entirely before it is executed. AOT compilation for SpMM faces three key limitations: unnecessary memory access, additional branch overhead, and redundant instructions. These limitations stem from the fact that crucial information pertaining to SpMM is not known until runtime. In this paper, we propose JITSpMM, a just-in-time (JIT) assembly code generation framework to accelerated SpMM computation on multi-core CPUs with SIMD extensions. First, JITSpMM integrates the JIT assembly code generation technique into three widely-used workload division methods for SpMM to achieve balanced workload distribution among CPU threads. Next, with the availability of runtime information, JITSpMM employs a novel technique, coarse-grain column merging, to maximize instruction-level parallelism by unrolling the performance-critical loop. Furthermore, JITSpMM intelligently allocates registers to cache frequently accessed data to minimizing memory accesses, and employs selected SIMD instructions to enhance arithmetic throughput. We conduct a performance evaluation of JITSpMM and compare it two AOT baselines. The first involves existing SpMM implementations compiled using the Intel icc compiler with auto-vectorization. The second utilizes the highly-optimized SpMM routine provided by Intel MKL. Our results show that JITSpMM provides an average improvement of 3.8× and 1.4×, respectively.
Thomas B. Rolinger, H. Howie Huang
CGO2
2024 Adaptive Prefetching for Fine-grain Communication in PGAS Programs
abstract
Applications that require distributed-memory systems and exhibit irregular memory access patterns present both productivity and performance challenges. This is largely due to the fine-grain communication that arises from irregular memory accesses to distributed data. The Partitioned Global Address Space (PGAS) model provides a globally shared address space and one-sided communication, making it well suited for implementing irregular codes. However, while the PGAS model provides productivity advantages, irregular applications have difficulty achieving high performance due to the cost of fine-grain remote accesses. One way to improve the performance of such PGAS programs is to prefetch remote data before it is needed, thereby hiding communication latency. The challenges of applying prefetching are computing the prefetch distance (i.e., how far ahead to issue prefetches), determining when prefetching will be profitable, and modifying the program to perform prefetching. In this work, we present an adaptive prefetching optimization that can be applied to PGAS programs with irregular memory access patterns. We target the Chapel parallel programming language, which implements a PGAS model. Our optimization leverages runtime information to adjust the prefetch distance and to pause/resume prefetching when that is likely to provide performance gains. Furthermore, we develop compiler support to automatically apply the optimization without requiring user intervention. We evaluate the optimization across five different distributed-memory systems and four workloads. We observe runtime speed-ups of 0.5 – 377x for the Chapel workloads when compared to not prefetching, and 3 – 215x when compared to implementations written in UPC, OpenSHMEM, and one-sided MPI.
Thomas B. Rolinger, Alan Sussman
IPDPS1
2021 Optimizing Memory-Compute Colocation for Irregular Applications on a Migratory Thread Architecture
abstract
The movement of data between memory and processors has become a performance bottleneck for many applications. This is made worse for applications with sparse and irregular memory accesses, as they exhibit weak locality and make poor utilization of cache. As a result, colocating memory and compute is crucial for achieving high performance on irregular applications. There are two paradigms for memory-compute colocation. The first is the conventional approach of moving the data to the compute. The second paradigm is to move the compute to the data, which is less conventional and not as well understood. An example are migratory threads, which physically relocate upon remote accesses to the compute resource that hosts the data. In this paper, we explore the paradigm of moving compute to the data by optimizing memory-compute colocation for irregular applications on a migratory thread architecture. Our optimization method includes both initial data placement as well as data replication. We evaluate our optimization on sparse matrix-vector multiply (SpMV) and sparse matrix-matrix multiply (SpGEMM). Our results show that we can achieve speed-ups as high as 4.2x on SpMV and 6x on SpGEMM when compared to the default data layout. We also highlight that our optimization to improve memory-compute colocation can be applicable to both migratory threads and more conventional systems. To this end, we evaluate our optimization approach on a conventional compute cluster using the Chapel programming language. We demonstrate speed-ups as high as 18x for SpMV.
Thomas B. Rolinger, Christopher D. Krieger, Alan Sussman
IPDPS1
2019 Performance considerations for scalable parallel tensor decomposition
Thomas B. Rolinger, Tyler A. Simon, Christopher D. Krieger
J. Parallel Distributed Comput.1
2018 An Empirical Evaluation of Allgatherv on Multi-GPU Systems
abstract
Applications for deep learning and big data analytics have compute and memory requirements that exceed the limits of a single GPU. However, effectively scaling out an application to multiple GPUs is challenging due to the complexities of communication between the GPUs, particularly for collective communication with irregular message sizes. In this work, we provide a performance evaluation of the Allgatherv routine on multi-GPU systems, focusing on GPU network topology and the communication library used. We present results from the OSU-micro benchmark as well as conduct a case study for sparse tensor factorization, one application that uses Allgatherv with highly irregular message sizes. We extend our existing tensor factorization tool to run on systems with different node counts and varying number of GPUs per node. We then evaluate the communication performance of our tool when using traditional MPI, CUDA-aware MVAPICH and NCCL across a suite of real-world data sets on three different systems: a 16-node cluster with one GPU per node, NVIDIA's DGX-1 with 8 GPUs and Cray's CS-Storm with 16 GPUs. Our results show that irregularity in the tensor data sets produce trends that contradict those in the OSU micro-benchmark, as well as trends that are absent from the benchmark.
Thomas B. Rolinger, Tyler A. Simon, Christopher D. Krieger
CCGrid1