Rahulkumar Gayatri

dblp:117/6956 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Leveraging LLVM OpenMP GPU Offload Optimizations for Kokkos Applications
abstract
OpenMP provides a cross-vendor API for GPU offload that can serve as an implementation layer under performance portability frameworks like the Kokkos C++ library. However, recent work identified some impediments to performance with this approach arising from limitations in the API or in the available implementations. Advanced programming concepts such as hierarchical parallelism and use of dynamic shared memory were a particular area of concern. In this paper, we apply recent improvements and extensions in the LLVM/Clang OpenMP compiler and runtime library to the Kokkos backend that targets GPUs via OpenMP offload. We focus on efficient hierarchical parallelism and use of fast GPU scratch memory. We compare the performance of applications written using the Kokkos library with this improved OpenMP backend against the same programs using the CUDA and HIP backends. This evaluation shows progress toward closing the performance gaps between native and OpenMP backends and offers insights that may be useful to users and implementers of other runtime systems and programming frameworks for GPUs.
Rahulkumar Gayatri, Shilei Tian, Stephen Olivier, Eric Wright, Johannes Doerfert
HiPC1
2022 Kokkos 3: Programming Model Extensions for the Exascale Era
abstract
As the push towards exascale hardware has increased the diversity of system architectures, performance portability has become a critical aspect for scientific software. We describe the Kokkos Performance Portable Programming Model that allows developers to write single source applications for diverse high-performance computing architectures. Kokkos provides key abstractions for both the compute and memory hierarchy of modern hardware. We describe the novel abstractions that have been added to Kokkos version 3 such as hierarchical parallelism, containers, task graphs, and arbitrary-sized atomic operations to prepare for exascale era architectures. We demonstrate the performance of these new features with reproducible benchmarks on CPUs and GPUs.
Christian Trott, Damien Lebrun-Grandié, Daniel Arndt 0003, Jan Ciesko, Vinh Q. Dang, Nathan D. Ellingwood, Rahulkumar Gayatri, Evan Harvey, Daisy S. Hollman, Daniel Ibanez, Nevin Liber, Jonathan R. Madsen, Jeff Miles, David Poliakoff, Amy Powell, Sivasankaran Rajamanickam, Mikael Simberg, Daniel Sunderland, Bruno Turcksin, Jeremiah J. Wilke
IEEE Trans. Parallel Distributed Syst.7
2021 Non-recurring engineering (NRE) best practices: a case study with the NERSC/NVIDIA OpenMP contract
abstract
The NERSC supercomputer, Perlmutter, consists of AMD CPUs and NVIDIA GPUs. NERSC users expect to be able to use OpenMP to take advantage of the highly capable GPUs. This paper describes how NERSC/NVIDIA constructed a Non-Recurring Engineering (NRE) contract to add OpenMP GPU-offload support to the NVIDIA HPC compilers. The paper describes how the contract incorporated the strengths of both parties and encouraged collaboration to improve the quality of the final deliverable. We include our best practices and how this particular contract took into account emerging OpenMP specifications, NERSC workload requirements, and how to use OpenMP most efficiently on GPU hardware. This paper includes OpenMP application performance results obtained with the NVIDIA compilers distributed in the NVIDIA HPC SDK.
Christopher S. Daley, Annemarie Southwell, Rahulkumar Gayatri, Scott Biersdorfff, Craig Toepfer, Güray Özen, Nicholas J. Wright
SC3
2021 Billion atom molecular dynamics simulations of carbon at extreme conditions and experimental time and length scales
abstract
Billion atom molecular dynamics (MD) using quantum-accurate machine-learning Spectral Neighbor Analysis Potential (SNAP) observed long-sought high pressure BC8 phase of carbon at extreme pressure (12 Mbar) and temperature (5,000 K). 24-hour, 4650 node production simulation on OLCF Summit demonstrated an unprecedented scaling and unmatched real-world performance of SNAP MD while sampling 1 nanosecond of physical time. Efficient implementation of SNAP force kernel in LAMMPS using the Kokkos CUDA backend on NVIDIA GPUs combined with excellent strong scaling (better than 97% parallel efficiency) enabled a peak computing rate of 50.0 PFLOPs (24.9% of theoretical peak) for a 20 billion atom MD simulation on the full Summit machine (27,900 GPUs). The peak MD performance of 6.21 Matom-steps/node-s is 22.9 times greater than a previous record for quantum-accurate MD. Near perfect weak scaling of SNAP MD highlights its excellent potential to advance the frontier of quantum-accurate MD to trillion atom simulations on upcoming exascale platforms.
Kien Nguyen-Cong, Jonathan T. Willman, Stan G. Moore, Anatoly B. Belonoshko, Rahulkumar Gayatri, Evan Weinberg, Mitchell A. Wood, Aidan P. Thompson, Ivan I. Oleynik
SC5
2020 Experiences in porting mini-applications to OpenACC and OpenMP on heterogeneous systems
abstract
Summary This article studies mini‐applications—Minisweep, GenASiS, GPP, and FF—that use computational methods commonly encountered in HPC. We have ported these applications to develop OpenACC and OpenMP versions, and evaluated their performance on Titan (Cray XK7 with K20x GPUs), Cori (Cray XC40 with Intel KNL), Summit (IBM AC922 with Volta GPUs), and Cori‐GPU (Cray CS‐Storm 500NX with Intel Skylake and Volta GPUs). Our goals are for these new ports to be useful to both application and compiler developers, to document and describe the lessons learned and the methodology to create optimized OpenMP and OpenACC versions, and to provide a description of possible migration paths between the two specifications. Cases where specific directives or code patterns result in improved performance for a given architecture are highlighted. We also include discussions of the functionality and maturity of the latest compilers available on the above platforms with respect to OpenACC or OpenMP implementations.
Verónica G. Vergara Larrea, Reuben D. Budiardja, Rahulkumar Gayatri, Christopher S. Daley, Oscar R. Hernandez, Wayne Joubert
Concurr. Comput. Pract. Exp.3
2013 Loop level speculation in a task based programming model
abstract
Uncountable loops (such as while loops in C) and if-conditions are some of the most common constructs in programming. While-loops are widely used to determine the convergence in linear algebra algorithms or goal finding problems from graph algorithms, to name a few. In general while-loops are used whenever the loop iteration space, the number of iterations a loop executes is unknown. Usually in while-loops, the execution of the next iteration is decided inside the current loop iteration (i.e. the execution of iteration i depends on the values computed in iteration i-1). This precludes their parallel execution in today's ubiquitous multi-core architectures. In this paper a technique to speculatively create parallel tasks from the next iterations before the current one completes is proposed. If consecutive loop-iterations are only control dependent, then multiple iterations can be executed simultaneously; later in the execution path, the runtime system will decide to either commit the results of such speculatively executed iterations or undo the changes made by them. Data dependences within or between non-speculative and speculative work are honored to guarantee correctness. The proposed technique is implemented in SMPSs, a task-based dataflow programming model for shared-memory multiprocessor architectures. The approach is evaluated on a set of applications from graph algorithms and linear algebra. Results are promising with an average increase in the speedup of 1.2x with 16 threads when compared to non speculative execution of the applications. The increase in the speedup is significant, since the performance gain is achieved over an already parallelized version of the benchmarks.
Rahulkumar Gayatri, Rosa M. Badia, Eduard Ayguadé
HiPC1
2012 Transactional Access to Shared Memory in StarSs, a Task Based Programming Model
Rahulkumar Gayatri, Rosa M. Badia, Eduard Ayguadé, Mikel Luján, Ian Watson
Euro-Par1