EDBT 2026 Demo / reviewers in the wild / expert
Mathialakan Thavappiragasam
dblp:161/5704
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-1240-3475ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI and HPC Applications on Leadership Computing Platforms: Performance and Scalability StudiesabstractAs HPC systems move into the exascale era an increasing diversity of processing hardware is being deployed. The last decade saw the ascendance of NVIDIA GPU-accelerated systems among the largest scale HPC systems and spurred the need for application developers to consider approaches to performance portability that preserved developer productivity. This challenge has been compounded in the last several years by the introduction of the first two exascale systems, Frontier and Aurora (\#2 and \#3 on the November 2024 Top 500 list respectively). These systems utilize new and different GPUs, with the AMD MI-250X GPU on Frontier and the Intel Data Center GPU Max 1550 on Aurora. This study investigates the performance and qualitative performance portability of$\mathbf{1 2}$HPC and ML applications on three large scale HPC systems that utilize GPUs from the three different vendors: Frontier (AMD), Aurora (Intel), and Polaris (NVIDIA A100). The performance of these applications is evaluated on single GPU, single node, and multinode scales on each of the systems. We show that the figures-of-merit (FOMs) of the applications on a single GPU of Aurora and Frontier ranged from$0.9-4 x$and$0.8-2.5 x$, respectively, the performance on a GPU of Polaris. We also show that the FOMs on a single node of Aurora and Frontier ranged from 1.3-6.3x and 0.8-2.6x, respectively, a single node of Polaris. The applications were scaled up to 512 nodes showing good scaling efficiency across the board. Finally, we discuss useful concepts and experiences gained in running diverse applications on diverse HPC systems. JaeHyuk Kwack, Colleen Bertoni, Umesh Unnikrishnan, Riccardo Balin, Khalid Hossain, Yasaman Ghadar, Timothy J. Williams, Abhishek Bagusetty, Mathialakan Thavappiragasam, Väinö Hatanpää, Archit Vasan, John R. Tramm, Scott Parker |
IPDPS | 9 |
| 2023 | CPU-GPU Tuning for Modern Scientific Applications using Node-Level HeterogeneityabstractScientific applications must be tuned for performance to run efficiently on supercomputers having nodes with a CPU (or, a general-purpose host processor) and GPUs (or, accelerator device processors). Conventional wisdom suggests focusing tuning of applications for a GPU and making the CPU only have the role of offloading computation to the GPU, given the CPU's relatively miniscule amount of computational power. However, this is overly conservative for modern scientific applications, which include those using scientific workflows with real-time data constraints and AI/ML with low numerical precision requirements. This work identifies new performance opportunities for modern scientific applications via CPU-GPU tuning, a strategy that unifies and integrates tuning of the CPU and GPU performance parameters. Applying CPU-GPU tuning to a dot product representative of these applications run on the widely-used Summit supercomputer results in up to an 8.15x speedup. These results provide groundwork for auto-tuning software for applications run on supercomputers having node-level heterogeneity. Mathialakan Thavappiragasam, Vivek Kale |
HiPC | 1 |
| 2022 | Portability for GPU-accelerated molecular docking applications for cloud and HPC: can portable compiler directives provide performance across all platforms?abstractHigh-throughput structure-based screening of drug-like molecules has become a common tool in biomedical research. Recently, acceleration with graphics processing units (GPUs) has provided a large performance boost for molecular docking programs. Both cloud and high-performance computing (HPC) resources have been used for large screens with molecular docking programs; while NVIDIA GPUs have dominated cloud and HPC resources, new vendors such as AMD and Intel are now entering the field, creating the problem of software portability across different GPUs. Ideally, software productivity could be maximized with portable programming models that are able to maintain high performance across architectures. While in many cases compiler directives have been used as an easy way to offload parallel regions of a CPU-based program to a GPU accelerator, they may also be an attractive programming model for providing portability across different GPU vendors, in which case the porting process may proceed in the reverse direction: from low-level, architecture-specific code to higher-level directive-based abstractions. MiniMDock is a new mini-application (miniapp) designed to capture the essential computational kernels found in molecular docking calculations, such as are used in phar-maceutical drug discovery efforts, in order to test different solutions for porting across GPU architectures. Here we extend MiniMDock to GPU offloading with OpenMP directives, and compare to performance of kernels using CUDA and HIP on NVIDIA and AMD GPUs, respectively, as well as across different compilers, exploring performance bottlenecks. We document this reverse-porting process, from highly optimized device code to a higher-level version using directives, compare code structure, and describe barriers that were overcome in this effort. Mathialakan Thavappiragasam, Wael R. Elwasif, Ada Sedova |
CCGRID | 1 |
| 2022 | OpenMDlr: parallel, open-source tools for general protein structure modeling and refinement from pairwise distancesabstractSUMMARY: Easy-to-use, open-source, general-purpose programs for modeling a protein structure from inter-atomic distances are needed for modeling from experimental data and refinement of predicted protein structures. OpenMDlr is an open-source Python package for modeling protein structures from pairwise distances between any atoms, and optionally, dihedral angles. We provide a user-friendly input format for harnessing modern biomolecular force fields in an easy-to-install package that can efficiently make use of multiple compute cores. AVAILABILITY AND IMPLEMENTATION: OpenMDlr is available at https://github.com/BSDExabio/OpenMDlr-amber. The package is written in Python (versions 3.x). All dependencies are open-source and can be installed with the Conda package management system. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Russell B. Davidson, Jess Woods, T. Chad Effler, Mathialakan Thavappiragasam, Julie C. Mitchell, Jerry M. Parks, Ada Sedova |
Bioinform. | 4 |
| 2021 | Addressing Load Imbalance in Bioinformatics and Biomedical Applications: Efficient Scheduling across Multiple GPUsabstractComputational bioinformatics and biomedical applications frequently contain heterogeneously sized units of work or tasks, for instance due to variability in the sizes of biological sequences and molecules. Variable-sized workloads lead to load imbalances in parallel implementations which detract from efficiency and performance. Many modern computing resources now have multiple graphics processing units(GPUs) per computer for acceleration. These multiple GPU resources need to be used efficiently through balancing of workloads across the GPUs. OpenMP is a portable directive-based parallel programming API used ubiquitously in bioscience applications to program CPUs; recently, the use of OpenMP directives for GPU acceleration has become possible. Here, motivated by experiences with imbalanced loads in GPU-accelerated bioinformatics applications, we address the load balancing problem using OpenMP task-to-GPU scheduling combined with OpenMP GPU offloading for multiply heterogeneous workloads – loads with both variable input sizes, and simultaneously, variable convergence rates for algorithms with a stochastic component – scheduled across multiple GPUs. We aim to develop strategies which are both easy to use and have lower overheads, and may be incorporated incrementally in existing programs which already make use of OpenMP for CPU-based threading in order to make use of multi-GPU computers. We test different combinations of input size variability and convergence rate variability, and characterize the effects of these different scenarios on the performance of scheduling strategies across multiple GPUs with OpenMP. We present several dynamic scheduling solutions for different parallel patterns, explore optimizations, and provide publicly available example computational kernels to make these strategies easy to use in programs. This work will enable application developers to efficiently and easily use multiple GPUs for imbalanced workloads found in bioinformatics and biomedical applications. Mathialakan Thavappiragasam, Vivek Kale, Oscar R. Hernandez, Ada Sedova |
BIBM | 1 |