EDBT 2026 Demo / reviewers in the wild / expert
Maarit J. Korpi-Lagg
dblp:173/2941 · also Maarit J. Käpylä, Maarit Käpylä
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-9614-2200ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Stencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning StrategiesabstractABSTRACT Over the last ten years, graphics processors have become the de facto accelerator for data‐parallel tasks in various branches of high‐performance computing, including machine learning and computational sciences. However, with the recent introduction of AMD‐manufactured graphics processors to the world's fastest supercomputers, tuning strategies established for previous hardware generations must be re‐evaluated. In this study, we evaluate the performance and energy efficiency of stencil computations on modern datacenter graphics processors and propose a tuning strategy for fusing cache‐heavy stencil kernels. The studied cases comprise both synthetic and practical applications, which involve the evaluation of linear and nonlinear stencil functions in one to three dimensions. Our experiments reveal that AMD and Nvidia graphics processors exhibit key differences in both hardware and software, necessitating platform‐specific tuning to reach their full computational potential. Johannes Pekkilä, Oskar Lappi, Fredrik Robertsen, Maarit J. Korpi-Lagg |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Supporting Opportunistic Data Operations for Data-Intensive Computational ApplicationsabstractA long running data-intensive computational application acquires costly computing resources. With the emerging new architectures, like computing systems with multiple nodes of many-core CPUs and accelerators, while domain-specific tools and libraries employed in such an application leverage high parallelism on accelerators for intensive computations, the remaining resources can potentially be utilized for other application-related data operations. Such data operations, called opportunistic data operations in this work, must usually be carried out for post-processing or follow-up analytics based on results produced during the runtime of the application. These operations are not easily backfilled or preempted under the guidance of the domain scientist or by common task scheduling systems due to their complex dependencies.In this paper, we introduce a framework for domain scientists to identify and execute opportunistic data operation tasks. With a minimal specification or modification of the main application, the scientists can specify, monitor, and execute opportunistic tasks independently from the main application and the framework will detect underutilized resources to execute these tasks, thereby, optimizing utilization efficiency within the allocated resources. We present experiments to demonstrate the applicability of our framework on a magnetic field modeling running on the LUMI computing system. Minh-Tri Nguyen, Anh-Dung Nguyen, Jarno Rantaharju, Touko Puro, Matthias Rheinhardt, Maarit J. Korpi-Lagg, Hong Linh Truong 0001 |
IEEE Big Data | 6 |
| 2024 | SOMA: Observability, monitoring, and in situ analytics for exascale applicationsabstractSummary With the rise of exascale systems and large, data‐centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service‐based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application‐specific diagnostic data and performance data addresses these needs. We present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi‐GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring. Dewi Yokelson, Oskar Lappi, Srinivasan Ramesh, Miikka S. Väisälä, Kevin A. Huck, Touko Puro, Boyana Norris, Maarit J. Korpi-Lagg, Keijo Heljanko, Allen D. Malony |
Concurr. Comput. Pract. Exp. | 8 |
| 2022 | Scalable communication for high-order stencil computations using CUDA-aware MPIabstractModern compute nodes in high-performance computing provide a tremendous level of parallelism and processing power. However, as arithmetic performance has been observed to increase at a faster rate relative to memory and network bandwidths, optimizing data movement has become critical for achieving strong scaling in many communication-heavy applications. This performance gap has been further accentuated with the introduction of graphics processing units, which can provide by multiple factors higher throughput in data-parallel tasks than central processing units. In this work, we explore the computational aspects of iterative stencil loops and implement a generic communication scheme using CUDA-aware MPI, which we use to accelerate magnetohydrodynamics simulations based on high-order finite differences and third-order Runge–Kutta integration. We put particular focus on improving intra-node locality of workloads. Our GPU implementation scales strongly from one to 64 devices at 50%–87% of the expected efficiency based on a theoretical performance model. Compared with a multi-core CPU solver, our implementation exhibits 20–60× speedup and 9–12× improved energy efficiency in compute-bound benchmarks on 16 nodes. Johannes Pekkilä, Miikka S. Väisälä, Maarit J. Korpi-Lagg, Matthias Rheinhardt, Oskar Lappi |
Parallel Comput. | 3 |
| 2016 | Singular Value Decomposition update and its application to (Inc)-OP-ELM
Alexander Grigorievskiy, Yoan Miché, Maarit J. Korpi-Lagg, Amaury Lendasse |
Neurocomputing | 3 |