EDBT 2026 Demo / reviewers in the wild / expert
Krzysztof Rojek
dblp:06/8297 · also Krzysztof Andrzej Rojek
· DBLP profile ↗
12ranked-venue papers
6as first author
3since 2021 · last 2026
0000-0002-2635-7345ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resource optimization with MPI process malleability for dynamic workloads in HPC clustersabstractDynamic resource management is essential for optimizing computational efficiency in modern high-performance computing (HPC) environments, particularly as systems scale. While research has demonstrated the benefits of malleability in resource management systems (RMS), the adoption of such techniques in production environments remains limited due to challenges in standardization, interoperability, and usability. Addressing these gaps, this paper extends our prior work on the Dynamic Management of Resources (DMR) framework, which provides a modular and user-friendly approach to dynamic resource allocation. Building upon the original DMRlib reconfiguration runtime, this work integrates new methodology from the Malleability Module (MaM) of the Proteo framework, further enhancing reconfiguration capabilities with new spawning strategies and data redistribution methods. In this paper, we explore new malleability strategies in HPC dynamic workloads, such as merging MPI communicators and asynchronous reconfigurations, which offer new opportunities for dramatically reducing memory overhead. The proposed enhancements are rigorously evaluated on a world-class supercomputer, demonstrating improved resource utilization and workload efficiency. Results show that dynamic resource management can reduce the workload completion time by 40% and increase the resource utilization by over 20%, compared to static resource allocation. Sergio Iserte, Iker Martín-Álvarez, Krzysztof Rojek, José Ignacio Aliaga, María Isabel Castillo, Weronika Folwarska, Antonio J. Peña |
Future Gener. Comput. Syst. | 3 |
| 2024 | Single- and multi-GPU computing on NVIDIA- and AMD-based server platforms for solidification modeling applicationabstractSummary This work explores the performance of single‐ and multi‐GPU computing on state‐of‐the‐art NVIDIA‐ and AMD‐based server‐class hardware using various programming interfaces to accelerate a real‐world scientific application for solidification modeling based on the phase‐field method. The main computations of this memory‐bound application correspond to 20 stencils computed across grid nodes. We investigate the application's scalability for two basic schemes of organizing computation: without and with hiding data transfers behind computation, combined with using either peer‐to‐peer inter‐GPU data transfers through NVIDIA NVLink and AMD Infinity interconnects or communication over the PCIe and main memory. Among the studied programming interfaces is CUDA, HIP, and OpenMP Accelerator Model. While the first two are designed to write the codes for a specific hardware platform, OpenMP enables code portability between NVIDIA and AMD GPUs. The resulting performance is experimentally assessed on computing platforms containing NVIDIA V100 (up to 8 GPUs) and A100 (one GPU), as well as AMD MI210 (one device) and MI250 (up to 8 logical GPUs). Kamil Halbiniak, Norbert Meyer, Krzysztof Rojek |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Optimizing throughput of Seq2Seq model training on the IPU platform for AI-accelerated CFD simulations
Pawel Rosciszewski, Adam Krzywaniak, Sergio Iserte, Krzysztof Rojek, Pawel Gepner |
Future Gener. Comput. Syst. | 4 |
| 2020 | An study of the effect of process malleability in the energy efficiency on GPU-based clusters
Sergio Iserte, Krzysztof Rojek |
J. Supercomput. | 2 |
| 2019 | Machine learning method for energy reduction by utilizing dynamic mixed precision on GPU-based supercomputersabstractSummary In this work, we propose a method that allows us to reduce energy consumption of an application executed on supercomputing centers. The proposed method is based on a mixed precision arithmetic where the precision of data is calibrated at runtime. For this reason, we develop a modified version of the random forest algorithm. The effectiveness of the proposed approach is validated with a real‐life scientific application called MPDATA, which is part of the numerical model used in weather forecasting. The energy efficiency of the proposed method is examined using two GPU‐based clusters. The first of them is the Piz Daint supercomputer, currently ranked 3rd at the TOP500 list (November 2017). It is equipped with NVIDIA Tesla P100 GPU accelerators based on the Pascal architecture. The second is the MICLAB cluster containing NVIDIA Tesla K80 based on the Kepler architecture. The achieved results show that the proposed machine learning method allows us to provide the accuracy of computation comparable with that achieved double precision and reduce the energy consumption up to 36% compared to the double precision version of MPDATA. Krzysztof Rojek |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Energy-aware mechanism for stencil-based MPDATA algorithm with constraintsabstractSummary In this paper, we propose an energy‐aware task management mechanism designed for the forward‐in‐time algorithms running on multicore central processing units (CPUs), where the multidimensional positive definite advection transport algorithm stencil‐based algorithm is one of the representative examples. This mechanism is based on the dynamic voltage and frequency scaling technique and allows the reduction of energy consumption for an existing algorithm (or application) such that the predefined execution time is respected, without requiring any modifications in the algorithm itself. This paper also provides the formulation of a method for minimizing the energy consumption with time constraints, which is based on the adaptive scheduling with online modeling. Finally, using the autotuning technique, we provide the automation of the process for creation and determination of the best energy profile at runtime, even in the presence of additional CPU workloads. The experimental results on a 6‐core computing platform show that the proposed mechanism provides the energy savings of up to 1.43x when compared to the default Linux scaling governor. Also, we confirm the effectiveness of the self‐adaptive feature of the proposed mechanism, by showing its ability to maintain the requested execution time in spite of additional CPU workloads imposed by other applications. Krzysztof Rojek, Aleksandar Ilic, Roman Wyrzykowski, Leonel Sousa |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Systematic adaptation of stencil-based 3D MPDATA to GPU architecturesabstractSummary In this work, we focus on a systematic adaptation of the stencil‐based multidimensional positive definite advection transport algorithm (MPDATA) to different graphics processing unit (GPU)‐based computing platforms. Another objective of this work is to compare the performance of MPDATA on several platforms, including a multi‐GPU system with two NVIDIA Tesla K80 cards, and single‐card platforms with Tesla K20X, GeForce GTX TITAN, and GeForce GTX 980. The usage of the following optimization methods is proposed to improve the overall performance: (i) reducing the number of operations by the subexpression elimination when implementing 2.5D blocking; (ii) reorganization of boundary conditions for reducing branch instructions; (iii) advanced memory management to increase the coalesced memory access; and (iv) warps rearrangement for optimizing the data access to GPU global memory. The presented methods of the MPDATA adaptation to GPU architectures allow us to efficiently use many graphics processors within a single node by applying peer‐to‐peer data transfers between GPU global memories. We propose an auto‐tuning procedure to compensate architectural differences between the considered platforms. This procedure takes into account algorithm/GPU‐specific parameters. The proposed approach to adaptation of MPDATA to GPU architectures allows us to achieve up to 482.5 Gflop/s for the platform equipped with two NVIDIA K80 GPUs. Copyright © 2016 John Wiley & Sons, Ltd. Krzysztof Rojek, Roman Wyrzykowski, Lukasz Kuczynski |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Modeling power consumption of 3D MPDATA and the CG method on ARM and Intel multicore architecturesabstractWe propose an approach to estimate the power consumption of algorithms, as a function of the frequency and number of cores, using only a very reduced set of real power measures. In addition, we also provide the formulation of a method to select the voltage–frequency scaling–concurrency throttling configurations that should be tested in order to obtain accurate estimations of the power dissipation. The power models and selection methodology are verified using two real scientific application: the stencil-based 3D MPDATA algorithm and the conjugate gradient (CG) method for sparse linear systems. MPDATA is a crucial component of the EULAG model, which is widely used in weather forecast simulations. The CG algorithm is the keystone for iterative solution of sparse symmetric positive definite linear systems via Krylov subspace methods. The reliability of the method is confirmed for a variety of ARM and Intel architectures, where the estimated results correspond to the real measured values with the average error being slightly below 5% in all cases. Krzysztof Rojek, Enrique S. Quintana-Ortí, Roman Wyrzykowski |
J. Supercomput. | 1 |
| 2017 | Performance modeling of 3D MPDATA simulations on GPU clusterabstractThe goal of this study is to parallelize the multidimensional positive definite advection transport algorithm (MPDATA) across a computational cluster equipped with GPUs. Our approach permits us to provide an extensive overlapping GPU computations and data transfers, both between computational nodes, as well as between the GPU accelerator and CPU host within a node. For this aim, we decompose a computational domain into two unequal parts which correspond to either data dependent or data independent parts. Then, data transfers can be performed simultaneously with computations corresponding to the second part. Our approach allows for achieving 16.372 Tflop/s using 136 GPUs. To estimate the scalability of the proposed approach, a performance model dedicated to MPDATA simulations is developed. We focus on the analysis of computation and communication execution times, as well as the influence of overlapping data transfers and GPU computations, with regard to the number of nodes. Krzysztof Rojek, Roman Wyrzykowski |
J. Supercomput. | 1 |
| 2015 | Adaptation of fluid model EULAG to graphics processing unit architectureabstractSummary The goal of this study is to adapt the multiscale fluid solver EULerian or LAGrangian framewrok (EULAG) to future graphics processing units (GPU) platforms. The EULAG model has the proven record of successful applications, and excellent efficiency and scalability on conventional supercomputer architectures. Currently, the model is being implemented as the new dynamical core of the COSMO weather prediction framework. Within this study, two main modules of EULAG, namely the multidimensional positive definite advection transport algorithm (MPDATA) and the variational generalized conjugate residual, elliptic pressure solver Generalized Conjugate Residual (GCR) are analyzed and optimized. In this paper, a method is proposed, which ensures a comprehensive analysis of the resource consumption including registers, shared, and global memories. This method allows us to identify bottlenecks of the algorithm, including data transfers between host and global memory, global and shared memories, as well as GPU occupancy. We put the emphasis on providing a fixed memory access pattern, padding as well as organizing computation in the MPDATA algorithm. The testing and validation of the new GPU implementation have been carried out based on modeling decaying turbulence of a homogeneous incompressible fluid in a triply‐periodic cube. Simulations performed using the standard version of EULAG and its new GPU implementation give similar solutions. Preliminary results show a promising increase in terms of computational efficiency. Copyright © 2014 John Wiley & Sons, Ltd. Krzysztof Rojek, Milosz Ciznicki, Bogdan Rosa, Piotr Kopta, Michal Kulczewski, Krzysztof Kurowski, Zbigniew Pawel Piotrowski, Lukasz Szustak, Damian Karol Wójcik, Roman Wyrzykowski |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Parallelization of 2D MPDATA EULAG algorithm on hybrid architectures with GPU accelerators
Roman Wyrzykowski, Lukasz Szustak, Krzysztof Rojek |
Parallel Comput. | 3 |
| 2012 | Model-driven adaptation of double-precision matrix multiplication to the Cell processor architecture
Roman Wyrzykowski, Krzysztof Rojek, Lukasz Szustak |
Parallel Comput. | 2 |