EDBT 2026 Demo / reviewers in the wild / expert
Pawel Czarnul
dblp:73/4985
· DBLP profile ↗
21ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-4918-9196ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Energy-Performance Optimization Tool Extension with Periodic Power Cap TuningabstractThis paper presents an improved solution for optimizing energy usage energy-performance trade-offs by means of power capping. We propose a new version of the DEPO tool that supports periodic tuning for dynamically changing workloads, targeted ultimately at deployment in the cloud at the Centre of Informatics Tricity Academic Supercomputer and networK, Gdańsk, Poland. We validated the approach using 27 experimental scenarios by executing three different applications in succession, generated as permutations with repetition, and compared the results to base runs without power capping. Across the experiments, periodic tuning delivers greater savings in energy consumption or energy-delay product (EDP) than singleshot tuning with a fixed tuning duration, eliminating the need for manual selection of that duration. The method also exhibits higher overall stability across all tested application orderings: Integer Sort (IS), Scalar Pentadiagonal Solver (SP), and Unstructured Adaptive (UA) from the NAS Parallel Benchmarks (NPB) suite. Evaluations were performed on two dual-socket servers equipped with 2x Intel Xeon Silver 4316 (Ice Lake) and 2x Intel Xeon Gold 6130 (Skylake) CPUs. The software is available as open source. Dawid Szmidka, Adam Krzywaniak, Pawel Czarnul, Jerzy Proficz |
ICPADS | 3 |
| 2025 | Optimization of resource-aware parallel and distributed computing: a reviewabstractThis paper presents a review of state-of-the-art solutions concerning the optimization of computing in the field of parallel and distributed systems. Firstly, we contribute by identifying resources and quality metrics in this context including servers, network interconnects, storage systems, computational devices as well as execution time/performance, energy, security, and error vulnerability, respectively. We subsequently identify commonly used problem formulations and algorithms for integer linear programming, greedy algorithms, dynamic programming, genetic algorithms, particle swarm optimization, ant colony optimization, game theory, and reinforcement learning. Afterward, we characterize frequently considered optimization problems by stating these terms in domains such as data centers, cloud, fog, blockchain, high performance, and volunteer computing. Based on the extensive analysis, we identify how particular resources and corresponding quality metrics are considered in these domains and which problem formulations are used for which system types, either parallel or distributed environments. This allows us to formulate open research problems and challenges in this field and analyze research interest in problem formulations/domains in recent years. Pawel Czarnul, Marcel Antal, Hamza Baniata, Dalvan Griebler, Attila Kertész, Christoph W. Kessler, Andreas Kouloumpris, Salko Kovacic, András Márkus, Maria K. Michael, Panagiota Nikolaou, Isil Öz, Radu Prodan, Gordana Rakic |
J. Supercomput. | 1 |
| 2024 | Dataset Characteristics and Their Impact on Offline Policy Learning of Contextual Multi-Armed Bandits
Piotr Januszewski, Dominik Grzegorzek, Pawel Czarnul |
ICAART (2) | 3 |
| 2023 | Performance assessment of OpenMP constructs and benchmarks using modern compilers and multi-core CPUsabstractConsidering ongoing developments of both modern CPUs, especially in the context of increasing numbers of cores, cache memory and architectures as well as compilers there is a constant need for benchmarking representative and frequently run workloads.The key metric is speed-up as the computational power of modern CPUs stems mainly from using multiple cores.In this paper, we show and discuss results from running codes such as: batch normalization, convolution, linear function, matrix multiplication, prime number test and wave equation; using compilers such as: GNU gcc, LLVM clang, icx, icc; run on four different 1 or 2-socket systems: 1 x Intel Core i7-5960X, 1 x Intel Core i9-9940X, 2 x Intel Xeon Platinum 8280L, 2 x Intel Xeon Gold 6130.Results can be regarded as suggestions concerning scaling on particular CPUs including recommended thread number configurations. Bartlomiej Gawrych, Pawel Czarnul |
FedCSIS | 2 |
| 2023 | UNRES-GPU for physics-based coarse-grained simulations of protein systems at biological time- and size-scalesabstractSUMMARY: The UNited RESisdue (UNRES) package for coarse-grained simulations, which has recently been optimized to treat large protein systems, has been implemented on Graphical Processor Units (GPUs). An over 100-time speed-up of the GPU code (run on an NVIDIA A100) with respect to the sequential code and an 8.5 speed-up with respect to the parallel Open Multi-Processing (OpenMP) code (run on 32 cores of 2 AMD EPYC 7313 Central Processor Units (CPUs)) has been achieved for large proteins (with size over 10 000 residues). Due to the averaging over the fine-grain degrees of freedom, 1 time unit of UNRES simulations is equivalent to about 1000 time units of laboratory time; therefore, millisecond time scale of large protein systems can be reached with the UNRES-GPU code. AVAILABILITY AND IMPLEMENTATION: The source code of UNRES-GPU along with the benchmarks used for tests is available at https://projects.task.gda.pl/eurohpcpl-public/unres. Krzysztof M. Ocetkiewicz, Cezary Czaplewski, Henryk Krawczyk, Agnieszka G. Lipska, Adam Liwo, Jerzy Proficz, Adam K. Sieradzan, Pawel Czarnul |
Bioinform. | 8 |
| 2023 | A multithreaded CUDA and OpenMP based power-aware programming framework for multi-node GPU systemsabstractSummary In the article, we have proposed a framework that allows programming a parallel application for a multi‐node system, with one or more graphical processing units (GPUs) per node, using an OpenMP+extended CUDA API. OpenMP is used for launching threads responsible for management of particular GPUs and extended CUDA calls allow to transfer data and launch kernels on local and remote GPUs. The framework hides inter‐node MPI communication from the programmer. For optimization, the implementation takes advantage of the MPI_THREAD_MULTIPLE mode allowing: multiple threads handling distinct GPUs as well as overlapping communication and computations transparently using multiple CUDA streams. The solution allows data parallelization across available GPUs in order to minimize execution time and supports a power‐aware mode in which GPUs are automatically selected for computations using a greedy approach in order not to exceed an imposed power limit. We have implemented and benchmarked three parallel applications including: finding the largest divisors; verification of the Collatz conjecture; finding patterns in vectors. These were tested on three various systems: a GPU cluster with 16 nodes, each with NVIDIA GTX 1060 GPU; a powerful 2‐node system—one node with 8 NVIDIA Quadro RTX 6000 GPUs, the second with 4 NVIDIA Quadro RTX 5000 GPUs; a heterogeneous environment with one node with 2 NVIDIA RTX 2080 and 2 nodes with NVIDIA GTX 1060 GPUs. We demonstrated effectiveness of the framework through execution times versus power caps within ranges of 100–1400 W, 250–3000 W, and 125–600 W for these systems respectively as well as gains from using two versus one CUDA streams per GPU. Finally, we have shown that for the testbed applications the solution allows to obtain high speed‐ups between 89.3% and 97.4% of the theoretically assessed ideal ones, for 16 nodes and 2 CUDA streams, demonstrating very good parallel efficiency. Pawel Czarnul |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Dynamic GPU power capping with online performance tracing for energy efficient GPU computing using DEPO tool
Adam Krzywaniak, Pawel Czarnul, Jerzy Proficz |
Future Gener. Comput. Syst. | 2 |
| 2022 | Optimization of Data Assignment for Parallel Processing in a Hybrid Heterogeneous Environment Using Integer Linear ProgrammingabstractAbstract In the paper we investigate a practical approach to application of integer linear programming for optimization of data assignment to compute units in a multi-level heterogeneous environment with various compute devices, including CPUs, GPUs and Intel Xeon Phis. The model considers an application that processes a large number of data chunks in parallel on various compute units and takes into account computations, communication including bandwidths and latencies, partitioning, merging, initialization, overhead for computational kernel launch and cleanup. We show that theoretical results from our model are close to real results as differences do not exceed 5% for larger data sizes, with up to 16.7% for smaller data sizes. For an exemplary workload based on solving systems of equations of various sizes with various compute-to-communication ratios we demonstrate that using an integer linear programming solver (lp_solve) with timeouts allows to obtain significantly better total (solver+application) run times than runs without timeouts, also significantly better than arbitrary chosen ones. We show that OpenCL 1.2’s device fission allows to obtain better performance in heterogeneous CPU+GPU environments compared to the GPU-only and the default CPU+GPU configuration, where a whole device is assigned for computations leaving no resources for GPU management. Tomasz Boinski, Pawel Czarnul |
Comput. J. | 2 |
| 2022 | DEPO: A dynamic energy-performance optimizer tool for automatic power capping for energy efficient high-performance computingabstractAbstract In the article we propose an automatic power capping software tool DEPO that allows one to perform runtime optimization of performance and energy related metrics. For an assumed application model with an initialization phase followed by a running phase with uniform compute and memory intensity, the tool performs automatic tuning engaging one of the two exploration algorithms—linear search (LS) and golden section search (GSS), finds a power cap optimizing a given metric and sets it for the remaining computations. The considered metrics include energy (E), energy‐delay sum, energy‐delay product. We present experimental results obtained for a set of benchmarks that differ in compute and memory intensity—parallel custom built OpenMP implementations of: numerical integration, heat distribution simulation (HEAT), fast Fourier transform (FFT), and additionally NAS parallel benchmarks: CG, MG, BT, SP, and LU. Tests were performed using multi‐core CPUs that are representatives of modern servers and the desktop family: 2 Intel Xeon E5‐2670 v3 CPU (Haswell‐EP) and Intel i7‐9700K CPU (Coffee Lake). The results show that our approach enabled considerable improvements for the tested metrics, for example, for HEAT and Coffee Lake we minimized energy by 50% at the cost of a 15% increase in execution time (LS), for FFT energy was minimized by 40% at a 25.5% increase in execution time (GSS), for SP and Haswell energy was minimized by 25% at the cost of an 18.5% time increase and for Coffee Lake energy was decreased by 56% with a 12% time increase. Adam Krzywaniak, Pawel Czarnul, Jerzy Proficz |
Softw. Pract. Exp. | 2 |
| 2021 | Human awareness versus Autonomous Vehicles view: comparison of reaction times during emergenciesabstractHuman safety is one of the most critical factors when a new technology is introduced to the everyday use. It was no different in the case of Autonomous Vehicles (AV), designed to replace generally available Conventional Vehicles (CV) in the future. AV rules, from the start, focus on guaranteeing safety for passengers and other road users, and these assumptions usually work during normal traffic conditions. However, there is still a problem with proper reaction time to sudden, dangerous and unexpected scenarios like a running animal on a rural road during the night. In this paper, we compare human and AV responses to sudden scenarios and accidents. As the AV topic can be analyzed as an ICT system, we review modern sensors, computer architectures and algorithms designed for this type of problems. Beside regular analysis, we also show which algorithms can run simultaneously and if vehicles have proper tools to guarantee safety during regular system delays. As a final result, we present a diagram which depicts Autonomous Vehicle logic and allows to identify bottlenecks. Additionally, the analysis shows how different refresh rates and algorithm execution times can affect the braking distance thus safety of other road users. Aleksander Rydzewski, Pawel Czarnul |
IV | 2 |
| 2020 | Development and benchmarking a parallel Data AcQuisition framework using MPI with hash and hash+tree structures in a cluster environmentabstractThe following topics are dealt with: parallel processing; multiprocessing systems; resource allocation; multi-threading; message passing; application program interfaces; cloud computing; optimisation; program compilers; parallel architectures. Pawel Czarnul, Grzegorz Golaszewski, Grzegorz Jereczek, Maciej Maciejewski |
ISPDC | 1 |
| 2019 | Performance evaluation of Unified Memory with prefetching and oversubscription for selected parallel CUDA applications on NVIDIA Pascal and Volta GPUsabstractThe paper presents assessment of Unified Memory performance with data prefetching and memory oversubscription. Several versions of code are used with: standard memory management, standard Unified Memory and optimized Unified Memory with programmer-assisted data prefetching. Evaluation of execution times is provided for four applications: Sobel and image rotation filters, stream image processing and computational fluid dynamic simulation, performed on Pascal and Volta architecture GPUs—NVIDIA GTX 1080 and NVIDIA V100 cards. Furthermore, we evaluate the possibility of allocating more memory than available on GPUs and assess performance of codes using the three aforementioned implementations, including memory oversubscription available in CUDA. Results serve as recommendations and hints for other similar codes regarding expected performance on modern and already widely available GPUs. Marcin Knap, Pawel Czarnul |
J. Supercomput. | 2 |
| 2018 | Analyzing energy/performance trade-offs with power capping for parallel applications on modern multi and many core processorsabstractIn the paper we present extensive results from analyzing energy/performance trade-offs with power capping observed on four different modern CPUs, for three different parallel applications such as 2D heat distribution, numerical integration and Fast Fourier Transform.The CPU tested represent both multi-core type CPUs such as Intel R Xeon R E5, desktop and mobile i7 as well as many-core Intel R Xeon Phi TM x200 but also server, desktop and mobile solutions used widely nowadays.We show that using enforced power caps we can find points of lower than default energy consumption but mostly for desktop and mobile solutions at the cost of increased execution time.We show with particular numbers how energy consumed, power consumption and execution time change for the point of minimum energy used versus the default configuration with no power limit, for each application and each tested CPU. Adam Krzywaniak, Jerzy Proficz, Pawel Czarnul |
FedCSIS | 3 |
| 2018 | Parallelization of large vector similarity computations in a hybrid CPU+GPU environmentabstractThe paper presents design, implementation and tuning of a hybrid parallel OpenMP+CUDA code for computation of similarity between pairs of a large number of multidimensional vectors. The problem has a wide range of applications, and consequently its optimization is of high importance, especially on currently widespread hybrid CPU+GPU systems targeted in the paper. The following are presented and tested for computation of all vector pairs: tuning of a GPU kernel with consideration of memory coalescing and using shared memory, minimization of GPU memory allocation costs, optimization of CPU–GPU communication in terms of size of data sent, overlapping CPU–GPU communication and kernel execution, concurrent kernel execution, determination of best sizes for data batches processed on CPUs and GPUs along with best GPU grid sizes. It is shown that all codes scale in hybrid environments with various relative performances of compute devices, even for a case when comparisons of various vector pairs take various amounts of time. Tests were performed on two high-performance hybrid systems with: 2 x Intel Xeon E5-2640 CPU + 2 x NVIDIA Tesla K20m and latest generation 2 x Intel Xeon CPU E5-2620 v4 + NVIDIA’s Pascal generation GTX 1070 cards. Results demonstrate expected improvements and beneficial optimizations important for users incorporating such types of computations into their parallel codes run on similar systems. Pawel Czarnul |
J. Supercomput. | 1 |
| 2017 | Performance evaluation of unified memory and dynamic parallelism for selected parallel CUDA applicationsabstractThe aim of this paper is to evaluate performance of new CUDA mechanisms—unified memory and dynamic parallelism for real parallel applications compared to standard CUDA API versions. In order to gain insight into performance of these mechanisms, we decided to implement three applications with control and data flow typical of SPMD, geometric SPMD and divide-and-conquer schemes, which were then used for tests and experiments. Specifically, tested applications include verification of Goldbach’s conjecture, 2D heat transfer simulation and adaptive numerical integration. We experimented with various ways of how dynamic parallelism can be deployed into an existing implementation and be optimized further. Subsequently, we compared the best dynamic parallelism and unified memory versions to respective standard API counterparts. It was shown that usage of dynamic parallelism resulted in improvement in performance for heat simulation, better than static but worse than an iterative version for numerical integration and finally worse results for Golbach’s conjecture verification. In most cases, unified memory results in decrease in performance. On the other hand, both mechanisms can contribute to simpler and more readable codes. For dynamic parallelism, it applies to algorithms in which it can be naturally applied. Unified memory generally makes it easier for a programmer to enter the CUDA programming paradigm as it resembles the traditional memory allocation/usage pattern. Lukasz Jarzabek, Pawel Czarnul |
J. Supercomput. | 2 |
| 2016 | Modeling energy consumption of parallel applicationsabstractThe paper presents modeling and simulation of energy consumption of two types of parallel applications: geometric Single Program Multiple Data (SPMD) and divide-and-conquer (DAC).Simulation is performed in a new MERPSYS (Modeling Efficiency, Reliability and Power consumption of multilevel parallel HPC SYStems using CPUs and GPUs) environment.Model of an application uses the Java language with extensions representing message exchange between processes working in parallel.Simulation is performed by running threads representing distinct process codes of an application, with consideration of process counts.Instead of running time consuming calculations, their times are simulated using functions representing computational time dependent on input data sizes.The simulator considers performance and power consumption values for compute devices stored in its database.We performed verification of running the two applications on up to 512 and 1024 processes respectively on a large cluster from Academic Computer Center in Gdansk demonstrating a high degree of accuracy between simulated and measured results. Pawel Czarnul, Jaroslaw Kuchta, Pawel Rosciszewski, Jerzy Proficz |
FedCSIS | 1 |
| 2016 | KernelHive: a new workflow-based framework for multilevel high performance computing using clusters and workstations with CPUs and GPUsabstractSummary The paper presents a new open‐source framework called KernelHive for multilevel parallelization of computations among various clusters, cluster nodes, and finally, among both CPUs and GPUs for a particular application. An application is modeled as an acyclic directed graph with a possibility to run nodes in parallel and automatic expansion of nodes (called node unrolling) depending on the number of computation units available. A methodology is proposed for parallelization and mapping of an application to the environment that includes selection of devices using a chosen optimizer, selection of best grid configurations for compute devices, optimization of data partitioning and the execution. One of possibly many scheduling algorithms can be selected considering execution time, power consumption, and so on. An easy‐to‐use GUI is provided for modeling and monitoring with a repository of ready‐to‐use constructs and computational kernels. The methodology, execution times, and scalability have been demonstrated for a distributed and parallel password‐breaking example run in a heterogeneous environment with a cluster and servers with different numbers of nodes and both CPUs and GPUs. Additionally, performance of the framework has been compared with an MPI + OpenCL implementation using a parallel geospatial interpolation application employing up to 40 cluster nodes and 320 cores. Copyright © 2015 John Wiley & Sons, Ltd. Pawel Rosciszewski, Pawel Czarnul, Rafal Lewandowski, Marcel Schally-Kacprzak |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Comparison of selected algorithms for scheduling workflow applications with dynamically changing service availabilityabstractThis paper compares the quality and execution times of several algorithms for scheduling service based workflow applications with changeable service availability and parameters. A workflow is defined as an acyclic directed graph with nodes corresponding to tasks and edges to dependencies between tasks. For each task, one out of several available services needs to be chosen and scheduled to minimize the workflow execution time and keep the cost of service within the budget. During the execution of a workflow, some services may become unavailable, new ones may appear, and costs and execution times may change with a certain probability. Rescheduling is needed to obtain a better schedule. A solution is proposed on how integer linear programming can be used to solve this problem to obtain optimal solutions for smaller problems or suboptimal solutions for larger ones. It is compared side-by-side with GAIN, divide-and-conquer, and genetic algorithms for various probabilities of service unavailability or change in service parameters. The algorithms are implemented and subsequently tested in a real BeesyCluster environment. Pawel Czarnul |
J. Zhejiang Univ. Sci. C | 1 |
| 2013 | Design of a distributed system using mobile devices and workflow management for measurement and control of a smart home and healthabstractThe paper presents design of a distributed system for measurements and control of a smart home including temper-atures, light, fire danger, health problems of inhabitants such as increased body temperature, a person falling etc. This is done by integration of mobile devices and standards, distributed service based middleware BeesyCluster and a workflow management system. Mobile devices are used to measure the parameters and are coupled with computers for remote access. The latter is made possible by exposing services in the distributed middleware. Such services can then be integrated into complex workflow applications for constant monitoring and analysis of input data, instant feedback such as turning off lights or heating or notification of medical emergency. Interfaces are defined for particular system components and implementation hints for measurements using the Java Mobile Sensor JSR 256 API are provided. Pawel Czarnul |
HSI | 1 |
| 2013 | Modeling, run-time optimization and execution of distributed workflow applications in the JEE-based BeesyCluster environmentabstractThe paper presents a complete solution for modeling scientific and business workflow applications, static and just-in-time QoS selection of services and workflow execution in a real environment. The workflow application is modeled as an acyclic directed graph where nodes denote tasks and edges denote dependencies between the tasks. The BeesyCluster middleware is used to allow providers to publish services from sequential or parallel applications, from their servers or clusters. Optimization algorithms are proposed to select a capable service for each task so that a global criterion is optimized such as a product of workflow execution time and cost, a linear combination of those or minimization of the time with a cost constraint. The paper presents implementation details of the multithreaded workflow execution engine implemented in JEE. Several tests were performed for three different optimization goals for two business and scientific workflow applications. Finally, the overhead of the solution is presented. Pawel Czarnul |
J. Supercomput. | 1 |
| 2013 | A model, design, and implementation of an efficient multithreaded workflow execution engine with data streaming, caching, and storage constraintsabstractThe paper proposes a model, design, and implementation of an efficient multithreaded engine for execution of distributed service-based workflows with data streaming defined on a per task basis. The implementation takes into account capacity constraints of the servers on which services are installed and the workflow data footprint if needed. Furthermore, it also considers storage space of the workflow execution engine and its cost. Caching service output data is implemented to speed up the execution of the workflow. Input data is partitioned into data packets, which are passed and processed by services previously selected for workflow tasks so that the aforementioned constraints are met. Performance impact of the proposed mechanisms is investigated for workflow structures common in acyclic directed graph workflow applications. It is shown for a real workflow with distributed processing of digital media content that the initial budget needs to be properly distributed between both the cost of services, but also the cost of intermediate storage to obtain good workflow execution times. Pawel Czarnul |
J. Supercomput. | 1 |