Jerzy Proficz

dblp:41/4044 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-2975-9339ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2
YearPublicationVenuePosition
2025 Dynamic Energy-Performance Optimization Tool Extension with Periodic Power Cap Tuning
abstract
This paper presents an improved solution for optimizing energy usage energy-performance trade-offs by means of power capping. We propose a new version of the DEPO tool that supports periodic tuning for dynamically changing workloads, targeted ultimately at deployment in the cloud at the Centre of Informatics Tricity Academic Supercomputer and networK, Gdańsk, Poland. We validated the approach using 27 experimental scenarios by executing three different applications in succession, generated as permutations with repetition, and compared the results to base runs without power capping. Across the experiments, periodic tuning delivers greater savings in energy consumption or energy-delay product (EDP) than singleshot tuning with a fixed tuning duration, eliminating the need for manual selection of that duration. The method also exhibits higher overall stability across all tested application orderings: Integer Sort (IS), Scalar Pentadiagonal Solver (SP), and Unstructured Adaptive (UA) from the NAS Parallel Benchmarks (NPB) suite. Evaluations were performed on two dual-socket servers equipped with 2x Intel Xeon Silver 4316 (Ice Lake) and 2x Intel Xeon Gold 6130 (Skylake) CPUs. The software is available as open source.
Dawid Szmidka, Adam Krzywaniak, Pawel Czarnul, Jerzy Proficz
ICPADS4
2023 UNRES-GPU for physics-based coarse-grained simulations of protein systems at biological time- and size-scales
abstract
SUMMARY: The UNited RESisdue (UNRES) package for coarse-grained simulations, which has recently been optimized to treat large protein systems, has been implemented on Graphical Processor Units (GPUs). An over 100-time speed-up of the GPU code (run on an NVIDIA A100) with respect to the sequential code and an 8.5 speed-up with respect to the parallel Open Multi-Processing (OpenMP) code (run on 32 cores of 2 AMD EPYC 7313 Central Processor Units (CPUs)) has been achieved for large proteins (with size over 10 000 residues). Due to the averaging over the fine-grain degrees of freedom, 1 time unit of UNRES simulations is equivalent to about 1000 time units of laboratory time; therefore, millisecond time scale of large protein systems can be reached with the UNRES-GPU code. AVAILABILITY AND IMPLEMENTATION: The source code of UNRES-GPU along with the benchmarks used for tests is available at https://projects.task.gda.pl/eurohpcpl-public/unres.
Krzysztof M. Ocetkiewicz, Cezary Czaplewski, Henryk Krawczyk, Agnieszka G. Lipska, Adam Liwo, Jerzy Proficz, Adam K. Sieradzan, Pawel Czarnul
Bioinform.6
2023 Dynamic GPU power capping with online performance tracing for energy efficient GPU computing using DEPO tool
Adam Krzywaniak, Pawel Czarnul, Jerzy Proficz
Future Gener. Comput. Syst.3
2022 DEPO: A dynamic energy-performance optimizer tool for automatic power capping for energy efficient high-performance computing
abstract
Abstract In the article we propose an automatic power capping software tool DEPO that allows one to perform runtime optimization of performance and energy related metrics. For an assumed application model with an initialization phase followed by a running phase with uniform compute and memory intensity, the tool performs automatic tuning engaging one of the two exploration algorithms—linear search (LS) and golden section search (GSS), finds a power cap optimizing a given metric and sets it for the remaining computations. The considered metrics include energy (E), energy‐delay sum, energy‐delay product. We present experimental results obtained for a set of benchmarks that differ in compute and memory intensity—parallel custom built OpenMP implementations of: numerical integration, heat distribution simulation (HEAT), fast Fourier transform (FFT), and additionally NAS parallel benchmarks: CG, MG, BT, SP, and LU. Tests were performed using multi‐core CPUs that are representatives of modern servers and the desktop family: 2 Intel Xeon E5‐2670 v3 CPU (Haswell‐EP) and Intel i7‐9700K CPU (Coffee Lake). The results show that our approach enabled considerable improvements for the tested metrics, for example, for HEAT and Coffee Lake we minimized energy by 50% at the cost of a 15% increase in execution time (LS), for FFT energy was minimized by 40% at a 25.5% increase in execution time (GSS), for SP and Haswell energy was minimized by 25% at the cost of an 18.5% time increase and for Coffee Lake energy was decreased by 56% with a 12% time increase.
Adam Krzywaniak, Pawel Czarnul, Jerzy Proficz
Softw. Pract. Exp.3
2021 All-gather Algorithms Resilient to Imbalanced Process Arrival Patterns
abstract
Two novel algorithms for the all-gather operation resilient to imbalanced process arrival patterns (PATs) are presented. The first one, Background Disseminated Ring (BDR), is based on the regular parallel ring algorithm often supplied in MPI implementations and exploits an auxiliary background thread for early data exchange from faster processes to accelerate the performed all-gather operation. The other algorithm, Background Sorted Linear synchronized tree with Broadcast (BSLB), is built upon the already existing PAP-aware gather algorithm, that is, Background Sorted Linear Synchronized tree (BSLS), followed by a regular broadcast distributing gathered data to all participating processes. The background of the imbalanced PAP subject is described, along with the PAP monitoring and evaluation topics. An experimental evaluation of the algorithms based on a proposed mini-benchmark is presented. The mini-benchmark was performed over 2,000 times in a typical HPC cluster architecture with homogeneous compute nodes. The obtained results are analyzed according to different PATs, data sizes, and process numbers, showing that the proposed optimization works well for various configurations, is scalable, and can significantly reduce the all-gather elapsed times, in our case, up to factor 1.9 or 47% in comparison with the best state-of-the-art solution.
Jerzy Proficz
ACM Trans. Archit. Code Optim.1
2021 Improving Clairvoyant: reduction algorithm resilient to imbalanced process arrival patterns
abstract
Abstract The Clairvoyant algorithm proposed in “A novel MPI reduction algorithm resilient to imbalances in process arrival times” was analyzed, commented and improved. The comments concern handling certain edge cases in the original pseudocode and description, i.e., adding another state of a process, improved cache friendliness more precise complexity estimations and some other issues improving the robustness of the algorithm implementation. The proposed improvements include skipping of idle loop rounds, simplifying generation of the ready set and management of the state array and an about 90-fold reduction in memory usage. Finally an extension enabling process arrival times (PATs) prediction was added: an additional background thread used to exchange the data with the PAT estimations. The performed tests, with a dedicated mini-benchmark executed in an HPC environment, showed correctness and improved performance of the solution, with comparison to the original or other state-of-the-art algorithms.
Jerzy Proficz, Krzysztof M. Ocetkiewicz
J. Supercomput.1
2020 Investigation into MPI All-Reduce Performance in a Distributed Cluster with Consideration of Imbalanced Process Arrival Patterns
Jerzy Proficz, Piotr Sumionka, Jaroslaw Skomial, Marcin Semeniuk, Karol Niedzielewski, Maciej Walczak
AINA1
2018 Analyzing energy/performance trade-offs with power capping for parallel applications on modern multi and many core processors
abstract
In the paper we present extensive results from analyzing energy/performance trade-offs with power capping observed on four different modern CPUs, for three different parallel applications such as 2D heat distribution, numerical integration and Fast Fourier Transform.The CPU tested represent both multi-core type CPUs such as Intel R Xeon R E5, desktop and mobile i7 as well as many-core Intel R Xeon Phi TM x200 but also server, desktop and mobile solutions used widely nowadays.We show that using enforced power caps we can find points of lower than default energy consumption but mostly for desktop and mobile solutions at the cost of increased execution time.We show with particular numbers how energy consumed, power consumption and execution time change for the point of minimum energy used versus the default configuration with no power limit, for each application and each tested CPU.
Adam Krzywaniak, Jerzy Proficz, Pawel Czarnul
FedCSIS2
2018 Improving all-reduce collective operations for imbalanced process arrival patterns
abstract
Two new algorithms for the all-reduce operation optimized for imbalanced process arrival patterns (PAPs) are presented: (1) sorted linear tree, (2) pre-reduced ring as well as a new way of online PAP detection, including process arrival time estimations, and their distribution between cooperating processes was introduced. The idea, pseudo-code, implementation details, benchmark for performance evaluation and a real case example for machine learning are provided. The results of the experiments were described and analyzed, showing that the proposed solution has high scalability and improved performance in comparison with the usually used ring and Rabenseifner algorithms.
Jerzy Proficz
J. Supercomput.1
2016 Modeling energy consumption of parallel applications
abstract
The paper presents modeling and simulation of energy consumption of two types of parallel applications: geometric Single Program Multiple Data (SPMD) and divide-and-conquer (DAC).Simulation is performed in a new MERPSYS (Modeling Efficiency, Reliability and Power consumption of multilevel parallel HPC SYStems using CPUs and GPUs) environment.Model of an application uses the Java language with extensions representing message exchange between processes working in parallel.Simulation is performed by running threads representing distinct process codes of an application, with consideration of process counts.Instead of running time consuming calculations, their times are simulated using functions representing computational time dependent on input data sizes.The simulator considers performance and power consumption values for compute devices stored in its database.We performed verification of running the two applications on up to 512 and 1024 processes respectively on a large cluster from Academic Computer Center in Gdansk demonstrating a high degree of accuracy between simulated and measured results.
Pawel Czarnul, Jaroslaw Kuchta, Pawel Rosciszewski, Jerzy Proficz
FedCSIS4
2010 The Task Graph Assignment for KASKADA Platform
Henryk Krawczyk, Jerzy Proficz
ICSOFT (1)2
2000 Suitability of the time controlled environment for race detection in distributed applications
Henryk Krawczyk, Bartosz Krysztop, Jerzy Proficz
Future Gener. Comput. Syst.3