VLDB 2026 Research / reviewers in the wild / expert
Lukasz Szustak
dblp:91/8296
· DBLP profile ↗
19ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-7429-6981ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating AMD EPYC CPU architectures on CFD applications
Marcin Lawenda, Lukasz Szustak, László Környei, Flavio Cesar Cunha Galeazzo, Pawel Bratek |
Future Gener. Comput. Syst. | 2 |
| 2025 | Profiling and Optimization of Multicard GPU Machine Learning JobsabstractABSTRACT The article discusses various model optimization techniques, providing a comprehensive analysis of key performance indicators. Several parallelization strategies for image recognition are analyzed, adapted to different hardware and software configurations, including distributed data parallelism and distributed hardware processing. Changing the tensor layout in PyTorch DataLoader from NCHW to NHWC and enabling pin _ memory has proven to be very beneficial and easy to implement. Furthermore, the impact of different performance techniques (DPO, LoRA, QLoRA, and QAT) on the tuning process of LLMs was investigated. LoRA allows for faster tuning, while requiring less VRAM compared to DPO. On the other hand, QAT is the most resource‐intensive method, with the slowest processing times. A significant portion of LLM tuning time is attributed to initializing new kernels and synchronizing multiple threads when memory operations are not dominant. Marcin Lawenda, Kyrylo Khloponin, Krzesimir Samborski, Lukasz Szustak |
Concurr. Comput. Pract. Exp. | 4 |
| 2025 | Prediction model of performance-energy trade-off for CFD codes on AMD-based cluster
Marcin Lawenda, Lukasz Szustak, László Környei |
Future Gener. Comput. Syst. | 2 |
| 2024 | Special Issue on the pervasive nature of HPC (PN-HPC)abstractSummary This special issue on the Pervasive Nature of HPC (PN‐HPC) collects an extension of the most valuable works presented at the sixth Workshop on Models, Algorithms and Methodologies for Hybrid Parallelism in New HPC Systems (MAMHYP‐22), held in Gdansk (Poland) in September 2022, jointly with the 14th conference on Parallel Processing and Applied Mathematics (PPAM‐22). New original papers related to the workshop themes are also included. The final aim is to provide a glimpse of the current state of knowledge related to the development of efficient methodologies and algorithms for HPC systems with multiple forms of parallelism. Marco Lapegna, Valeria Mele, Raffaele Montella, Lukasz Szustak |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Large-Scale Parallelization of Human Migration SimulationabstractForced displacement of people worldwide, for example, due to violent conflicts, is common in the modern world, and today more than 82 million people are forcibly displaced. This puts the problem of migration at the forefront of the most important problems of humanity. The Flee simulation code is an agent-based modeling tool that can forecast population displacements in civil war settings, but performing accurate simulations requires nonnegligible computational capacity. In this article, we present our approach to Flee parallelization for fast execution on multicore platforms, as well as discuss the computational complexity of the algorithm and its implementation. We benchmark parallelized code using supercomputers equipped with AMD EPYC Rome 7742 and Intel Xeon Platinum 8268 processors and investigate its performance across a range of alternative rule sets, different refinements in the spatial representation, and various numbers of agents representing displaced persons. We find that Flee scales excellently to up to 8192 cores for large cases, although very detailed location graphs can impose a large initialization time overhead. Derek Groen, Nikela Papadopoulou, Petros Anastasiadis, Marcin Lawenda, Lukasz Szustak, Sergiy Gogolenko, Hamid Arabnejad, Alireza Jahani |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | Profiling and optimization of Python-based social sciences applications on HPC systems by means of task and data parallelismabstractThe article presents optimization techniques for two Python-based large-scale social sciences applications: SN (Social Network) Simulator and KPM (Kernel Polynomial Method). These applications use MPI technology to transfer data between computing processes, which in the regular implementation leads to load imbalance and performance degradation. To avoid this effect, we propose a 2-stage optimization. In the first step, the order of tasks is changed, and in the second step, the tasks are divided into smaller ones for easier allocation. In addition, we focus on mitigating performance and memory bottlenecks using modern ccNUMA systems with multiple NUMA domains. As part of the performance analysis, the limitations of communication in data traffic between and within the processor were revealed and resolved through appropriate data allocation. Benchmarking was carried out, examining various environments, including vendors of traditional x86-64 and ARM-based processors for HPC. Lukasz Szustak, Marcin Lawenda, Sebastian Arming, Gregor Bankhamer, Christoph Schweimer, Robert Elsässer |
Future Gener. Comput. Syst. | 1 |
| 2023 | Reducing energy consumption using heterogeneous voltage frequency scaling of data-parallel applications for multicore systems
Pawel Bratek, Lukasz Szustak, Roman Wyrzykowski, Tomasz Olas |
J. Parallel Distributed Comput. | 2 |
| 2021 | About the granularity portability of block-based Krylov methods in heterogeneous computing environmentsabstractSummary Large‐scale problems in engineering and science often require the solution of sparse linear algebra problems and the Krylov subspace iteration methods (KM) have led to a major change in how users deal with them. But, for these solvers to use extreme‐scale hardware efficiently a lot of work was spent to redesign both the KM algorithms and their implementations to address challenges like extreme concurrency, complex memory hierarchies, costly data movement, and heterogeneous node architectures. All the redesign approaches bases the KM algorithm on block‐based strategies which lead to the Block‐KM (BKM) algorithm which has high granularity (i.e., the ratio of computation time to communication time). The work proposes novel parallel revisitation of the modules used in BKM which are based on the overlapping of communication and computation. Such revisitation is evaluated by a model of their granularity and verified on the basis of a case study related to a classical problem from numerical linear algebra. Luisa Carracciuolo, Valeria Mele, Lukasz Szustak |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Dynamic workload prediction and distribution in numerical modeling of solidification on multi-/manycore architecturesabstractSummary This work is a part of the global tendency to use modern computing systems for modeling the phase‐field phenomena. The main goal of this article is to improve the performance of a parallel application for the solidification modeling, assuming the dynamic intensity of computations in successive time steps when calculations are performed using a carefully selected group of nodes in the grid. A two‐step method is proposed to optimize the application for multi‐/manycore architectures. In the first step, the loop fusion is used to execute all kernels in a single nested loop and reduce the number of conditional operators. These modifications are vital to implementing the second step, which includes an algorithm for the dynamic workload prediction and load balancing across cores of a computing platform. Two versions of the algorithm are proposed—with the 1D and 2D maps used for predicting the computational domain within the grid. The proposed optimizations allow increasing the application performance significantly for all tested configurations of computing resources. The highest performance gain is achieved for two Intel Xeon Platinum 8180 CPUs, where the new code based on the 2D map yields the speedup of up to 2.74 times, while the usage of the proposed method with the 2D map for a single KNL accelerator permits reducing the execution time up to 1.91 times. Kamil Halbiniak, Tomasz Olas, Lukasz Szustak, Adam Kulawik, Marco Lapegna |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Exploration of OpenCL Heterogeneous Programming for Porting Solidification Modeling to CPU-GPU PlatformsabstractSummary This article provides a comprehensive study of OpenCL heterogeneous programming for porting applications to CPU–GPU computing platforms, with a real‐life application for the solidification modeling. The aim is to achieve a flexible workload distribution between available CPU–GPU resources and optimize application performance. Considering the solidification application as a use case, we explore the necessary steps required for (i) adaptation of an application to CPU–GPU platforms, and (ii) mapping the application workload onto the OpenCL programming model. The adaptation is based on a reformulation of steps developed previously for CPU–MIC architectures. The mapping process allows us to utilize OpenCL for harnessing CPU and GPU cores using data parallelism, as well as for the management of available compute devices with task parallelism. The resulting OpenCL code's performance and energy efficiency is experimentally studied for two platforms with powerful GPUs of various generations (with Kepler and Volta architectures). The experiments confirm the performance advantage of using computing resources of both GPUs and CPUs. The achieved benefit depends on the relationship between the computing power of CPUs and GPUs. Moreover, this gain entails the growth of the average power that increases the energy consumed during the application execution. Kamil Halbiniak, Lukasz Szustak, Tomasz Olas, Roman Wyrzykowski, Pawel Gepner |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | Architectural Adaptation and Performance-Energy Optimization for CFD Application on AMD EPYC RomeabstractThe advantages of the second-generation AMD EPYC Rome processors can be successfully used in the race to Exascale. However, the novel architecture's complexity makes it challenging to adapt demanding scientific codes - like stencil ones - to platforms with Rome CPUs. This article tackles this challenge by exploring the adaptation of the stencil-based CFD (computational fluid dynamics) application called MPDATA to these processors' influential features. We show that the previously proposed parametric adaptation methodology can be profitably applied to extend the performance portability of the memory-bound MPDATA on the AMD EPYC architecture. The extension of the parametric adaptation on the novel architecture requires careful consideration of two relevant aspects that reflect splitting the Rome architecture into multiple dies - features of the cache hierarchy and partitioning cores into work teams. The article also investigates the correlation between the performance optimizations and energy efficiency for a ccNUMA platform powered by top-of-the-line 64-core AMD Rome 7742 CPUs, comparing the results against two servers with Intel Xeon Scalable processors of different generations. Even without appealing to prices, the achieved performance and energy efficiency results are a solid argument confirming the competitiveness of AMD Rome processors against Intel Xeon CPUs in scientific applications. Lukasz Szustak, Roman Wyrzykowski, Lukasz Kuczynski, Tomasz Olas |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Performance enhancement of a dynamic K-means algorithm through a parallel adaptive strategy on multicore CPUs
Giuliano Laccetti, Marco Lapegna, Valeria Mele, Diego Romano, Lukasz Szustak |
J. Parallel Distributed Comput. | 5 |
| 2020 | Correlation of Performance Optimizations and Energy Consumption for Stencil-Based Application on Intel Xeon Scalable ProcessorsabstractThis article provides a comprehensive study of the impact of performance optimizations on the energy efficiency of a real-world CFD application called MPDATA, as well as an insightful analysis of performance-energy interaction of these optimizations with the underlying hardware that represents the first generation of Intel Xeon Scalable processors. Considering the MPDATA iterative application as a use case, we explore the fundamentals of energy and performance analysis for a memory-bound application when exposed to a set of optimization steps that increase the application performance, by improving the operational intensity of code and utilizing resources more efficiently. It is shown that for memory-bound applications, optimizing toward high performance could be a powerful strategy for improving the energy efficiency as well. In fact, for the considered performance optimizations, the energy gain is correlated with the performance gain but with varying degrees. As a result, these optimizations allow improving both performance and energy consumption radically, up to about 10.9 and 8.8 times, respectively. The impact of the Intel AVX-512 SIMD extension on the energy consumption and performance is demonstrated. Also, we discover limitations on the usability of CPU frequency scaling as a tool for balancing energy savings with admissible performance losses. Lukasz Szustak, Roman Wyrzykowski, Tomasz Olas, Valeria Mele |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Unleashing the performance of ccNUMA multiprocessor architectures in heterogeneous stencil computationsabstractThis paper meets the challenge of harnessing the heterogeneous communication architecture of ccNUMA multiprocessors for heterogeneous stencil computations, an important example of which is the Multidimensional Positive Definite Advection Transport Algorithm (MPDATA). We propose a method for optimization of parallel implementation of heterogeneous stencil computations that is a combination of the islands-of-core strategy and ( \(3{+}1\) )D decomposition. The method allows a flexible management of the trade-off between computation and communication costs in accordance with features of modern ccNUMA architectures. Its efficiency is demonstrated for the implementation of MPDATA on the SGI UV 2000 and UV 3000 servers, as well as for 2- and 4-socket ccNUMA platforms based on various Intel CPU architectures, including Skylake, Broadwell, and Haswell. Lukasz Szustak, Kamil Halbiniak, Roman Wyrzykowski, Ondrej Jakl |
J. Supercomput. | 1 |
| 2018 | Strategy for data-flow synchronizations in stencil parallel computations on multi-/manycore systemsabstractIn this paper, an innovative strategy for the data-flow synchronization in shared-memory systems is proposed. This strategy assumes to synchronize only interdependent threads instead of using the barrier approach that—in contrast to our approach—synchronize all threads. We demonstrate the adaptation of the data-flow synchronization strategy to two complex scientific applications based on stencil codes. An algorithm for the data-flow synchronization is developed and successfully used for both applications. The proposed approach is evaluated for various Intel microarchitectures released in the last 5 years, including the newest processors: Skylake and Knights Landing. The important part of this assessment is the performance comparison of the proposed data-flow synchronization with the OpenMP barrier. The experimental results show that the performance of the studied applications can be accelerated up to 1.3 times using the proposed data-flow synchronizations strategy. Lukasz Szustak |
J. Supercomput. | 1 |
| 2017 | Model-Based Optimization of EULAG Kernel on Intel Xeon Phi Through Load ImbalancingabstractLoad balancing is a widely accepted technique for performance optimization of scientific applications on parallel architectures. Indeed, balanced applications do not waste processor cycles on waiting at points of synchronization and data exchange, maximizing this way the utilization of processors. In this paper, we challenge the universality of the load-balancing approach to optimization of the performance of parallel applications. First, we formulate conditions that should be satisfied by the performance profile of an application in order for the application to achieve its best performance via load balancing. Then we use a real-life scientific application, EULAG MPDATA kernel, to demonstrate that its performance profile on a modern parallel architecture, Intel Xeon Phi, significantly deviates from these conditions. Based on this observation, we propose a method of performance optimization of scientific applications through load imbalancing. In the case of data parallel application, the method uses functional performance models of the application to find partitioning that minimizes its computation time but not necessarily balances the load of processors. We apply this method to optimization of MPDATA on Intel Xeon Phi. Experimental results demonstrate that the performance of this carefully optimized load-balanced application can be further improved by 15percent using the proposed load-imbalancing technique. Alexey L. Lastovetsky, Lukasz Szustak, Roman Wyrzykowski |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Adaptation of fluid model EULAG to graphics processing unit architectureabstractSummary The goal of this study is to adapt the multiscale fluid solver EULerian or LAGrangian framewrok (EULAG) to future graphics processing units (GPU) platforms. The EULAG model has the proven record of successful applications, and excellent efficiency and scalability on conventional supercomputer architectures. Currently, the model is being implemented as the new dynamical core of the COSMO weather prediction framework. Within this study, two main modules of EULAG, namely the multidimensional positive definite advection transport algorithm (MPDATA) and the variational generalized conjugate residual, elliptic pressure solver Generalized Conjugate Residual (GCR) are analyzed and optimized. In this paper, a method is proposed, which ensures a comprehensive analysis of the resource consumption including registers, shared, and global memories. This method allows us to identify bottlenecks of the algorithm, including data transfers between host and global memory, global and shared memories, as well as GPU occupancy. We put the emphasis on providing a fixed memory access pattern, padding as well as organizing computation in the MPDATA algorithm. The testing and validation of the new GPU implementation have been carried out based on modeling decaying turbulence of a homogeneous incompressible fluid in a triply‐periodic cube. Simulations performed using the standard version of EULAG and its new GPU implementation give similar solutions. Preliminary results show a promising increase in terms of computational efficiency. Copyright © 2014 John Wiley & Sons, Ltd. Krzysztof Rojek, Milosz Ciznicki, Bogdan Rosa, Piotr Kopta, Michal Kulczewski, Krzysztof Kurowski, Zbigniew Pawel Piotrowski, Lukasz Szustak, Damian Karol Wójcik, Roman Wyrzykowski |
Concurr. Comput. Pract. Exp. | 8 |
| 2014 | Parallelization of 2D MPDATA EULAG algorithm on hybrid architectures with GPU accelerators
Roman Wyrzykowski, Lukasz Szustak, Krzysztof Rojek |
Parallel Comput. | 2 |
| 2012 | Model-driven adaptation of double-precision matrix multiplication to the Cell processor architecture
Roman Wyrzykowski, Krzysztof Rojek, Lukasz Szustak |
Parallel Comput. | 3 |