Marcin Lawenda

dblp:l/MarcinLawenda · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-4844-3655ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Evaluating AMD EPYC CPU architectures on CFD applications
Marcin Lawenda, Lukasz Szustak, László Környei, Flavio Cesar Cunha Galeazzo, Pawel Bratek
Future Gener. Comput. Syst.1
2025 Profiling and Optimization of Multicard GPU Machine Learning Jobs
abstract
ABSTRACT The article discusses various model optimization techniques, providing a comprehensive analysis of key performance indicators. Several parallelization strategies for image recognition are analyzed, adapted to different hardware and software configurations, including distributed data parallelism and distributed hardware processing. Changing the tensor layout in PyTorch DataLoader from NCHW to NHWC and enabling pin _ memory has proven to be very beneficial and easy to implement. Furthermore, the impact of different performance techniques (DPO, LoRA, QLoRA, and QAT) on the tuning process of LLMs was investigated. LoRA allows for faster tuning, while requiring less VRAM compared to DPO. On the other hand, QAT is the most resource‐intensive method, with the slowest processing times. A significant portion of LLM tuning time is attributed to initializing new kernels and synchronizing multiple threads when memory operations are not dominant.
Marcin Lawenda, Kyrylo Khloponin, Krzesimir Samborski, Lukasz Szustak
Concurr. Comput. Pract. Exp.1
2025 Prediction model of performance-energy trade-off for CFD codes on AMD-based cluster
Marcin Lawenda, Lukasz Szustak, László Környei
Future Gener. Comput. Syst.1
2024 Large-Scale Parallelization of Human Migration Simulation
abstract
Forced displacement of people worldwide, for example, due to violent conflicts, is common in the modern world, and today more than 82 million people are forcibly displaced. This puts the problem of migration at the forefront of the most important problems of humanity. The Flee simulation code is an agent-based modeling tool that can forecast population displacements in civil war settings, but performing accurate simulations requires nonnegligible computational capacity. In this article, we present our approach to Flee parallelization for fast execution on multicore platforms, as well as discuss the computational complexity of the algorithm and its implementation. We benchmark parallelized code using supercomputers equipped with AMD EPYC Rome 7742 and Intel Xeon Platinum 8268 processors and investigate its performance across a range of alternative rule sets, different refinements in the spatial representation, and various numbers of agents representing displaced persons. We find that Flee scales excellently to up to 8192 cores for large cases, although very detailed location graphs can impose a large initialization time overhead.
Derek Groen, Nikela Papadopoulou, Petros Anastasiadis, Marcin Lawenda, Lukasz Szustak, Sergiy Gogolenko, Hamid Arabnejad, Alireza Jahani
IEEE Trans. Comput. Soc. Syst.4
2023 Profiling and optimization of Python-based social sciences applications on HPC systems by means of task and data parallelism
abstract
The article presents optimization techniques for two Python-based large-scale social sciences applications: SN (Social Network) Simulator and KPM (Kernel Polynomial Method). These applications use MPI technology to transfer data between computing processes, which in the regular implementation leads to load imbalance and performance degradation. To avoid this effect, we propose a 2-stage optimization. In the first step, the order of tasks is changed, and in the second step, the tasks are divided into smaller ones for easier allocation. In addition, we focus on mitigating performance and memory bottlenecks using modern ccNUMA systems with multiple NUMA domains. As part of the performance analysis, the limitations of communication in data traffic between and within the processor were revealed and resolved through appropriate data allocation. Benchmarking was carried out, examining various environments, including vendors of traditional x86-64 and ARM-based processors for HPC.
Lukasz Szustak, Marcin Lawenda, Sebastian Arming, Gregor Bankhamer, Christoph Schweimer, Robert Elsässer
Future Gener. Comput. Syst.2
2007 Multi-installment divisible load processing in heterogeneous distributed systems
abstract
Abstract Divisible loads are parallel applications with fine granularity and negligible data dependencies. Such computations can be divided into parts of arbitrary sizes and processed independently in parallel. The load distribution process incurs considerable communication delays. To reduce processor waiting time during the computation initialization phase, the load is distributed in multiple small installments rather than in one big chunk. In this paper we analyze multi‐installment divisible load processing in heterogeneous distributed systems. Scheduling divisible loads in heterogeneous systems is hard because the sizes of the installments should be adjusted to the communication and computation capabilities of the system. We show that ignoring heterogeneity of the distributed system may result in arbitrarily bad solutions. Two algorithms are proposed to gear the load chunk sizes to different communication and computation speeds: an optimization branch‐and‐bound algorithm and a heuristic based on a genetic search method. The running times of both methods and the quality of the solutions are compared. Then, we use these algorithms to study the features of the multi‐installment divisible load scheduling problem. We demonstrate that it has both combinatorial and algebraic nature, and that optimum solutions are harder to find with the growing heterogeneity of the system. Copyright © 2007 John Wiley & Sons, Ltd.
Maciej Drozdowski, Marcin Lawenda
Concurr. Comput. Pract. Exp.2
2006 Virtual Laboratory as a Remote and Interactive Access to the Scientific Instrumentation Embedded in Grid Environment
abstract
The following paper describes the key architecture elements of the Virtual Laboratory Project, which is developed in the Poznan Supercomputing and Networking Center. Project started in 2002, and since then the work was carried under different statesponsored grants and now is carried under EU FP6 IST projects. The authors will explain briefly the main ideas and key design elements behind this original concept. The current state of the project will be presented, as the project reached its production phase, and plans for the near and far future will be revealed. A special attention will be given to the concept and implementation of Digital Science Library, as it plays an important role in the scope of the whole system. Although the various aspects of this system were presented on workshops and conferences, this paper brings together all the pieces, and presents the overall look on the entire scope of this Virtual Laboratory System.
Marcin Okon, Damian Kaliszan, Marcin Lawenda, Dominik Stoklosa, Tomasz Rajtar, Norbert Meyer, Maciej Stroinski
e-Science3
2005 On Optimum Multi-installment Divisible Load Processing in Heterogeneous Distributed Systems
Maciej Drozdowski, Marcin Lawenda
Euro-Par2