Mateusz Gruzewski

dblp:276/8172 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-9419-2749ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 NPDP programming for RISC multi-core processors
abstract
In recent years, parallel architectures have become ubiquitous due to advancements in AI and cloud computing.However, parallel processing is not limited to x86 CISC CPUs and advanced graphics cards, GPUs; it also includes computations on ARM-based and RISC-V devices.In recent years, ARM processors have been adapted to incorporate an increasing number of cores.Today, mobile devices feature at least eight execution units, typically divided into energy-efficient and performance-oriented groups.RISC-based parallel processors are also integrated on development boards supported by the Linux kernel.In this article, we tested our NPDP Benchmark Suite for non-serial polyadic dynamic programming, primarily in the field of computer algorithms and bioinformatics, to evaluate the performance of the RISC processors under study, as well as code locality and cache efficiency.The benchmark consists of 10 kernels written in C++ and OpenMP.In the Android environment, we used the JAVA NDK (Native Development Kit) to port the application.For Apple machines, we used a port to OpenMP for parallelization.For the RISC-V native Linux environment, we applied the native Linux setup for efficient execution.Finally, we summarized the article and outlined future work.
Marek Palkowski, Mateusz Gruzewski
FedCSIS2
2025 Automated Transformation of OpenMP to CUDA Kernels Using AI Models
abstract
The increasing demand for computational efficiency in high-performance computing (HPC) has driven research into automating the transformation of parallel programming paradigms. This paper investigates an AI-driven approach to translating OpenMP-based CPU parallel programs into CUDA-based GPU programs. Using omniCUDA, a custom fine-tuned large language model (LLM), functional CUDA kernels can be generated directly from OpenMP code, bypassing the need for traditional compiler optimization techniques. The training dataset consists of synthetic OpenMP-to-CUDA pairs and a selected subset of manually optimized algorithms from the PolyBench suite. Performance was evaluated on kernels not included in the training set, with only partial overlap, allowing me to assess the model’s ability to generalize to unseen algorithms. Experimental results confirm that the model produces syntactically correct and compilable CUDA code, successfully replicating functional behavior across parallel loop structures. Performance evaluation on four benchmark algorithms, three of which were not included in the training dataset, shows that the model consistently outperforms OpenMP implementations and, in some cases, surpasses even manually optimized CUDA kernels from the PolyBench suite. The presented approach demonstrates the feasibility and competitiveness of AI-assisted OpenMP-to-CUDA transformation. The model exhibits generalization capabilities beyond the training set, and ongoing work focuses on refining memory access strategies and kernel configurations to further enhance performance across diverse parallel workloads.
Mateusz Gruzewski
KES1
2025 Knowledge Extraction for RNA Secondary Structure Prediction Using Heterogeneous and Dynamic Programming
abstract
RNA folding can be compared to the process of knowledge extraction, as it involves extracting hidden information from the RNA nucleotide sequence to predict its 3D structure. Just like in knowledge extraction from large datasets, RNA folding uses various algorithms and mathematical models to analyze input data (RNA sequence) and uncover how the RNA molecule adopts its final form. Thus, RNA folding aims to understand how genetic information encoded in the nucleotide sequence translates into a functional molecular structure. RNA folding algorithms focus on constructing a large matrix using dynamic programming, which accounts for the majority of the computations. The Nussinov-like algorithms for RNA folding are fundamental non-serial polyadic dynamic programming codes (NPDP) for testing non-uniform dependency analysis, multi-core CPUs and GPUs, loop tiling transformations, and source-to-source techniques, as well as manual approaches. In this article, we first analyze the achievements in optimizing this code over the past two decades, discussing fully automated methods based on loop tiling and the polyhedral model, manual methods involving transposition, row- and square blocking, as well as the Four Russians technique, its variations, and hybrids that reduce the algorithm’s complexity. Second, we propose a novel method that extends the polyhedral model through dependency analysis, combines features of loop tiling and manual techniques, and can be applied to both GPUs and CPUs. Moreover, the adopted approach allows for optimizing other similar bioinformatics codes. Our research is based on multi-core processors and the OpenMP and OpenCL standards to demonstrate usability across multiple platforms. We will also discuss its simplicity and performance compared to other reference models. Finally, we conduct an experimental study on six multi-core machines, including both CPUs and GPUs, to demonstrate that the proposed approach outperforms related methods.
Mateusz Gruzewski, Marek Palkowski
KES1
2025 Cross-platform and polyhedral programming for Nussinov RNA folding
abstract
This article addresses the use of codes from polyhedral compilers with tiled and parallel code designed for CPU processors, automatically generated as source-to-source OpenMP for NVIDIA GPU graphics cards using CUDA. In previous publications, we demonstrated that it is possible to use large language models (LLM) to translate code, generate kernels, and correctly manage memory transfers between the host and the device without manual effort. Unfortunately, when the target architecture is not taken into account in detail, the performance of code designed for CPUs leaves much to be desired when running on GPUs. The architectural differences between these two platforms like cores, cache, and the dimensionality of computations require careful attention to performance portability . In this article, we address the Nussinov algorithm, a popular benchmark in bioinformatics, to achieve higher performance on the NVIDIA platform than automatically generated codes by LLM. Nussinov’s loop nests are a non-trivial kernel from the non-serial polyadic dynamic programming (NPDP) benchmark with non-uniform loops. We will utilize a polyhedral code framework that tiles and then manually modifies the most nested loop nest containing the majority of the computations, using the two-dimensional thread blocks. To accelerate the computations, shared memory within blocks is utilized. The resulting codes were tested on two modern NVIDIA devices for various RNA sequence lengths , compared to parallel and tiled CPU codes, and previously generated Nussinov’s GPU codes using LLMs. The correctness of these codes and their scalability were analyzed. Comparison to related approaches and future work are outlined.
Mateusz Gruzewski, Marek Palkowski
Future Gener. Comput. Syst.1
2024 Automatic Generation of OpenCL Code through Polyhedral Compilation with LLM
abstract
In recent years, a multitude of AI solutions has emerged to facilitate code generation, commonly known as Language Model-based Programming (LLM).These tools empower programmers to automate their work.Automatic programming also falls within the domain of optimizing compilers, primarily based on the polyhedral model, which processes loop nests concentrating most computations.This article focuses on harnessing LLM tools to generate OpenCL code for non-serial polyadic dynamic programming kernels.[1] We have chosen the Nussinov RNA folding computational task, previously employed to test polyhedral compilers in optimizing kernels with non-uniform dependences.The code generated in OpenMP by polyhedral optimizers is limited to CPU computations.We automatically convert it into the OpenCL standard using ChatGPT-3.5 through its source-to-source queries to extend the number of possible platforms.The validity and efficiency of the generated code were verified on various CPUs and GPUs from different manufacturers.
Marek Palkowski, Mateusz Gruzewski
FedCSIS2