EDBT 2026 Demo / reviewers in the wild / expert
Przemyslaw Stpiczynski
dblp:s/PStpiczynski
· DBLP profile ↗
19ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-8661-414XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Special issue on advances in techniques for assessment performance portability of HPC applicationsabstractThis special issue aims to present new developments and advances in techniques for assessment performance portability of high performance computing applications. It contains revised and extended versions of selected papers presented at the 10th Workshop on Language-Based Parallel Programming Models, WLPP 2024, which was a part of 15th International Conference on Parallel Processing and Applied Mathematics, PPAM 2024, held on September 8–11, 2024, in Ostrava, Czech Republic. Ami Marowka, Przemyslaw Stpiczynski, Roman Wyrzykowski |
Future Gener. Comput. Syst. | 2 |
| 2025 | Performance portability of sparse matrix-vector multiplication implemented using OpenMP, OpenACC and SYCLabstractThe aim of this paper is to study the performance portability of OpenMP, OpenACC and SYCL implementations of sparse matrix–vector product (SpMV) and its extended version in which the dot product of the input vector and the result is also calculated, for CSR and BSR storage formats, on Intel and AMD CPUs and NVIDIA GPU platforms. We compare it with the performance portability of much more sophisticated implementations provided by the vendors in their Intel oneAPI MKL and NVIDIA cuSPARSE libraries. Using the reformulated performance portability metric we show how it changes for various sparse matrices and which portable implementation and format achieve better performance portability. Numerical experiments show that the considered portable implementations for the CSR format usually achieve better performance than for the BSR format. On GPU , CSR OpenACC implementations for SpMV and SpMV-DOT tend to be the best. On CPU, CSR OpenMP implementation usually gives the best results for SpMV-DOT, while CSR OpenMP and BSR MKL achieve the best results for a similar number of matrices. Kinga Stec, Przemyslaw Stpiczynski |
Future Gener. Comput. Syst. | 2 |
| 2024 | Fast slope algorithm with the use of vectorization and parallelization for multicore architecturesabstractAbstract The slope calculation algorithm is one of the most widely used geospatial algorithms employing the 3x3 moving window technique (along with calculation of aspect, curvature and flow direction). This work presents an approach consisting of transforming a slope algorithm from a sequential form into a version that can exploit vector and parallel traits of multicore architectures with vector instructions. This approach allows us to take advantage of the potential of the modern multicore processors. The basic idea for optimizing the 3x3 moving window computation is to split the equation used to calculate the result into parts that operate on data that are known to exist in adjacent memory locations. The research was conducted on two multicore architectures without the change in the code — the older architecture was Sandy Bridge and the newer one was Haswell (with more cores). The efficiency of the developed slope algorithm was verified in practice with the use of DEM files of the same resolution but of different sizes. We showed through the numerical experiments that our approach gives better time performance than the original algorithm (and other tools) — and with no loss of accuracy. Beata Bylina, Jaroslaw Bylina, Lukasz Chabudzinski, Karol Karpowicz, Michal Klisowski, Piotr Oleszczuk, Joanna Potiopa, Przemyslaw Stpiczynski |
GeoInformatica | 8 |
| 2023 | Performance of Portable Sparse Matrix-Vector Product Implemented Using OpenACCabstractThe aim of this paper is to study the performance of OpenACC implementations of sparse matrix-vector product for several storage formats: CSR, ELL, JAD, pJAD, and BSR, achieved on Intel CPU and NVIDIA GPU platforms to compare them with the performance of SpMV implementations using the BSR storage format provided by Intel MKL and NVIDIA cuSPARSE libraries.Numerical experiments show that vendorprovided BSR is the best format for CPUs but in the case of GPUs, the pJAD storage format allows to achieve better performance. Kinga Stec, Przemyslaw Stpiczynski |
FedCSIS | 2 |
| 2023 | Improving accuracy of summation using parallel vectorized Kahan's and Gill-Møller algorithmsabstractAbstract The aim of this paper is to show that Kahan's and Gill‐Møller compensated summation algorithms that allow to achieve high accuracy of summing long sequences of floating‐point numbers can be efficiently vectorized and parallelized using Intel AVX‐512 intrinsics together with OpenMP constructs in order to utilize SIMD extension of modern multicore processors. Numerical experiments show that the new implementations of the algorithms achieve much better accuracy than ordinary summation in both double and single precision and their performance is comparable with the performance of the ordinary summation algorithm optimized automatically. The vectorized Gill‐Møller algorithm is faster than the vectorized Kahan's algorithm. However, in case of single precision, the accuracy of the Gill‐Møller algorithm is worse than Kahan's but it can be fixed by the use of mixed‐precision. Then the accuracy of both compensated summation algorithms is the same and the Gill‐Møller algorithm is still faster than Kahan's. Beata Dmitruk, Przemyslaw Stpiczynski |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Editorial on Advances in High Performance Programming
Ami Marowka, Przemyslaw Stpiczynski |
Parallel Comput. | 2 |
| 2022 | Solving tridiagonal Toeplitz systems of linear equations on GPU-accelerated computersabstractAbstract The aim of this article is to show that solvers for tridiagonal Toeplitz systems of linear equations can be efficiently implemented for a variety of modern GPU‐accelerated and multicore architectures using OpenACC. We consider two parallel algorithms for solving such systems with special assumptions about coefficient matrices. As the first algorithm, we propose a new, faster implementation of the divide and conquer method. The next algorithm is a new, vectorizable algorithm based on a recently introduced sequential method. We consider the use of both column‐wise and row‐wise storage formats for two‐dimensional arrays and show how to efficiently convert between these two formats using cache memory and improve the overall performance of our implementations. We also show how to tune the performance by predicting the best values of the methods' parameters. Numerical experiments performed on Intel CPUs and Nvidia GPUs show that our new implementations achieve relatively good performance and accuracy. Beata Dmitruk, Przemyslaw Stpiczynski |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Editorial on the special issue on advances in parallel programming: Languages, models and algorithms
Ami Marowka, Przemyslaw Stpiczynski |
J. Parallel Distributed Comput. | 2 |
| 2020 | Algorithmic and language-based optimization of Marsa-LFIB4 pseudorandom number generator using OpenMP, OpenACC and CUDAabstractThe aim of this paper is to present new high-performance implementations of Marsa-LFIB4 which is an example of high-quality multiple recursive pseudorandom number generators. We propose an algorithmic approach that combines language-based vectorization techniques together with a new divide-and-conquer parallel method that exploits a special sparse structure of the matrix obtained from the recursive formula that defines the generator. Our portable OpenACC implementation achieves the performance comparable to those achieved by our CUDA-based and OpenMP-based implementations on GPUs and multicore CPUs, respectively. Przemyslaw Stpiczynski |
J. Parallel Distributed Comput. | 1 |
| 2018 | Special section on parallel programmingabstractWLPP 2017 was a two-day workshop focusing on high-level programming for large-scale parallel systems and multicore processors, with special emphasis on component architectures and models.Its goal was to bring together researchers working in the areas of applications, computational models, language design, compilers, system architecture, and programming tools to discuss new developments in programming Clouds and parallel systems.Papers in this section cover the most important topics presented during the workshop.The first two deal with parallel programming models.Thoman et al. [4] provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms and discuss its usefulness by classifying state-of-the-art task-based environments that are used today.Posner and Fohry [7] propose a hybrid work stealing scheme, which combines the lifeline-based variant of distributed task pools with the node-internal load balancing implemented as an extension of the APGAS library for Java. Ami Marowka, Przemyslaw Stpiczynski |
J. Supercomput. | 2 |
| 2018 | Vectorized algorithm for multidimensional Monte Carlo integration on modern GPU, CPU and MIC architecturesabstractThe aim of this paper is to show that the multidimensional Monte Carlo integration can be efficiently implemented on computers with modern multicore CPUs and manycore accelerators including Intel MIC and GPU architectures using a new vectorized version of LCG pseudorandom number generator which requires limited amount of memory. We introduce two new implementations of the algorithm based on directive-based parallel programming standards OpenMP and OpenACC and consider their performance using Hockney–Jesshope theoretical model of vector computations. We also present and discuss the results of experiments performed on dual-processor Intel Xeon E5-2670 computers with Intel Xeon Phi 7120P and NVIDIA K40m. Przemyslaw Stpiczynski |
J. Supercomput. | 1 |
| 2018 | Language-based vectorization and parallelization using intrinsics, OpenMP, TBB and Cilk PlusabstractThe aim of this paper is to evaluate OpenMP, TBB and Cilk Plus as basic language-based tools for simple and efficient parallelization of recursively defined computational problems and other problems that need both task and data parallelization techniques. We show how to use these models of parallel programming to transform a source code of Adaptive Simpson’s Integration to programs that can utilize multiple cores of modern processors. Using the example of Belman–Ford algorithm for solving single-source shortest path problems, we advise how to improve performance of data parallel algorithms by tuning data structures for better utilization of vector extensions of modern processors. Manual vectorization techniques based on Cilk array notation and intrinsics are presented. We also show how to simplify such optimization using Intel SIMD Data Layout Template containers. Przemyslaw Stpiczynski |
J. Supercomput. | 1 |
| 2015 | Using distributed memory parallel computers and GPU clusters for multidimensional Monte Carlo integrationabstractSummary The aim of this paper is to show that the multidimensional Monte Carlo integration can be efficiently implemented on various distributed memory parallel computers and clusters of multicore nodes using recently developed parallel versions of linear congruential generator and lagged Fibonacci generator pseudorandom number generators. We show how to accelerate the overall performance by offloading some computations to Graphics Processing Units (GPUs), and we discuss how to transform Message Passing Interface (MPI) + OpenMP programs to MPI + OpenMP + CUDA model. We explain how to utilize multiple cores of CPUs together with multiple GPU accelerators within a single node and how to achieve reasonable load balancing of all computational resources of GPU‐accelerated multicore nodes. We present and discuss the results of experiments performed on the following target architectures: IBM Blue Gene/Q parallel computer, a cluster of Intel Xeon E5‐2660 servers, and a Tesla‐based GPU cluster with Intel Xeon X5650 multicore processors. The results are presented from two points of view: strong scaling and weak scaling. We also compare the performance of all considered architectures. Copyright © 2014 John Wiley & Sons, Ltd. Dominik Szalkowski, Przemyslaw Stpiczynski |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Performance Analysis of Multicore and Multinodal Implementation of SpMV OperationabstractAbstract—In this paper we present two algorithms for perform-ing sparse matrix-dense vector multiplication (known as SpMV operation). We show parallel (multicore) version of algorithm, which can be efficiently implemented on the contemporary multicore architectures. Next, we show distributed (so-called multinodal) version targeted at high performance clusters. Both versions are thoroughly tested using different architectures, compiler tools and sparse matrices of different sizes. Considered matrices comes from The University of Florida Sparse Matrix Collection. The performance of the algorithms is compared to the performance of SpMV routine from widely used Intel Math Kernel Library. Beata Bylina, Jaroslaw Bylina, Przemyslaw Stpiczynski, Dominik Szalkowski |
FedCSIS | 3 |
| 2013 | Template Library for Multi-GPU Pseudorandom Number Generation
Dominik Szalkowski, Przemyslaw Stpiczynski |
FedCSIS | 2 |
| 2012 | Parallel GPU-accelerated recursion-based generators of pseudorandom numbers
Przemyslaw Stpiczynski, Dominik Szalkowski, Joanna Potiopa |
FedCSIS | 1 |
| 2011 | Solving Linear Recurrences on Hybrid GPU Accelerated Manycore Systems
Przemyslaw Stpiczynski |
FedCSIS | 1 |
| 1993 | Error Analysis of Two Parallel Algorithms for Solving Linear Recurrence Systems
Przemyslaw Stpiczynski |
Parallel Comput. | 1 |
| 1992 | Parallel Cholesky factorization on orthogonal multiprocessors
Przemyslaw Stpiczynski |
Parallel Comput. | 1 |