VLDB 2026 Research / reviewers in the wild / expert
Jaroslaw Bylina
dblp:03/3577
· DBLP profile ↗
18ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-0319-2525ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 16 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An environment model in multi-agent reinforcement learning with decentralized trainingabstractIn multi-agent reinforcement learning scenarios, independent learning, where agents learn independently based on their observations, is often preferred for its scalability and simplicity compared to centralized training.However, it faces significant challenges due to the non-stationary nature of the environment from each agent's perspective.We investigate if incorporating an environment model in multi-agent reinforcement learning with decentralized training can alleviate the non-stationarity effects caused by the adaptive behaviors of other agents.To do this, we design and implement a custom model-based algorithm and compare its performance with the well-known model-free algorithm (Deep Q-Network).Our algorithm uses an environment model to plan and select actions.However, we do not require the model to be perfect for action selection, allowing it to be learned and improved during training.Our results suggest that integrating environment models into MARL offers a viable solution to mitigate non-stationarity. Rafal Niedziólka-Domanski, Jaroslaw Bylina |
FedCSIS | 2 |
| 2024 | Fast slope algorithm with the use of vectorization and parallelization for multicore architecturesabstractAbstract The slope calculation algorithm is one of the most widely used geospatial algorithms employing the 3x3 moving window technique (along with calculation of aspect, curvature and flow direction). This work presents an approach consisting of transforming a slope algorithm from a sequential form into a version that can exploit vector and parallel traits of multicore architectures with vector instructions. This approach allows us to take advantage of the potential of the modern multicore processors. The basic idea for optimizing the 3x3 moving window computation is to split the equation used to calculate the result into parts that operate on data that are known to exist in adjacent memory locations. The research was conducted on two multicore architectures without the change in the code — the older architecture was Sandy Bridge and the newer one was Haswell (with more cores). The efficiency of the developed slope algorithm was verified in practice with the use of DEM files of the same resolution but of different sizes. We showed through the numerical experiments that our approach gives better time performance than the original algorithm (and other tools) — and with no loss of accuracy. Beata Bylina, Jaroslaw Bylina, Lukasz Chabudzinski, Karol Karpowicz, Michal Klisowski, Piotr Oleszczuk, Joanna Potiopa, Przemyslaw Stpiczynski |
GeoInformatica | 2 |
| 2023 | Impact of processor frequency scaling on performance and energy consumption for WZ factorization on multicore architectureabstractWith the growing demand for computing power, new multicore architectures have emerged to provide better performance.Reducing their energy consumption is one of the main challenges in achieving high performance computing.Current research trends develop new software and hardware techniques to achieve the best performance and energy compromise.In this work, we investigate the effect of processor frequency scaling using Dynamic Voltage Frequency Scaling on performance and energy consumption for the WZ factorization.This factorization is implemented both without optimization techniques and with strip mining.This technique involves transforming the program loop to improve program performance.Based on time and energy tests, we have shown that for the WZ factorization algorithm, regardless of the presence of manual optimization, it pays to reduce the frequency to save energy without losing performance.The conclusion can be extended to analogous algorithmsalso having a high ratio of memory access to computational operations. Beata Bylina, Jaroslaw Bylina, Monika Piekarz |
FedCSIS | 2 |
| 2022 | Influence of loop transformations on performance and energy consumption of the multithreded WZ factorizationabstractHigh-level loop transformations are a key instrument to effectively exploit the resource in modern architectures.Energy consumption on multi-core architectures is one of the major issues connected with high-performance computing.We examine the impact of four loop transformation strategies on performance and energy consumption.The investigated strategies include: loop fission, loop interchange (permutation), stripmining, and loop tiling.Additionally, a column-wise and row-wise store formats for dense matrices are considered.Parallelization and vectorization are implemented using OpenMP directives.As a test, the WZ factorization algorithm is used.The comparison of selected strategies of the loop transformation is done for Intel architecture, namely Cascade Lake.It has been shown that for WZ factorization, which is an example of an application in which we can use the loop transformation, optimization towards highperformance can also be an effective strategy for improving energy efficiency.Our results show also that block size selection in loop tilling has a significant impact on energy consumption. Beata Bylina, Jaroslaw Bylina, Monika Piekarz |
FedCSIS | 2 |
| 2021 | The impact of vectorization and parallelization of the slope algorithm on performance and energy efficiency on multi-core architectureabstractCalculation of land-surface parameters (e.g.slope, aspect, curvature) is an important part of many geospatial analyses.Current research trends are aimed at developing new software techniques to achieve the best performance and energy trade-off.In our work, we concentrate on the vectorization and parallelization to improve overall energy efficiency and performance of the neighborhood raster algorithms for the computation of land-surface parameters.We chose the slope calculation algorithm as the basis for our investigation.The parallelization was achieved through redesigning the the original sequential code with OpenMP SIMD vectorization hints for compiler, OpenMP loop parallelization, and the hybrid of these techniques.To evaluate both performance and energy savings, we tested our vector-parallel implementations on a multi-core computer for various data sizes.RAPL interface was used to measure energy consumption.The results showed that optimization towards high performance can also be an effective strategy for improving energy efficiency. Beata Bylina, Joanna Potiopa, Michal Klisowski, Jaroslaw Bylina |
FedCSIS | 4 |
| 2018 | An effective sparse storage scheme for GPU-enabled uniformization methodabstractThe authors developed a GPU approach to the uniformization method for the computing transient solution of Markov models.The authors use two techniques to reduce the memory size of storing matrices.One of them is a modification of a storage sparse matrix format HYB; second is to utilize two GPU cards and the multicore CPU.The modified HYB format is suitable for sparse Markovian transition rate matrices and oversized matrices on single GPU, also improving computation performance at the same time.The use of two GPUs enables processing matrices of even bigger sizes. Beata Bylina, Jaroslaw Bylina, Marek Karwacki |
FedCSIS | 2 |
| 2018 | Parallelization of stochastic bounds for Markov chains on multicore and manycore platformsabstractThe author demonstrates the methodology for parallelizing of finding stochastic bounds for Markov chains on multicore and manycore platforms. The stochastic bounds algorithm for Markov chains with the sparse matrices is investigated, thus needing a lot of irregular memory access. Its parallel implementations should scale across multiple threads and characterize with a high performance and performance portability between multicore and manycore platforms. The presented methods are built on the usage of two parallelization extensions of the C++ language: OpenMP and Cilk Plus. For this two extensions, we use two programming models, namely loop parallelism and task-based parallelism. The numerical experiments show the execution time of the implementations and the scalability on multicore and manycore platforms. This work provides the parallel implementations and at the same time presents an educational example of how computer science problems with irregular memory access can be implemented for high performance using OpenMP and Cilk Plus. Jaroslaw Bylina |
J. Supercomput. | 1 |
| 2017 | A Framework for Generating and Evaluating Parallelized CodeabstractThe work describes a flexible framework built to generate various (parallel) software versions and to benchmark them.The framework is written with the use of the Python language with some support of the gnuplot plotting program.An example of the use of this tool shows the tuning of a matrix factorization on different architectures (Intel Haswell and Intel Knights Corner) with various parameters of parallelization, vectorization, blocking etc. Jaroslaw Bylina |
FedCSIS | 1 |
| 2017 | OpenMP Thread Affinity for Matrix Factorization on Multicore SystemsabstractThe aim of this paper is to investigate the impact of thread affinity on computing performance for matrix factorization on shared memory multicore systems with hierarchical memory.We consider two parallel block matrix factorizations (LU and WZ) and employ thread affinity to improve their performance.We study decomposition without pivoting and we compare differences between various affinity strategies for diagonally dominant matrices.Our results show that the choice of thread affinity has the measurable impact on the performance of the matrice factorizations. Beata Bylina, Jaroslaw Bylina |
FedCSIS | 2 |
| 2016 | Parallelizing nested loops on the Intel Xeon Phi on the example of the dense WZ factorizationabstractIn this article we evaluate some strategies of parallelizing nested loops on Intel Xeon Phi on the example of the WZ factorization for dense matrices.We employ both parallelism and vectorization to accelerate nested loops on manycore coprocessor.For random dense square matrices with the dominant diagonal we report the execution time and the performance of the nested loops.Numerical experiments show that the vectorization that is efficiently exploiting SIMD vector units do not always improve the application performance on the coprocessor. Jaroslaw Bylina, Beata Bylina |
FedCSIS | 1 |
| 2015 | Strategies of parallelizing nested loops on the multicore architectures on the example of the WZ factorization for the dense matricesabstractIn the WZ factorization the outermost parallel loop decreases the number of iterations executed at each step and this changes the amount of parallelism in each step.The aim of the paper is to present four strategies of parallelizing nested loops on multicore architectures on the example of the WZ factorization.For random dense square matrices with the dominant diagonal we report the execution time, the performance, the speedup of the WZ factorization for these four strategies of parallelizing nested loops and we investigate the accuracy of such solutions.It is possible to shorten the runtime when utlilizing the appropriate strategies with the use of good scheduling. Beata Bylina, Jaroslaw Bylina |
FedCSIS | 2 |
| 2014 | Performance analysis of the WZ factorization in MATLABabstractIn the paper the authors present the WZ factorization in MATLAB.MATLAB is an environment for matrix computations, therefore in the paper there are presented both the sequential WZ factorization and a block-wise version of the WZ factorization (called here VWZ).Both the algorithms were implemented and their performance was investigated.For random dense square matrices with the dominant diagonal we report the execution time of the WZ factorization in MATLAB and we investigate the accuracy of such solutions.Additionally, the results (time and accuracy) for our WZ implementations were compared to the similar ones based on the LU factorization. Beata Bylina, Jaroslaw Bylina |
FedCSIS | 2 |
| 2014 | Performance Analysis of Multicore and Multinodal Implementation of SpMV OperationabstractAbstract—In this paper we present two algorithms for perform-ing sparse matrix-dense vector multiplication (known as SpMV operation). We show parallel (multicore) version of algorithm, which can be efficiently implemented on the contemporary multicore architectures. Next, we show distributed (so-called multinodal) version targeted at high performance clusters. Both versions are thoroughly tested using different architectures, compiler tools and sparse matrices of different sizes. Considered matrices comes from The University of Florida Sparse Matrix Collection. The performance of the algorithms is compared to the performance of SpMV routine from widely used Intel Math Kernel Library. Beata Bylina, Jaroslaw Bylina, Przemyslaw Stpiczynski, Dominik Szalkowski |
FedCSIS | 2 |
| 2013 | Mixed precision iterative refinement techniques for the WZ factorization
Beata Bylina, Jaroslaw Bylina |
FedCSIS | 2 |
| 2012 | GPU-accelerated WZ Factorization with the Use of the CUBLAS Library
Beata Bylina, Jaroslaw Bylina |
FedCSIS | 2 |
| 2012 | Multi-GPU Implementation of the Uniformization Method for Solving Markov Models
Marek Karwacki, Beata Bylina, Jaroslaw Bylina |
FedCSIS | 3 |
| 2011 | The incomplete factorization preconditioners applied to the GMRES(m) method for solving Markov chains
Beata Bylina, Jaroslaw Bylina |
FedCSIS | 2 |
| 2011 | The influence of a matrix condition number on iterative methods' convergence
Anna Pyzara, Beata Bylina, Jaroslaw Bylina |
FedCSIS | 3 |