Beata Bylina

dblp:04/1338 · DBLP profile ↗
← Back
19ranked-venue papers
15as first author
7since 2021 · last 2025
0000-0002-1327-9747ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 17 · 14 first-author · 5 since 2021Software engineering, systems software and programming languages · 17 · 14 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Teaching Parallel Programming on the CPU Based on Matrix Multiplication Using MKL, OpenMP and SYCL Libraries
Emilia Bober, Beata Bylina
CSEDU (2)2
2025 A Multithreaded Java-Based Video Encoder for Multicore Systems
abstract
The growing demand for efficient video compression solutions underscores the increasing significance of research into coding optimization on multi-core systems.This paper presents the implementation and performance analysis of a multithreaded video encoder developed in Java and optimized for modern multi-core processors.The encoder performs intra-frame compression of raw YUV 4:2:0 video using a frame-based parallelism approach.This method maximizes CPU core utilization while minimizing thread management overhead.Tests were conducted on five platforms equipped with multicore processors from Intel and AMD.The application's execution time was measured for varying numbers of frames.The obtained results demonstrate effective scalability of the encoder as the number of processed frames increases, confirming that Java can be an effective tool for implementing parallel video compression on multi-core systems.
Beata Bylina, Maciej Okon
FedCSIS1
2024 Fast slope algorithm with the use of vectorization and parallelization for multicore architectures
abstract
Abstract The slope calculation algorithm is one of the most widely used geospatial algorithms employing the 3x3 moving window technique (along with calculation of aspect, curvature and flow direction). This work presents an approach consisting of transforming a slope algorithm from a sequential form into a version that can exploit vector and parallel traits of multicore architectures with vector instructions. This approach allows us to take advantage of the potential of the modern multicore processors. The basic idea for optimizing the 3x3 moving window computation is to split the equation used to calculate the result into parts that operate on data that are known to exist in adjacent memory locations. The research was conducted on two multicore architectures without the change in the code — the older architecture was Sandy Bridge and the newer one was Haswell (with more cores). The efficiency of the developed slope algorithm was verified in practice with the use of DEM files of the same resolution but of different sizes. We showed through the numerical experiments that our approach gives better time performance than the original algorithm (and other tools) — and with no loss of accuracy.
Beata Bylina, Jaroslaw Bylina, Lukasz Chabudzinski, Karol Karpowicz, Michal Klisowski, Piotr Oleszczuk, Joanna Potiopa, Przemyslaw Stpiczynski
GeoInformatica1
2023 Impact of processor frequency scaling on performance and energy consumption for WZ factorization on multicore architecture
abstract
With the growing demand for computing power, new multicore architectures have emerged to provide better performance.Reducing their energy consumption is one of the main challenges in achieving high performance computing.Current research trends develop new software and hardware techniques to achieve the best performance and energy compromise.In this work, we investigate the effect of processor frequency scaling using Dynamic Voltage Frequency Scaling on performance and energy consumption for the WZ factorization.This factorization is implemented both without optimization techniques and with strip mining.This technique involves transforming the program loop to improve program performance.Based on time and energy tests, we have shown that for the WZ factorization algorithm, regardless of the presence of manual optimization, it pays to reduce the frequency to save energy without losing performance.The conclusion can be extended to analogous algorithmsalso having a high ratio of memory access to computational operations.
Beata Bylina, Jaroslaw Bylina, Monika Piekarz
FedCSIS1
2023 The scalability in terms of the time and the energy for several matrix factorizations on a multicore machine
abstract
Scalability is an important aspect related to time and energy savings on modern multicore architectures.In this paper, we investigate and analyze scalability in terms of time and energy.We compare the execution time and consumption energy of the LU factorization (without pivoting) and Cholesky, both with Math Kernel Library (MKL) on a multicore machine.In order to save the energy of these multithreaded factorizations, the dynamic voltage and frequency scaling (DVFS) technique was used.This technique allows the clock frequency to be scaled without changing the implementation.An experimental scalability evaluation was performed on an Intel Xeon Gold multicore machine, depending on the number of threads and the clock frequency.Our test results show that scalability in terms of the execution time expressed by the Speedup metric has values close to a linear function with an increase in the number of threads.In contrast, scalability in terms of the energy consumed expressed by the Greenup metric has values close to a logarithmic function with an increase in the number of threads.Both kinds of scalability depend on the clock frequency settings and the number of threads.
Beata Bylina, Monika Piekarz
FedCSIS1
2022 Influence of loop transformations on performance and energy consumption of the multithreded WZ factorization
abstract
High-level loop transformations are a key instrument to effectively exploit the resource in modern architectures.Energy consumption on multi-core architectures is one of the major issues connected with high-performance computing.We examine the impact of four loop transformation strategies on performance and energy consumption.The investigated strategies include: loop fission, loop interchange (permutation), stripmining, and loop tiling.Additionally, a column-wise and row-wise store formats for dense matrices are considered.Parallelization and vectorization are implemented using OpenMP directives.As a test, the WZ factorization algorithm is used.The comparison of selected strategies of the loop transformation is done for Intel architecture, namely Cascade Lake.It has been shown that for WZ factorization, which is an example of an application in which we can use the loop transformation, optimization towards highperformance can also be an effective strategy for improving energy efficiency.Our results show also that block size selection in loop tilling has a significant impact on energy consumption.
Beata Bylina, Jaroslaw Bylina, Monika Piekarz
FedCSIS1
2021 The impact of vectorization and parallelization of the slope algorithm on performance and energy efficiency on multi-core architecture
abstract
Calculation of land-surface parameters (e.g.slope, aspect, curvature) is an important part of many geospatial analyses.Current research trends are aimed at developing new software techniques to achieve the best performance and energy trade-off.In our work, we concentrate on the vectorization and parallelization to improve overall energy efficiency and performance of the neighborhood raster algorithms for the computation of land-surface parameters.We chose the slope calculation algorithm as the basis for our investigation.The parallelization was achieved through redesigning the the original sequential code with OpenMP SIMD vectorization hints for compiler, OpenMP loop parallelization, and the hybrid of these techniques.To evaluate both performance and energy savings, we tested our vector-parallel implementations on a multi-core computer for various data sizes.RAPL interface was used to measure energy consumption.The results showed that optimization towards high performance can also be an effective strategy for improving energy efficiency.
Beata Bylina, Joanna Potiopa, Michal Klisowski, Jaroslaw Bylina
FedCSIS1
2018 An effective sparse storage scheme for GPU-enabled uniformization method
abstract
The authors developed a GPU approach to the uniformization method for the computing transient solution of Markov models.The authors use two techniques to reduce the memory size of storing matrices.One of them is a modification of a storage sparse matrix format HYB; second is to utilize two GPU cards and the multicore CPU.The modified HYB format is suitable for sparse Markovian transition rate matrices and oversized matrices on single GPU, also improving computation performance at the same time.The use of two GPUs enables processing matrices of even bigger sizes.
Beata Bylina, Jaroslaw Bylina, Marek Karwacki
FedCSIS1
2017 OpenMP Thread Affinity for Matrix Factorization on Multicore Systems
abstract
The aim of this paper is to investigate the impact of thread affinity on computing performance for matrix factorization on shared memory multicore systems with hierarchical memory.We consider two parallel block matrix factorizations (LU and WZ) and employ thread affinity to improve their performance.We study decomposition without pivoting and we compare differences between various affinity strategies for diagonally dominant matrices.Our results show that the choice of thread affinity has the measurable impact on the performance of the matrice factorizations.
Beata Bylina, Jaroslaw Bylina
FedCSIS1
2016 Parallelizing nested loops on the Intel Xeon Phi on the example of the dense WZ factorization
abstract
In this article we evaluate some strategies of parallelizing nested loops on Intel Xeon Phi on the example of the WZ factorization for dense matrices.We employ both parallelism and vectorization to accelerate nested loops on manycore coprocessor.For random dense square matrices with the dominant diagonal we report the execution time and the performance of the nested loops.Numerical experiments show that the vectorization that is efficiently exploiting SIMD vector units do not always improve the application performance on the coprocessor.
Jaroslaw Bylina, Beata Bylina
FedCSIS2
2016 Data Structures for Markov Chain Transition Matrices on Intel Xeon Phi
abstract
We employ Intel Xeon Phi as a high-performance coprocessor to solve Markov chains.Matrices arising from Markov models are very sparse with short rows.In this paper, the authors research two storage formats of Markov chain transition matrices on Intel Xeon Phi.In this work CSR and HYB (modification ELL) formats for such matrices are studied.Numerical experiments results for transition matrices of Markov chains from wireless networks and call-center models show that HYB format in offload version is more effective than CSR format.The obtained performance for HYB format is even 1.45 times better in comparison to multi-threaded CPU (dual Intel Xeon E5-2670) with the use of the CSR format (SpMV from the MKL library on CPU).
Beata Bylina, Joanna Potiopa
FedCSIS1
2015 Strategies of parallelizing nested loops on the multicore architectures on the example of the WZ factorization for the dense matrices
abstract
In the WZ factorization the outermost parallel loop decreases the number of iterations executed at each step and this changes the amount of parallelism in each step.The aim of the paper is to present four strategies of parallelizing nested loops on multicore architectures on the example of the WZ factorization.For random dense square matrices with the dominant diagonal we report the execution time, the performance, the speedup of the WZ factorization for these four strategies of parallelizing nested loops and we investigate the accuracy of such solutions.It is possible to shorten the runtime when utlilizing the appropriate strategies with the use of good scheduling.
Beata Bylina, Jaroslaw Bylina
FedCSIS1
2014 Performance analysis of the WZ factorization in MATLAB
abstract
In the paper the authors present the WZ factorization in MATLAB.MATLAB is an environment for matrix computations, therefore in the paper there are presented both the sequential WZ factorization and a block-wise version of the WZ factorization (called here VWZ).Both the algorithms were implemented and their performance was investigated.For random dense square matrices with the dominant diagonal we report the execution time of the WZ factorization in MATLAB and we investigate the accuracy of such solutions.Additionally, the results (time and accuracy) for our WZ implementations were compared to the similar ones based on the LU factorization.
Beata Bylina, Jaroslaw Bylina
FedCSIS1
2014 Performance Analysis of Multicore and Multinodal Implementation of SpMV Operation
abstract
Abstract—In this paper we present two algorithms for perform-ing sparse matrix-dense vector multiplication (known as SpMV operation). We show parallel (multicore) version of algorithm, which can be efficiently implemented on the contemporary multicore architectures. Next, we show distributed (so-called multinodal) version targeted at high performance clusters. Both versions are thoroughly tested using different architectures, compiler tools and sparse matrices of different sizes. Considered matrices comes from The University of Florida Sparse Matrix Collection. The performance of the algorithms is compared to the performance of SpMV routine from widely used Intel Math Kernel Library.
Beata Bylina, Jaroslaw Bylina, Przemyslaw Stpiczynski, Dominik Szalkowski
FedCSIS1
2013 Mixed precision iterative refinement techniques for the WZ factorization
Beata Bylina, Jaroslaw Bylina
FedCSIS1
2012 GPU-accelerated WZ Factorization with the Use of the CUBLAS Library
Beata Bylina, Jaroslaw Bylina
FedCSIS1
2012 Multi-GPU Implementation of the Uniformization Method for Solving Markov Models
Marek Karwacki, Beata Bylina, Jaroslaw Bylina
FedCSIS2
2011 The incomplete factorization preconditioners applied to the GMRES(m) method for solving Markov chains
Beata Bylina, Jaroslaw Bylina
FedCSIS1
2011 The influence of a matrix condition number on iterative methods' convergence
Anna Pyzara, Beata Bylina, Jaroslaw Bylina
FedCSIS2