EDBT 2026 Demo / reviewers in the wild / expert
Roman Wyrzykowski
dblp:72/1812
· DBLP profile ↗
25ranked-venue papers
12as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 12 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advances in algorithms, models, hardware, and software for next-generations HPC systems, volume 2
Roman Wyrzykowski, Ewa Deelman |
Future Gener. Comput. Syst. | 1 |
| 2025 | Special issue on advances in techniques for assessment performance portability of HPC applicationsabstractThis special issue aims to present new developments and advances in techniques for assessment performance portability of high performance computing applications. It contains revised and extended versions of selected papers presented at the 10th Workshop on Language-Based Parallel Programming Models, WLPP 2024, which was a part of 15th International Conference on Parallel Processing and Applied Mathematics, PPAM 2024, held on September 8–11, 2024, in Ostrava, Czech Republic. Ami Marowka, Przemyslaw Stpiczynski, Roman Wyrzykowski |
Future Gener. Comput. Syst. | 3 |
| 2024 | Advances into exascale computingabstractThe landscape of high-performance computing (HPC) has been expanding with new technologies and increased system complexity. For hardware, this trend is driven by technological inventions increasing computing power capabilities while taming cost metrics. For applications, we are witnessing increasing growth in algorithms' complexity to accommodate constantly expanding data sizes and take full advantage of increasingly complex processing and storage characteristics. Successfully handling this expansion and maintaining performance and efficiency on adequate levels requires software that matches the emerging hardware and system innovations and can address concerns arising from evolving and new paradigms. An example of the evolving paradigms is heterogeneity-one of the most challenging properties of emerging parallel, distributed, and edge computing platforms. An appealing example of the new paradigms is a vigorous expansion of artificial intelligence (AI) and machine learning (ML) methods that have become pervasive in solving the most demanding problems across many science and engineering disciplines. Also, the approaching end of Moore's Law scaling requires exciting but complex innovations and advances in HPC hardware/software environments that need to be matched by advances in algorithms and software systems to meet the demands of the applications. Roman Wyrzykowski, Boleslaw K. Szymanski |
Concurr. Comput. Pract. Exp. | 1 |
| 2024 | Preface of special issue on advances in algorithms, models, hardware, and software for next-generations HPC systems
Roman Wyrzykowski, Ewa Deelman |
Future Gener. Comput. Syst. | 1 |
| 2023 | Reducing energy consumption using heterogeneous voltage frequency scaling of data-parallel applications for multicore systems
Pawel Bratek, Lukasz Szustak, Roman Wyrzykowski, Tomasz Olas |
J. Parallel Distributed Comput. | 3 |
| 2022 | Algorithmic and software development advances for next-generation heterogeneous platformsabstractHeterogeneity is emerging as one of the most profound and challenging characteristics of today's and tomorrow's parallel and distributed computing environments, presenting new and exciting opportunities for their development. Most modern computing systems are heterogeneous, either for organic reasons because components grew independently, as is the case of desktop grids, by design to leverage the strength of specific hardware, as is the case of accelerated systems, or both. The impact of heterogeneity on all forms of parallel and distributed computing is increasing rapidly. Traditional algorithms, programming environments, and tools designed for legacy homogeneous systems will at best achieve a small fraction of the efficiency and the potential performance expected from parallel computing in tomorrow's highly diversified and mixed architectures. Innovative ideas, fresh models, novel algorithms, and other specialized or unified programming environments and tools are needed to efficiently use these new and increasingly diverse computing systems—for accelerating scientific discovery and impactful innovation. The International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar) has been the premier forum over the last 20 years, bringing together researchers to discuss these challenges and the solutions. The wide range of topics includes achieving performance portability on heterogeneous architectures, advances in software environments that facilitate efficient use of heterogeneous systems, performance and energy optimization of numerical and machine learning algorithms on heterogeneous platforms, to name a few. The works presented at the HeteroPar'2020 workshop covered topics clearly exhibiting the significance and growth of the heterogeneous computing field. However, one general trend is apparent: the broad adoption of Graphics Processing Units (GPU) accelerators. Over the last decade, GPUs have been established as the main powerhouse in leadership supercomputers and an invaluable component to accelerate computations for a vast spectrum of applications—from numerical linear algebra libraries powering computational science to various machine learning workloads. This trend is evidenced by the increasing number of GPU-related publications submitted to HeterPar and supported by growing diversity within the GPU world, where AMD accelerator architectures start to compete with Nvidia's comprehensive solutions, along with the third GPU accelerator option—from Intel—available soon. This special issue of Concurrency and Computation: Practice and Experience contains six selected papers from the HeteroPar'2020 workshop. We hope you find them interesting and stimulating new ideas and future advancements for next-generation heterogeneous platforms. The 18th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2020) was held in Warsaw, Poland, on 25 August 2020. For the 11th time, this workshop was organized in conjunction with the Euro-Par annual series of international conferences. Because of the COVID-19 pandemic, HeteroPar'2020 was held as a virtual event. Sixteen articles were submitted for review, with authors from eight countries. Each paper secured at least three reviews from members of the program committee, whereas 12 submissions received at least four reviews. After a thorough peer-reviewing process that included discussion and agreement among reviewers whenever necessary, nine articles were selected for presentation at the workshop. The review process focused on the quality of the papers, their innovative ideas and applicability to heterogeneous computing. The topics addressed in the accepted papers include domain-specific languages for numerical algorithms, virtualization for CUDA applications, unified memory in CUDA, porting CUDA codes to AMD GPUs, management of heterogeneous cloud resources, GPU implementation of graph neural networks, GPU and CPU signal processing for a wildlife tracking system, parallelization of the k-means algorithm on CPU-GPU platforms, and a portable solver for systems of linear equations. An integral part of the workshop was two keynote talks given by Enrique S. Quintana-Orti (Technical University of Valencia, Spain) and Tal Ben-Nun (ETHZ Zurich, Switzerland) devoted to using approximate and transprecision computing in sparse linear solvers, and data-centric approach for performance portability on heterogeneous architecture, respectively. After the workshop, the program committee invited the authors of the presented works to submit revised and extended versions of their contributions as part of the papers submitted to this special issue. These new versions were reviewed independently again by at least three reviewers. Finally, six papers were accepted for publication in the special issue. They are summarized below. Aliaga et al.1 focus on optimizing the sparse matrix–vector product (SpMV), which dictates, to a large extent, the performance of a considerable variety of scientific applications. The proposed approach introduces a variant of the coordinate sparse matrix format that allows combining load-balancing with compressing both the indexing arrays and the numerical information to reduce the pressure on memory while using the available compute power of modern CPUs and GPUs efficiently. This approach is multi-platform, in the sense that the realizations are built upon common principles but differ in the implementation details, which are adapted either to avoid thread divergence in the GPU case or to maximize compression for multicore architectures. The evaluation on the two last generations of NVIDIA GPUs as well as Intel and AMD processors demonstrate the benefits of the new kernels compared with the optimized implementations of SpMV in Nvidia's cuSPARSE and Intel's MKL libraries. k-Means is a standard algorithm for clustering data used as the final step for high-quality spectral clustering. To overcome the scalability challenge when processing large datasets, the authors of paper2 propose to apply also the k-means algorithm as a preprocessing task to reduce the input data instances. Additionally, parallel optimization techniques are introduced to improve the efficiency of the k-means algorithm on CPU and GPU. Notably, a two-step summation method with package processing is used to handle the effect of rounding errors that may occur during the phase of updating cluster centroids. The extensive experiments on synthetic and real-world datasets containing millions of instances exhibit a speedup up to 7 for the k-means iteration time on GPU versus 20/40 CPU threads using AVX units while achieving double-precision accuracy with single-precision computations. Dmitruk et al.3 show how the OpenACC standard can be efficiently used to implement solvers for tridiagonal Toeplitz systems of linear equations for a variety of modern GPU-accelerated and multicore architectures. Two parallel algorithms are studied concerning particular assumptions about coefficient matrices. In the first case, a new, faster implementation of the divide and conquer method is proposed, while in the second one, a novel, vectorizable algorithm is introduced. Using both column-wise and row-wise matrix storage formats is studied, along with efficient conversion between them using cache memory to improve the overall performance. It is also shown how to tune the performance by predicting the best values of the methods' parameters. Numerical experiments performed on Intel CPUs and Nvidia GPUs confirm the excellent performance and accuracy of the developed implementations. Robust high-performance implementations of signal-processing tasks performed by a high-throughput wildlife tracking system are presented by Rubinpur et al.4 The system tracks radio transmitters attached to wild animals by estimating the time of arrival of radio packets to multiple receivers. The time-consuming estimation of wideband radio signals is a bottleneck that limits the system's throughput. A sequential high-performance CPU implementation has been developed first, and then a GPU implementation to overcome this bottleneck. The authors carefully evaluate the performance of these real-world codes. The evaluation indicates that the GPU version dramatically improves both performance and power-performance efficiency relative to a desktop CPU—a scenario typical for current base stations. Performance improves by more than 50 times on a high-end GPU and more than four times with a GPU platform that consumes almost five times less power than the CPU one. The desire to take advantage of virtualization in heterogeneous computing resources with GPU accelerators motivates Eiling et al.5 Currently, GPUs do not offer virtualization support that enables fine-grained control, increased flexibility, and fault tolerance. The authors present Cricket—a transparent and low-overhead solution to GPU virtualization that enables future research of various virtualization techniques, due to its open-source nature. Cricket supports remote execution and checkpoint/restart of CUDA applications. Both features allow the distribution of GPU tasks dynamically and flexibly across computing nodes and the multitenant usage of GPU resources, improving their flexibility and utilization in high-performance and cloud computing. Solving partial differential equations (PDEs) on unstructured grids is a cornerstone of engineering and scientific computing. Alhaddad et al.6 introduce the HighPerMeshes C++-embedded domain-specific language (DSL) that bridges the abstraction gap between the mathematical formulation of mesh-based algorithms for PDE problems and an increasing number of heterogeneous platforms with their various programming models. The HighPerMeshes DSL aims at higher productivity of the code development for multiple target platforms. For this aim, the OpenCL is used as a backend, targeting various GPUs and other heterogeneous architectures such as FPGAs. Apart from describing the basic structure of the DSL, its usage is demonstrated with three examples. The mapping of the abstract algorithmic description onto parallel hardware, including compute clusters, is also presented. Finally, the achievable performance and scalability are demonstrated for different example problems. The guest editors of this special issue wish to thank the authors of the submitted papers, the reviewers for the careful evaluation of the papers, and the valuable suggestions that helped the authors to improve their contributions. In addition, we would like to sincerely thank Prof. David W. Walker (Editor-in-Chief of Concurrency and Computation: Practice and Experience) for the opportunity to guest edit this special issue and for his guidance during this process. Data sharing is not applicable to this article as no datasets were generated or analyzed in this study. Roman Wyrzykowski, Florina M. Ciorba |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Exploration of OpenCL Heterogeneous Programming for Porting Solidification Modeling to CPU-GPU PlatformsabstractSummary This article provides a comprehensive study of OpenCL heterogeneous programming for porting applications to CPU–GPU computing platforms, with a real‐life application for the solidification modeling. The aim is to achieve a flexible workload distribution between available CPU–GPU resources and optimize application performance. Considering the solidification application as a use case, we explore the necessary steps required for (i) adaptation of an application to CPU–GPU platforms, and (ii) mapping the application workload onto the OpenCL programming model. The adaptation is based on a reformulation of steps developed previously for CPU–MIC architectures. The mapping process allows us to utilize OpenCL for harnessing CPU and GPU cores using data parallelism, as well as for the management of available compute devices with task parallelism. The resulting OpenCL code's performance and energy efficiency is experimentally studied for two platforms with powerful GPUs of various generations (with Kepler and Volta architectures). The experiments confirm the performance advantage of using computing resources of both GPUs and CPUs. The achieved benefit depends on the relationship between the computing power of CPUs and GPUs. Moreover, this gain entails the growth of the average power that increases the energy consumed during the application execution. Kamil Halbiniak, Lukasz Szustak, Tomasz Olas, Roman Wyrzykowski, Pawel Gepner |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Taming next-generation HPC systems: Run-time system and algorithmic advancementsabstractPPAM is a biennial series of international conferences dedicated to exchanging ideas between researchers involved in parallel and distributed computing, including theory and applications, as well as applied and computational mathematics.Twelve previous events have been held in different universities in Poland since 1994, when the first PPAM took place in Czestochowa.Thus, the event in Bialystok was an opportunity to celebrate the 25th anniversary of PPAM.The focus of PPAM 2019 was on models, algorithms, and software tools that facilitate efficient and convenient use of modern parallel and distributed computing systems, as well as on large-scale modern applications, including advances in machine learning and artificial intelligence.This meeting gathered more than 170 participants from 26 countries.The accepted papers were presented at the regular tracks of the PPAM 2019 conference and during the workshops.With each submission evaluated by at least three reviewers, a strict reviewing process resulted in the acceptance of 91 contributed papers for publication in the conference proceedings, while approximately 43% of the submissions were rejected.The Program Committee selected 41 papers for presentation in the regular conference track, resulting in an acceptance rate of about 46%.Based on the review results, 10 papers (11% of submissions) were selected for a special journal issue.Besides quality, another important criterion for selection was each paper's contribution to the thematic consistency of the issue.The focus of this special issue is on algorithmic advancements in matching the software properties to parallel architecture, including GPU accelerators and clusters.These advancements are crucial for successfully parallelizing such complex applications as simulating geophysical flows, solving ordinary differential equations (ODEs), structural analysis of nuclear reactor containment buildings, solving generalized eigenvalue problems, modeling of material science phenomena, and others.A complementary topic of this issue is advances in run-time systems since increasing levels of parallelism in multi-and many-core chips and the emerging heterogeneity of computational resources coupled with energy, resilience, and data movement constraints radically increase the importance of efficient run-time scheduling and execution control.After the conference, the Program Committee invited the authors of selected papers to submit revised and extended versions of their works.These new versions were reviewed independently again by at least three reviewers.Finally, nine contributions were accepted for publication.They are summarized below.Paper [1] focuses on the accurate assembly of the system matrix, which is an essential step in any code that solves partial differential equations on a mesh.This step can become costly in multigrid codes requiring cascades of matrices that depend upon each other, or dynamic adaptive mesh refinement.To reduce the time to solution, the authors propose that these constructions can be performed concurrently with the multigrid cycles.Furthermore, they desynchronize the assembly from the solution process.This non-trivial increase in the concurrency level improves the scalability.As assembly routines are notoriously memory-and bandwidth-demanding, the final algorithmic enhancement uses a hierarchical, lossy compression scheme that brings the memory footprint down aggressively even when the system matrix entries carry little information or are not yet available with high accuracy.An efficient algorithm for the parallel solution of indefinite saddle point systems with iterative solvers based on the Golub-Kahan bidiagonalization is presented in Reference [2].Such systems arise in many application fields, for example, in structural mechanics.A scalability study of the generalized solver shows improved performance for the two-dimensional (2D) Stokes equations compared to previous works.Furthermore, the authors investigate the performance of different parallel inner solvers in the outer Golub-Kahan iteration for a three-dimensional (3D) Stokes problem.When the number of cores is increasing for a fixed problem size, the solver exhibits good speedups of up to 50% with the 1024 cores.For the tests in which the problem size grows while the workload in each core stays constant, the performance of the solver scales almost linearly with the increase in the number of cores. Paper [3] proposes a locality optimization technique for the parallel solution on GPUs of large systems of ODEs by explicit one-step methods.This technique is based on tiling across the stages of a one-step method and is enabled by a special structure of the class of ODE systems-with the limited access distance.The paper focuses on increasing the range of access distances for which the tiling technique can provide a speedup Roman Wyrzykowski, Boleslaw K. Szymanski |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Architectural Adaptation and Performance-Energy Optimization for CFD Application on AMD EPYC RomeabstractThe advantages of the second-generation AMD EPYC Rome processors can be successfully used in the race to Exascale. However, the novel architecture's complexity makes it challenging to adapt demanding scientific codes - like stencil ones - to platforms with Rome CPUs. This article tackles this challenge by exploring the adaptation of the stencil-based CFD (computational fluid dynamics) application called MPDATA to these processors' influential features. We show that the previously proposed parametric adaptation methodology can be profitably applied to extend the performance portability of the memory-bound MPDATA on the AMD EPYC architecture. The extension of the parametric adaptation on the novel architecture requires careful consideration of two relevant aspects that reflect splitting the Rome architecture into multiple dies - features of the cache hierarchy and partitioning cores into work teams. The article also investigates the correlation between the performance optimizations and energy efficiency for a ccNUMA platform powered by top-of-the-line 64-core AMD Rome 7742 CPUs, comparing the results against two servers with Intel Xeon Scalable processors of different generations. Even without appealing to prices, the achieved performance and energy efficiency results are a solid argument confirming the competitiveness of AMD Rome processors against Intel Xeon CPUs in scientific applications. Lukasz Szustak, Roman Wyrzykowski, Lukasz Kuczynski, Tomasz Olas |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Correlation of Performance Optimizations and Energy Consumption for Stencil-Based Application on Intel Xeon Scalable ProcessorsabstractThis article provides a comprehensive study of the impact of performance optimizations on the energy efficiency of a real-world CFD application called MPDATA, as well as an insightful analysis of performance-energy interaction of these optimizations with the underlying hardware that represents the first generation of Intel Xeon Scalable processors. Considering the MPDATA iterative application as a use case, we explore the fundamentals of energy and performance analysis for a memory-bound application when exposed to a set of optimization steps that increase the application performance, by improving the operational intensity of code and utilizing resources more efficiently. It is shown that for memory-bound applications, optimizing toward high performance could be a powerful strategy for improving the energy efficiency as well. In fact, for the considered performance optimizations, the energy gain is correlated with the performance gain but with varying degrees. As a result, these optimizations allow improving both performance and energy consumption radically, up to about 10.9 and 8.8 times, respectively. The impact of the Intel AVX-512 SIMD extension on the energy consumption and performance is demonstrated. Also, we discover limitations on the usability of CPU frequency scaling as a tool for balancing energy savings with admissible performance losses. Lukasz Szustak, Roman Wyrzykowski, Tomasz Olas, Valeria Mele |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Algorithmic advances in parallel architectures and energy-efficient computingabstractin Lublin, Poland.PPAM is a biennial series of international conferences dedicated to exchanging ideas between researchers involved in parallel and distributed computing, including theory and applications, as well as applied and computational mathematics.The focus of PPAM 2017 was on models, algorithms, and software tools that facilitate efficient and convenient use of modern parallel and distributed computing systems, as well as on large-scale applications, including data-intensive and machine learning problems. Roman Wyrzykowski, Boleslaw K. Szymanski |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Unleashing the performance of ccNUMA multiprocessor architectures in heterogeneous stencil computationsabstractThis paper meets the challenge of harnessing the heterogeneous communication architecture of ccNUMA multiprocessors for heterogeneous stencil computations, an important example of which is the Multidimensional Positive Definite Advection Transport Algorithm (MPDATA). We propose a method for optimization of parallel implementation of heterogeneous stencil computations that is a combination of the islands-of-core strategy and ( \(3{+}1\) )D decomposition. The method allows a flexible management of the trade-off between computation and communication costs in accordance with features of modern ccNUMA architectures. Its efficiency is demonstrated for the implementation of MPDATA on the SGI UV 2000 and UV 3000 servers, as well as for 2- and 4-socket ccNUMA platforms based on various Intel CPU architectures, including Skylake, Broadwell, and Haswell. Lukasz Szustak, Kamil Halbiniak, Roman Wyrzykowski, Ondrej Jakl |
J. Supercomput. | 3 |
| 2017 | Energy-aware mechanism for stencil-based MPDATA algorithm with constraintsabstractSummary In this paper, we propose an energy‐aware task management mechanism designed for the forward‐in‐time algorithms running on multicore central processing units (CPUs), where the multidimensional positive definite advection transport algorithm stencil‐based algorithm is one of the representative examples. This mechanism is based on the dynamic voltage and frequency scaling technique and allows the reduction of energy consumption for an existing algorithm (or application) such that the predefined execution time is respected, without requiring any modifications in the algorithm itself. This paper also provides the formulation of a method for minimizing the energy consumption with time constraints, which is based on the adaptive scheduling with online modeling. Finally, using the autotuning technique, we provide the automation of the process for creation and determination of the best energy profile at runtime, even in the presence of additional CPU workloads. The experimental results on a 6‐core computing platform show that the proposed mechanism provides the energy savings of up to 1.43x when compared to the default Linux scaling governor. Also, we confirm the effectiveness of the self‐adaptive feature of the proposed mechanism, by showing its ability to maintain the requested execution time in spite of additional CPU workloads imposed by other applications. Krzysztof Rojek, Aleksandar Ilic, Roman Wyrzykowski, Leonel Sousa |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Systematic adaptation of stencil-based 3D MPDATA to GPU architecturesabstractSummary In this work, we focus on a systematic adaptation of the stencil‐based multidimensional positive definite advection transport algorithm (MPDATA) to different graphics processing unit (GPU)‐based computing platforms. Another objective of this work is to compare the performance of MPDATA on several platforms, including a multi‐GPU system with two NVIDIA Tesla K80 cards, and single‐card platforms with Tesla K20X, GeForce GTX TITAN, and GeForce GTX 980. The usage of the following optimization methods is proposed to improve the overall performance: (i) reducing the number of operations by the subexpression elimination when implementing 2.5D blocking; (ii) reorganization of boundary conditions for reducing branch instructions; (iii) advanced memory management to increase the coalesced memory access; and (iv) warps rearrangement for optimizing the data access to GPU global memory. The presented methods of the MPDATA adaptation to GPU architectures allow us to efficiently use many graphics processors within a single node by applying peer‐to‐peer data transfers between GPU global memories. We propose an auto‐tuning procedure to compensate architectural differences between the considered platforms. This procedure takes into account algorithm/GPU‐specific parameters. The proposed approach to adaptation of MPDATA to GPU architectures allows us to achieve up to 482.5 Gflop/s for the platform equipped with two NVIDIA K80 GPUs. Copyright © 2016 John Wiley & Sons, Ltd. Krzysztof Rojek, Roman Wyrzykowski, Lukasz Kuczynski |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Algorithmic advances for parallel architecturesabstractAlgorithmic advances for parallel architecturesThis special issue of Concurrency and Computation: Practice and Experience contains revised and extended versions of selected papers presented at the 11th International Conference on Parallel Processing and Applied Mathematics, PPAM 2015, which was held on September 6 to 9, 2015, in Krakow, Poland.PPAM is a biennial series of international conferences dedicated to exchanging ideas between researchers involved in parallel and distributed computing, including theory and applications, as well as applied and computational mathematics.The focus of PPAM 2015 was on models, algorithms, and software tools that facilitate efficient and convenient use of modern parallel and distributed computing systems, as well as on large-scale applications, including data-intensive problems.PPAM 2015 was organized by the Department of Computer and Information Science of the Czestochowa University of Technology together with the AGH University of Science and Technology, under the patronage of the Committee of Informatics of the Polish Academy of Sciences, in cooperation with the ICT COST Action IC1305 "Network for Sustainable Ultrascale Computing (NESUS)."This meeting gathered more than 190 participants from 33 countries.A strict reviewing process, with each submission reviewed at least three times, resulted in acceptance of 111 contributed papers with the acceptance rate of 57%.The accepted papers were presented at the regular tracks of the PPAM 2015 conference, as well as during the workshops that were important and integral parts of the PPAM 2015 meeting.Based on the results of the reviews, selected papers were recommended for a special journal issue.Besides quality, another important goal that influenced the paper selection was a maximum possible thematic consistency of the issue.The focus of this special issue is on algorithmic advances to better match the software properties to the targeted parallel architecture.These advances vary from general, like increasing potential reuse of a part of computation, to specific for a particular architecture, eg, using GPUs or with multi-core and many-core processors, or specific applications, such as spherical Delaunay triangulations, or visualization of complex networks.The authors were contacted after the conference and invited to submit revised and extended versions of their papers.These new versions were reviewed independently by three reviewers.Finally, nine contributions were accepted for publication.They are summarized below.Paper [1] focuses on LU factorization, which is a recurrent operation in scientific and engineering applications used to solve linear systems.In particular, the authors adapt the Gauss-Huard algorithm that has been introduced as an efficient solution for modern platforms equipped with accelerators to improve reusing of computations in this algorithm.This approach was evaluated on the solution of Lyapunov matrix equations via the LRCF-ADI method and validated on three benchmarks.In future work, the authors plan to design a heuristic for choosing the optimal block size and to integrate their solution with the mixed precision techniques to further accelerate the Lyapunov solver.An efficient algorithm for solving dense symmetric indefinite systems on GPUs is presented in the work of Baboulin et al. [2].The critical challenge on these types of computations on hybrid CPU/GPU is to keep to minimum the expensive data transfer and synchronization between the CPU and GPU, or within the GPU.The algorithmic advances here included selecting the solver using iterative refinements and the factorization without pivoting.This approach was then combined with the preprocessing technique based on Random Butterfly Transformations and with the mixed-precision algorithm.The authors demonstrated that an efficient solution can be obtained by avoiding the pivoting and using the lower precision arithmetic.The validation was made using an application in acoustics studied in this paper. Roman Wyrzykowski, Boleslaw K. Szymanski |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Modeling power consumption of 3D MPDATA and the CG method on ARM and Intel multicore architecturesabstractWe propose an approach to estimate the power consumption of algorithms, as a function of the frequency and number of cores, using only a very reduced set of real power measures. In addition, we also provide the formulation of a method to select the voltage–frequency scaling–concurrency throttling configurations that should be tested in order to obtain accurate estimations of the power dissipation. The power models and selection methodology are verified using two real scientific application: the stencil-based 3D MPDATA algorithm and the conjugate gradient (CG) method for sparse linear systems. MPDATA is a crucial component of the EULAG model, which is widely used in weather forecast simulations. The CG algorithm is the keystone for iterative solution of sparse symmetric positive definite linear systems via Krylov subspace methods. The reliability of the method is confirmed for a variety of ARM and Intel architectures, where the estimated results correspond to the real measured values with the average error being slightly below 5% in all cases. Krzysztof Rojek, Enrique S. Quintana-Ortí, Roman Wyrzykowski |
J. Supercomput. | 3 |
| 2017 | Performance modeling of 3D MPDATA simulations on GPU clusterabstractThe goal of this study is to parallelize the multidimensional positive definite advection transport algorithm (MPDATA) across a computational cluster equipped with GPUs. Our approach permits us to provide an extensive overlapping GPU computations and data transfers, both between computational nodes, as well as between the GPU accelerator and CPU host within a node. For this aim, we decompose a computational domain into two unequal parts which correspond to either data dependent or data independent parts. Then, data transfers can be performed simultaneously with computations corresponding to the second part. Our approach allows for achieving 16.372 Tflop/s using 136 GPUs. To estimate the scalability of the proposed approach, a performance model dedicated to MPDATA simulations is developed. We focus on the analysis of computation and communication execution times, as well as the influence of overlapping data transfers and GPU computations, with regard to the number of nodes. Krzysztof Rojek, Roman Wyrzykowski |
J. Supercomput. | 2 |
| 2017 | Model-Based Optimization of EULAG Kernel on Intel Xeon Phi Through Load ImbalancingabstractLoad balancing is a widely accepted technique for performance optimization of scientific applications on parallel architectures. Indeed, balanced applications do not waste processor cycles on waiting at points of synchronization and data exchange, maximizing this way the utilization of processors. In this paper, we challenge the universality of the load-balancing approach to optimization of the performance of parallel applications. First, we formulate conditions that should be satisfied by the performance profile of an application in order for the application to achieve its best performance via load balancing. Then we use a real-life scientific application, EULAG MPDATA kernel, to demonstrate that its performance profile on a modern parallel architecture, Intel Xeon Phi, significantly deviates from these conditions. Based on this observation, we propose a method of performance optimization of scientific applications through load imbalancing. In the case of data parallel application, the method uses functional performance models of the application to find partitioning that minimizes its computation time but not necessarily balances the load of processors. We apply this method to optimization of MPDATA on Intel Xeon Phi. Experimental results demonstrate that the performance of this carefully optimized load-balanced application can be further improved by 15percent using the proposed load-imbalancing technique. Alexey L. Lastovetsky, Lukasz Szustak, Roman Wyrzykowski |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Adaptation of fluid model EULAG to graphics processing unit architectureabstractSummary The goal of this study is to adapt the multiscale fluid solver EULerian or LAGrangian framewrok (EULAG) to future graphics processing units (GPU) platforms. The EULAG model has the proven record of successful applications, and excellent efficiency and scalability on conventional supercomputer architectures. Currently, the model is being implemented as the new dynamical core of the COSMO weather prediction framework. Within this study, two main modules of EULAG, namely the multidimensional positive definite advection transport algorithm (MPDATA) and the variational generalized conjugate residual, elliptic pressure solver Generalized Conjugate Residual (GCR) are analyzed and optimized. In this paper, a method is proposed, which ensures a comprehensive analysis of the resource consumption including registers, shared, and global memories. This method allows us to identify bottlenecks of the algorithm, including data transfers between host and global memory, global and shared memories, as well as GPU occupancy. We put the emphasis on providing a fixed memory access pattern, padding as well as organizing computation in the MPDATA algorithm. The testing and validation of the new GPU implementation have been carried out based on modeling decaying turbulence of a homogeneous incompressible fluid in a triply‐periodic cube. Simulations performed using the standard version of EULAG and its new GPU implementation give similar solutions. Preliminary results show a promising increase in terms of computational efficiency. Copyright © 2014 John Wiley & Sons, Ltd. Krzysztof Rojek, Milosz Ciznicki, Bogdan Rosa, Piotr Kopta, Michal Kulczewski, Krzysztof Kurowski, Zbigniew Pawel Piotrowski, Lukasz Szustak, Damian Karol Wójcik, Roman Wyrzykowski |
Concurr. Comput. Pract. Exp. | 10 |
| 2015 | 10th international conference on Parallel Processing and Applied Mathematics, PPAM 2013abstractThis special issue of Concurrency and Computation: Practice and Experience contains revised and extended versions of selected papers presented at the 10th International Conference on Parallel Processing and Applied Mathematics, PPAM 2013, which was held on September 8–11, 2013 in Warsaw, Poland. PPAM is a biennial series of international conferences dedicated to exchanging ideas between researchers involved in parallel and distributed computing, including theory and applications, as well as applied and computational mathematics. The focus of PPAM 2013 was on models, algorithms and software tools that facilitate efficient and convenient use of modern parallel and distributed computing systems, as well as on large-scale applications. PPAM 2013, the jubilee PPAM conference, was organized by the Department of Computer and Information Science of the Czestochowa University of Technology, under the patronage of the Committee of Informatics of the Polish Academy of Sciences, in cooperation with the Polish–Japanese Institute of Information Technology. This meeting gathered the largest number of participants in the history of PPAM conferences—more than 230 participants from 32 countries. A strict reviewing process, with each submission reviewed at least three times, resulted in acceptance of 143 contributed papers with the acceptance rate of 56%. The accepted papers were presented at the regular tracks of the PPAM 2013 conference, as well as during the workshops, which were important and integral components of PPAM meetings. Based on the results of the reviews, selected papers were recommended for a special journal issue. Besides quality, another important goal which influenced the paper selection was a maximum possible thematic consistency of the issue. The authors were contacted after the conference and invited to submit revised and extended versions of their papers. These new versions were reviewed independently by three reviewers. Finally, ten contributions were accepted for publication. They are summarized below. Paper 1 analyzes interaction occurring in the triangle: performance-power-energy for the execution of a pivotal numerical algorithm, the iterative Conjugate Gradient (CG) method, on a diverse collection of parallel multithreaded architectures. They range from general-purpose and digital signal multicore processors to GPUs. This analysis is especially timely in the decade where the power wall has arisen as a major obstacle to build faster processors. An alternative approach to the solution of a sequence of correlated eigenproblems is proposed in paper 2. The resulting eigensolver is optimized regarding the number of matrix–vector multiplications and parallelized for distributed memory architectures using the Elemental library framework. Numerical results show that the proposed solver achieves excellent scalability and is competitive with current dense linear algebra parallel eigensolvers. Paper 3 focuses on an efficient implementation of the multidimensional Monte-Carlo integration on various distributed-memory parallel computers and clusters of multi-/manycore nodes. In particular, it addresses the issue how to use multiple cores of CPUs together with multiple GPU accelerators within a single node, achieving a reasonable load balancing of available computational resources. The adaptation of the multi-scale fluid solver EULAG to modern GPU platforms is addressed in paper 4, which proposes a method ensuring a comprehensive analysis of the resource consumption, including registers, shared and global memories. This method allows for identifying bottlenecks of the underlying algorithms. In consequence, the two main modules of EULAG have been redesigned, and the new organization of computations has been implemented. The experimental results demonstrate a promising increase in terms of computational efficiency. Paper 5 presents GSWABE, a GPU-accelerated pairwise sequence alignment algorithm for a collection of short DNA sequences. The performance of GSWABE has been evaluated on a Kepler-based Tesla K40 GPU using a variety of datasets. In particular, compared to the CUDA-based gpu-pairAlign software, GSWABE demonstrate stable and consistent speedups with maximum values of 11.2, 10.7 and 10.6 for global, semi-global and local alignments, respectively. Paper 6 investigates using fused CPU-GPU systems for an efficient implementation of two population-based meta-heuristic algorithms: multi-swarm particle swarm optimization (MPSO) and genetic algorithm (GA). The paper develops a hybrid parallel algorithm that combines a slower convergent algorithm (GA) with a faster one (MPSO). The resulting algorithms are implemented on the AMD a8-3530MX APU that packs four x86 CPU cores and 80 GPU processing elements, providing effective use of the hierarchical memory structure, four-way vectorization and zero-copy buffers. Paper 7 studies how to exploit the computational power of future Extreme-Scale computers using the example of the Semi-Lagrangian code GYSELA which performs large simulations with up to 65-k CPU cores. Among the Exascale challenges, the less memory per core is one of the most critical issues. This paper develops a general method to understand the memory behavior of an application when dealing with very large meshes. Paper 8 presents a new algorithm for the problem of scheduling moldable tasks with precedence constraints assuming the makespan objective and arbitrary speedup functions. It is shown through simulation that the proposed algorithm not only creates competitive schedules for arbitrary speedup functions, but also outperforms other published heuristics and approximation algorithms for non-decreasing speedup functions. This study on parallel task scheduling is especially timely due to the increasing number of cores in current parallel machines, and the growing need for the concurrent execution of tasks. A novel distributed program design framework PEGASUS DA is described in paper 9. This framework supports designing the application program execution control based on evolved monitoring of distributed program global states. The paper presents how the infrastructure provided in the framework can be applied for an automated construction of strongly consistent application global states and control predicates, which are used as the basis for the distributed program execution control. The framework is graphically supported, and is oriented on program design for clusters of multicore processors with multithreading and message passing communication. The use of PEGASUS DA is illustrated using the example of the Travelling Salesman Problem solved by a branch and bound method. Paper 10 proposes a novel Partially Diagonal Network-on-Chip (PDNOC) design that takes advantages of both heterogeneous network topology and congestion-aware application mapping. This network is analyzed in terms of interconnection structure, silicon area usage, power consumption, routing algorithms and implementation complexity. Evaluation results show that on average the proposed PDNOC designs provide up to 36% improvement in execution time over concentrated mesh, and 3.6 times better energy delay product over fully connected diagonal network. The guest editors of this special issue wish to thank the reviewers for the careful reviewing of the papers, and useful suggestions that helped the authors to improve their contributions. Also, we would like to sincerely thank Professor Geoffrey Fox (Editor-in-Chief of the Concurrency and Computation: Practice and Experience) for opportunity to edit this special issue and his guidance. Roman Wyrzykowski, Marek S. Tudruj |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Parallelization of 2D MPDATA EULAG algorithm on hybrid architectures with GPU accelerators
Roman Wyrzykowski, Lukasz Szustak, Krzysztof Rojek |
Parallel Comput. | 1 |
| 2012 | Model-driven adaptation of double-precision matrix multiplication to the Cell processor architecture
Roman Wyrzykowski, Krzysztof Rojek, Lukasz Szustak |
Parallel Comput. | 1 |
| 1998 | Fault Tolerant QR-Decomposition Algorithm and Its Parallel Implementation
Oleg Maslennikov, Juri Kaniewski, Roman Wyrzykowski |
Euro-Par | 3 |
| 1997 | A Technique for Mapping Sparse Matrix Computations into Regular Processor Arrays
Roman Wyrzykowski, Juri Kanevski |
Euro-Par | 1 |
| 1992 | Dependence graph transformations in the design of processor arrays for matrix multiplications
Roman Wyrzykowski, Juri Kanevski, Sergej Ovramenko |
Microprocess. Microprogramming | 1 |