Lubomir Riha

dblp:50/7508 · also Lubomír Ríha · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-1017-5766ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 GPU acceleration of hybrid FETI solver for problems of transient nonlinear dynamics
Jakub Homola, Ondrej Meca, Lubomir Riha, Tomás Brzobohatý
Future Gener. Comput. Syst.3
2026 Out-of-core aware multi-GPU rendering for large-scale scene visualization
Milan Jaros, Lubomir Riha, Petr Strakos, Tomás Kozubek
Future Gener. Comput. Syst.2
2026 An in-depth study of GPU frequency-scaling latency and its optimization on modern architectures
Daniel Velicka, Ondrej Vysocky, Osman Yasal, Lubomir Riha
Future Gener. Comput. Syst.4
2026 Adaptive neural seizure detection through intelligent EEG channel selection and dynamic feature recognition
Muhammad Zeerak Awan, Adil Jhangeer, Petr Strakos, Lubomir Riha
Neurocomputing4
2025 High Performance Visualization for Astrophysics and Cosmology
abstract
Modern Astrophysics and Cosmology (A&C) projects produce immense data volumes, necessitating advanced software tools for data access, storage, and analysis. Visualization Interface for the Virtual Observatory (VisIVO) is one such tool enabling multi-dimensional data analysis and knowledge discovery across complex astrophysical datasets. Leveraging containerization and virtualization, VisIVO has been deployed on various distributed computing platforms. Additionally, Blender, an open-source 3D suite, provides robust tools for rendering and processing volumetric data, making it suitable for visualizing complex datasets. At the SPACE Center of Excellence these tools are being adapted for high-performance visualization of cosmological simulations performed with GADGET and ChaNGa on pre-exascale systems. However, implementing high-performance visualization on diverse HPC platforms presents several challenges, including hardware and software compatibility, data management, scalability, performance portability, and efficient resource allocation. This paper outlines strategies to integrate VisIVO with workflow frameworks and streaming platforms to address these challenges. Workflow frameworks enhance portability, scheduling, and reproducibility of visualization workflows on pre-exascale systems used in A&C simulations. We also discuss the use of streaming platforms to enable concurrent (i.e. in-situ) analysis and visualization of simulations, reducing the need to store full simulation data by leveraging distributed databases that stream the output data in real time. Lastly, we present an adaptation of Blender to handle large-scale particle-based astrophysical data, offering high-quality visualization with interactive exploration capabilities.
Nicola Tuccari, Eva Sciacca, Fabio Vitello, Iacopo Colonnelli, Yolanda Becerra 0001, Enric Sosa Cintero, Guillermo Marin, Milan Jaros, Lubomir Riha, Petr Strakos, Sebastian Trujillo-Gomez, Emiliano Tramontana, Robert Wissing
PDP9
2025 Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
abstract
Schur complement matrices emerge in many domain decomposition methods that can utilize supercomputers to solve complex engineering problems. As most of today’s high-performance clusters’ performance lies in GPUs, these methods should also be accelerated.
Jakub Homola, Ondrej Meca, Lubomir Riha, Tomás Brzobohatý
SC3
2023 Scalable Flow Simulations with the Lattice Boltzmann Method
abstract
The primary goal of the EuroHPC JU project SCALABLE is to develop an industrial Lattice Boltzmann Method (LBM)-based computational fluid dynamics (CFD) solver capable of exploiting current and future extreme scale architectures, expanding current capabilities of existing industrial LBM solvers by at least two orders of magnitude in terms of processor cores and lattice cells, while preserving its accessibility from both the end-user and software developer's point of view. This is accomplished by transferring technology and knowledge between an academic code (waLBerla) and an industrial code (LaBS). This paper briefly introduces the characteristics and main features of both software packages involved in the process. We also highlight some of the performance achievements in scales of up to tens of thousand of cores presented on one academic and one industrial benchmark case.
Markus Holzer 0005, Gabriel Staffelbach, Ilan Rocchi, Jayesh Badwaik, Andreas Herten, Radim Vavrík, Ondrej Vysocky, Lubomir Riha, Romain Cuidard, Ulrich Rüde
CF8
2021 Application instrumentation for performance analysis and tuning with focus on energy efficiency
abstract
Summary Profiling and tuning of parallel applications is an essential part of HPC. Analysis and elimination of application hot spots can be performed using many available tools, which also provides resource consumption measurements for instrumented parts of the code. Since complex applications show different behavior in each part of the code, it is essential to be able to insert instrumentation to analyse these parts. Because each performance analysis or autotuning tool can bring different insights into an application behavior, it is valuable to analyze and optimize an application using a variety of them. We present our on request inserted shared C/C++ API for the most common open‐source HPC performance analysis tools, which simplify the process of the manual instrumentation. Besides manual instrumentation, profiling libraries provide different methods for instrumentation. Of these, the binary patching is the most universal mechanism, and highly improves the user‐friendliness and robustness of the tool. We provide an overview of the most commonly used binary patching tools, and describe a workflow for how to use them to implement a binary instrumentation tool for any profiler or autotuner. We have also evaluated the minimum overhead of the manual and binary instrumentation.
Ondrej Vysocky, Lubomir Riha, Andrea Bartolini
Concurr. Comput. Pract. Exp.2
2021 GPU Accelerated Path Tracing of Massive Scenes
abstract
This article presents a solution to path tracing of massive scenes on multiple GPUs. Our approach analyzes the memory access pattern of a path tracer and defines how the scene data should be distributed across up to 16 GPUs with minimal effect on performance. The key concept is that the parts of the scene that have the highest amount of memory accesses are replicated on all GPUs. We propose two methods for maximizing the performance of path tracing when working with partially distributed scene data. Both methods work on the memory management level and therefore path tracer data structures do not have to be redesigned, making our approach applicable to other path tracers with only minor changes in their code. As a proof of concept, we have enhanced the open-source Blender Cycles path tracer. The approach was validated on scenes of sizes up to 169 GB. We show that only 1–5% of the scene data needs to be replicated to all machines for such large scenes. On smaller scenes we have verified that the performance is very close to rendering a fully replicated scene. In terms of scalability we have achieved a parallel efficiency of over 94% using up to 16 GPUs.
Milan Jaros, Lubomir Riha, Petr Strakos, Matej Spetko
ACM Trans. Graph.2
2020 Toward an End-to-End Auto-tuning Framework in HPC PowerStack
abstract
Efficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software stack to manage power and energy of high-performance computing systems and standardizes the interfaces between different components of the stack. This survey paper presents the findings of a working group focused on the end-to-end tuning of the PowerStack. First, we provide a background on the PowerStack layer-specific tuning efforts in terms of their high-level objectives, the constraints and optimization goals, layer-specific telemetry, and control parameters, and we list the existing software solutions that address those challenges. Second, we propose the PowerStack end-to-end auto-tuning framework, identify the opportunities in co-tuning different layers in the PowerStack, and present specific use cases and solutions. Third, we discuss the research opportunities and challenges for collective auto-tuning of two or more management layers (or domains) in the PowerStack. This paper takes the first steps in identifying and aggregating the important R&D challenges in streamlining the optimization efforts across the layers of the PowerStack.
Xingfu Wu, Aniruddha Marathe, Siddhartha Jana, Ondrej Vysocky, Jophin John, Andrea Bartolini, Lubomir Riha, Michael Gerndt, Valerie Taylor 0001, Sridutt Bhalachandra
CLUSTER7
2020 Batched transpose-free ADI-type preconditioners for a Poisson solver on GPGPUs
Peter Arbenz, Lubomir Riha
J. Parallel Distributed Comput.2
2019 HPC, Cloud and Big-Data Convergent Architectures: The LEXIS Approach
Alberto Scionti, Jan Martinovic, Olivier Terzo, Etienne Walter, Marc Levrier, Stephan Hachinger, Donato Magarielli, Thierry Goubier, Stéphane Louise, Antonio Parodi, Sean Murphy, Carmine D'Amico, Simone Ciccia, Emanuele Danovaro, Martina Lagasio, Frédéric Donnat, Martin Golasowski, Tiago Quintino, James Nicholas Hawkes, Tomás Martinovic, Lubomir Riha, Katerina Slaninová, Stefano Serra-Capizzano, Roberto Peveri
CISIS21
2019 An Approach for Parallel Loading and Pre-Processing of Unstructured Meshes Stored in Spatially Scattered Fashion
abstract
This paper presents a workflow for parallel loading of database files containing sequentially stored unstructured meshes that are not considered to be efficiently read in parallel. In such a file consecutive elements are not spatially located and their respective nodes are at unknown positions in the file. This makes parallel loading challenging since adjacent elements are on different MPI processes, and their respective nodes are on unknown MPI processes. These two facts lead to a high communication overhead and very poor scalability if not addressed properly. In a standard approach, a sequentially stored mesh is sequentially converted to a particular parallel format accepted by a solver. This represents a significant bottleneck. Our proposed algorithm demonstrates that this bottleneck can be overcome, since it is able to (i) efficiently recreate an arbitrary stored sequential mesh in the distributed memory of a supercomputer without gathering the information into a single MPI rank, and (ii) prepare the mesh for massively parallel solvers.
Ondrej Meca, Lubomir Riha, Tomás Brzobohatý
IPDPS2
2019 Domain knowledge specification for energy tuning
abstract
Summary To overcome the challenges of energy consumption of HPC systems, the European Union Horizon 2020 READEX (Runtime Exploitation of Application Dynamism for Energy‐efficient Exascale computing) project uses an online auto‐tuning approach to improve energy efficiency of HPC applications. The READEX methodology pre‐computes optimal system configurations at design‐time, such as the CPU frequency, for instances of program regions and switches at runtime to the configuration given in the tuning model when the region is executed. READEX goes beyond previous approaches by exploiting dynamic changes of a region's characteristics by leveraging region and characteristic specific system configurations. While the tool suite supports an automatic approach, specifying domain knowledge such as the structure and characteristics of the application and application tuning parameters can significantly help to create a more refined tuning model. This paper presents the means available for an application expert to provide domain knowledge and presents tuning results for some benchmarks.
Madhura Kumaraswamy, Anamika Chowdhury, Michael Gerndt, Zakaria Bendifallah, Othman Bouizi, Uldis Locans, Lubomir Riha, Ondrej Vysocky, Martin Beseda, Jan Zapletal
Concurr. Comput. Pract. Exp.7
2017 READEX: Linking two ends of the computing continuum to improve energy-efficiency in dynamic applications
abstract
In both the embedded systems and High Performance Computing domains, energy-efficiency has become one of the main design criteria. Efficiently utilizing the resources provided in computing systems ranging from embedded systems to current petascale and future Exascale HPC systems will be a challenging task. Suboptimal designs can potentially cause large amounts of underutilized resources and wasted energy. In both domains, a promising potential for improving efficiency of scalable applications stems from the significant degree of dynamic behaviour, e.g., runtime alternation in application resource requirements and workloads. Manually detecting and leveraging this dynamism to improve performance and energy-efficiency is a tedious task that is commonly neglected by developers. However, using an automatic optimization approach, application dynamism can be analysed at design time and used to optimize system configurations at runtime. The European Union Horizon 2020 READEX (Runtime Exploitation of Application Dynamism for Energy-efficient eX-ascale computing) project will develop a tools-aided auto-tuning methodology inspired by the system scenario methodology used in embedded systems. Dynamic behaviour of HPC applications will be exploited to achieve improved energy-efficiency and performance. Driven by a consortium of European experts from academia, HPC resource providers, and industry, the READEX project aims at developing the first of its kind generic framework to split design time and runtime automatic tuning while targeting heterogeneous system at the Exascale level. This paper describes plans for the project as well as early results achieved during its first year. Furthermore, it is shown how project results will be brought back into the embedded systems domain.
Per Gunnar Kjeldsberg, Andreas Gocht, Michael Gerndt, Lubomir Riha, Joseph Schuchart, Umbreen Sabir Mian
DATE4
2016 Implementation of the efficient communication layer for the highly parallel total FETI and hybrid total FETI solvers
Lubomir Riha, Tomás Brzobohatý, Alexandros Markopoulos, Marta Jarosová, Tomás Kozubek, David Horák, Vaclav Hapla
Parallel Comput.1
2015 Optimization of selected remote sensing algorithms for embedded Nvidia Kepler GPU architecture
abstract
This paper evaluates the potential of embedded Graphic Processing Units in the Nvidia's Tegra K1 for onboard processing. The performance is compared to a general purpose multi-core CPU and full fledge GPU accelerator. This study uses two algorithms: Wavelet Spectral Dimension Reduction of Hyperspectral Imagery and Automated Cloud-Cover Assessment (ACCA) Algorithm. Tegra K1 achieved 51% for ACCA algorithm and 20% for the dimension reduction algorithm, as compared to the performance of the high-end 8-core server Intel Xeon CPU with 13.5 times higher power consumption.
Lubomir Riha, Jacqueline LeMoigne-Stewart, Tarek A. El-Ghazawi
IGARSS1
2015 Communication efficient work distributions in stencil operation based applications
abstract
Summary In recent years, the use of accelerators in conjunction with CPUs, known as heterogeneous computing, has brought about significant performance increases for scientific applications. One of the best examples of this is lattice quantum chromodynamics (QCD), a stencil operation based simulation. These simulations have a large memory footprint necessitating the use of many graphics processing units (GPUs) in parallel. This requires the use of a heterogeneous cluster with one or more GPUs per node. In order to obtain optimal performance, it is necessary to determine an efficient communication pattern between GPUs on the same node and between nodes. In this paper, we present a performance model based method for minimizing the communication time of applications with stencil operations, such as lattice QCD, on heterogeneous computing systems with a non‐blocking InfiniBand interconnection network. The proposed method is able to increase the performance of the most computationally intensive kernel of lattice QCD by 25% due to improved overlapping of communication and computation. We also demonstrate that the aforementioned performance model and efficient communication patterns can be used to determine a cost efficient heterogeneous system design for stencil operation based applications. Copyright © 2014 John Wiley & Sons, Ltd.
Joseph Schneible, Lubomir Riha, Maria Malik, Tarek A. El-Ghazawi, Andrei Alexandru
Concurr. Comput. Pract. Exp.2
2013 Application-specific processors for web-browsing: An exploration and evaluation of the design space
abstract
The current trend in computing has been to add more and more to the CPU; especially bigger and bigger caches and more cache levels. Based on these observations, we sought to see if bigger is always better. We test this by performing an architectural design space exploration of various cache and frequency configurations for ARM processors. Analyzing the data, we made the surprising discovery that bigger is not always better and we should in fact be taking a step back in the architectural evolutionary roadmap for some applications. In this study, we performed an analysis of the performance of web-browsers versus the architectural configuration and related it to end-user satisfaction. In the end, we were able to determine that a scaled back modern core would not only be sufficient, but improve the performance of the web-browser. In doing this, we have also developed GW-GEM5 a set of tools for the creation, monitoring and analysis of concurrent gem5 simulations on computer clusters for use in design space parameter studies.
Gabriel Yessin, Lubomir Riha, Tarek A. El-Ghazawi, David Mayhew
ASAP2
2011 Real-time motion object tracking using GPU
abstract
Motion detection is an important computer vision problem that has been used in different applications including surveillance. Most of the applications require fast processing due to their real time nature. The GPU (Graphic Processing Unit) is used as a cost-efficient tool that provides great opportunities for parallel processing. This paper studies the performance of the GPU in implementing a realistic motion detection application and compares the performance of multiple GPUs and CPU for this class of applications. The experimental results have shown that the GPU based system demonstrated high speedup and it is therefore a good low cost solution for real-time scenarios such as in video surveillance systems.
Lubomir Riha, Hoda El-Sayed
AICCSA1
2011 GPU accelerated one-pass algorithm for computing minimal rectangles of connected components
abstract
The connected component labeling is an essential task for detecting moving objects and tracking them in video surveillance application. Since tracking algorithms are designed for real-time applications, efficiencies of the underlying algorithms become critical. In this paper we present a new one-pass algorithm for computing minimal binding rectangles of all the connected components of background foreground segmented video frames (binary data) using GPU accelerator. The given image frame is scanned once in raster scan mode and the background foreground transition information is stored in a directed-graph where each transition is represented by a node. This data structure contains the locations of object edges in every row, and it is used to detect connected components in the image and extract its main features, e.g. bounding box size and location, location of the centroid, real size, etc. Further we use GPU acceleration to speed up feature extraction from the image to a directed graph from which minimal bounding rectangles will be computed subsequently. Also we compare the performance of GPU acceleration (using Tesla C2050 accelerator card) with the performance of multi-core (up 24 cores) general purpose CPU implementation of the algorithm.
Lubomir Riha, Mareboyana Manohar
WACV1