Jiri Filipovic

dblp:19/6116 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-5703-9673ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Automatic tuning based on hardware performance counters and machine learning
abstract
This paper presents a Machine Learning (ML) methodology for automatically tuning parallel applications in heterogeneous High Performance Computing (HPC) environments using Hardware Performance Counters (HwPCs). The methodology addresses three critical challenges: counter quantity versus accessibility tradeoff, data interpretation complexity, and dynamic optimization needs. The introduced ensemble-based methodology automatically identifies minimal yet informative HwPC sets for code region identification and tuning parameter optimization. Experimental validation demonstrates high accuracy in predicting optimal thread allocation ( > 0.90 K-fold accuracy) and thread affinity ( > 0.95 accuracy) while requiring only 4–6 HwPCs. Compared to search-based methods like OpenTuner, the methodology achieves competitive performance with dramatically reduced optimization time. The architecture-agnostic design enables consistent performance across CPU and GPU platforms. These results establish a foundation for efficient, portable, automatic, and scalable tuning of parallel applications.
Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Jordi Alcaraz
Future Gener. Comput. Syst.4
2026 Towards analysis and refinement of auto-tuning spaces
abstract
Source code-level auto-tuning enables applications to adapt their implementation to maintain peak performance under varying execution environments (i. e.hardware, input, or application settings). However, the performance of the auto-tuned code is inherently tied to the design of the tuning space (the space of possible changes to the code). An ideal tuning space must include configurations diverse enough to ensure high performance across all targeted environments while simultaneously eliminating redundant or inefficient regions that slow the tuning space search process. Traditional research has focused primarily on identifying optimization opportunities in the code and on efficient tuning space search. However, there is no rigorous methodology or tool supporting analysis and refinement of the tuning spaces, allowing for the addition of configurations that perform well in an unseen environment or the removal of configurations that perform poorly in any realistic environment. In this short communication, we argue that hardware performance counters should be used to analyze tuning spaces, and that such an analysis would allow programmers to refine the tuning spaces by adding configurations that unlock additional performance in unseen environments and removing those unlikely to produce efficient code in any realistic environment. While our primary goal is to introduce this research question and foster discussion, we also present a preliminary methodology for tuning-space analysis. We validate our approach through a case study using a GPU implementation of an N-body simulation. Our results demonstrate that the proposed analysis can detect the weaknesses of a tuning space: based on its outcomes, we refined the tuning space, improving the average configuration performance 3 . 3 × , and the best-performing configuration by 2 − 18 % .
Jiri Filipovic, Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora
Parallel Comput.1
2025 Estimating resource budgets to ensure autotuning efficiency
Jaroslav Olha, Jana Hozzová, Matej Antol, Jiri Filipovic
Parallel Comput.4
2024 Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning Techniques
abstract
Abstract Leveraging hardware performance counters provides valuable insights into system resource utilization, aiding performance analysis and tuning for parallel applications. The available counters vary with architecture and are collected at execution time. Their abundance and the limited number of registers for measurement make gathering laborious and costly. Efficient characterization of parallel regions necessitates a dimension reduction strategy. While recent efforts have focused on manually reducing the number of counters for specific architectures, this paper introduces a novel approach: an automatic dimension reduction technique for efficiently characterizing parallel code regions across diverse architectures. The methodology is based on Machine Learning ensembles because of their precision and ability at capturing different relationships between the input features and the target variables. Evaluation results show that ensembles can successfully reduce the number of hardware performance counters that characterize a code region. We validate our approach on CPUs using a comprehensive dataset of OpenMP regions, showing that any region can be accurately characterized by 8 relevant hardware performance counters. In addition, we also apply the proposed methodology on GPUs using a reduced set of kernels, demonstrating its effectiveness across various hardware configurations and workloads.
Suren Harutyunyan Gevorgyan, Eduardo César, Anna Sikora, Jiri Filipovic, Akash Dutta, Ali Jannesari, Jordi Alcaraz
Euro-Par (1)4
2024 A methodology for comparing optimization algorithms for auto-tuning
abstract
Adapting applications to optimally utilize available hardware is no mean feat: the plethora of choices for optimization techniques are infeasible to maximize manually. To this end, auto-tuning frameworks are used to automate this task, which in turn use optimization algorithms to efficiently search the vast searchspaces. However, there is a lack of comparability in studies presenting advances in auto-tuning frameworks and the optimization algorithms incorporated. As each publication varies in the way experiments are conducted, metrics used, and results reported, comparing the performance of optimization algorithms among publications is infeasible. The auto-tuning community identified this as a key challenge at the 2022 Lorentz Center workshop on auto-tuning. The examination of the current state of the practice in this paper further underlines this. We propose a community-driven methodology composed of four steps regarding experimental setup, tuning budget, dealing with stochasticity, and quantifying performance. This methodology builds upon similar methodologies in other fields while taking into account the constraints and specific characteristics of the auto-tuning field, resulting in novel techniques. The methodology is demonstrated in a simple case study that compares the performance of several optimization algorithms used to auto-tune CUDA kernels on a set of modern GPUs. We provide a software tool to make the application of the methodology easy for authors, and simplifies reproducibility of results.
Floris-Jan Willemsen, Richard Schoonhoven, Jiri Filipovic, Jacob Odgård Tørring, Rob van Nieuwpoort, Ben van Werkhoven
Future Gener. Comput. Syst.3
2023 pyCaverDock: Python implementation of the popular tool for analysis of ligand transport with advanced caching and batch calculation support
abstract
SUMMARY: Access pathways in enzymes are crucial for the passage of substrates and products of catalysed reactions. The process can be studied by computational means with variable degrees of precision. Our in-house approximative method CaverDock provides a fast and easy way to set up and run ligand binding and unbinding calculations through protein tunnels and channels. Here we introduce pyCaverDock, a Python3 API designed to improve user experience with the tool and further facilitate the ligand transport analyses. The API enables users to simplify the steps needed to use CaverDock, from automatizing setup processes to designing screening pipelines. AVAILABILITY AND IMPLEMENTATION: pyCaverDock API is implemented in Python 3 and is freely available with detailed documentation and practical examples at https://loschmidt.chemi.muni.cz/caverdock/.
Ondrej Vavra, Jakub Beránek, Jan Stourac, Martin Surkovský, Jiri Filipovic, Jirí Damborský, Jan Martinovic, David Bednar
Bioinform.5
2022 Using hardware performance counters to speed up autotuning convergence on GPUs
Jiri Filipovic, Jana Hozzová, Amin Nezarat, Jaroslav Olha, Filip Petrovic
J. Parallel Distributed Comput.1
2020 Exploiting historical data: Pruning autotuning spaces and estimating the number of tuning steps
abstract
Summary Autotuning, the practice of automatic tuning of applications to provide performance portability, has received increased attention in the research community, especially in high performance computing. Ensuring high performance on a variety of hardware usually means modifications to the code, often via different values of a selected set of parameters, such as tiling size, loop unrolling factor, or data layout. However, the search space of all possible combinations of these parameters can be large, which can result in cases where the benefits of autotuning are outweighed by its cost, especially with dynamic tuning. Therefore, estimating the tuning time in advance or shortening the tuning time is very important in dynamic tuning applications. We have found that certain properties of tuning spaces do not vary much when hardware is changed. In this article, we demonstrate that it is possible to use historical data to reliably predict the number of tuning steps that is necessary to find a well‐performing configuration and to reduce the size of the tuning space. We evaluate our hypotheses on a number of HPC benchmarks written in CUDA and OpenCL, using several different generations of GPUs and CPUs.
Jaroslav Olha, Jana Hozzová, Jan Fousek, Jiri Filipovic
Concurr. Comput. Pract. Exp.4
2020 A benchmark set of highly-efficient CUDA and OpenCL kernels and its dynamic autotuning with Kernel Tuning Toolkit
Filip Petrovic, David Strelák, Jana Hozzová, Jaroslav Olha, Richard Trembecký, Siegfried Benkner, Jiri Filipovic
Future Gener. Comput. Syst.7
2020 CaverDock: A Novel Method for the Fast Analysis of Ligand Transport
abstract
Here we present a novel method for the analysis of transport processes in proteins and its implementation called CaverDock. Our method is based on a modified molecular docking algorithm. It iteratively places the ligand along the access tunnel in such a way that the ligand movement is contiguous and the energy is minimized. The result of CaverDock calculation is a ligand trajectory and an energy profile of transport process. CaverDock uses the modified docking program Autodock Vina for molecular docking and implements a parallel heuristic algorithm for searching the space of possible trajectories. Our method lies in between the geometrical approaches and molecular dynamics simulations. Contrary to the geometrical methods, it provides an evaluation of chemical forces. However, it is far less computationally demanding and easier to set up compared to molecular dynamics simulations. CaverDock will find a broad use in the fields of computational enzymology, drug design, and protein engineering. The software is available free of charge to the academic users at https://loschmidt.chemi.muni.cz/caverdock/.
Jiri Filipovic, Ondrej Vavra, Jan Plhak, David Bednar, Sérgio M. Marques, Jan Brezovsky, Ludek Matyska, Jirí Damborský
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 CaverDock: a molecular docking-based tool to analyse ligand transport through protein tunnels and channels
abstract
MOTIVATION: Protein tunnels and channels are key transport pathways that allow ligands to pass between proteins' external and internal environments. These functionally important structural features warrant detailed attention. It is difficult to study the ligand binding and unbinding processes experimentally, while molecular dynamics simulations can be time-consuming and computationally demanding. RESULTS: CaverDock is a new software tool for analysing the ligand passage through the biomolecules. The method uses the optimized docking algorithm of AutoDock Vina for ligand placement docking and implements a parallel heuristic algorithm to search the space of possible trajectories. The duration of the simulations takes from minutes to a few hours. Here we describe the implementation of the method and demonstrate CaverDock's usability by: (i) comparison of the results with other available tools, (ii) determination of the robustness with large ensembles of ligands and (iii) the analysis and comparison of the ligand trajectories in engineered tunnels. Thorough testing confirms that CaverDock is applicable for the fast analysis of ligand binding and unbinding in fundamental enzymology and protein engineering. AVAILABILITY AND IMPLEMENTATION: User guide and binaries for Ubuntu are freely available for non-commercial use at https://loschmidt.chemi.muni.cz/caverdock/. The web implementation is available at https://loschmidt.chemi.muni.cz/caverweb/. The source code is available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ondrej Vavra, Jiri Filipovic, Jan Plhak, David Bednar, Sérgio M. Marques, Jan Brezovsky, Jan Stourac, Ludek Matyska, Jirí Damborský
Bioinform.2
2015 OpenCL Kernel Fusion for GPU, Xeon Phi and CPU
abstract
Kernel fusion is an optimization method, in which the code from several kernels is composed to create a new, fused kernel. It can push the performance of kernels beyond limits given for their isolated, unfused form. In this paper, we introduce a classification of different types of kernel fusion for both data dependent and data independent kernels. We study kernel fusion on three types of OpenCL devices: GPU, Xeon Phi and CPU. Those hardware platforms have quite different properties, thus, kernel fusion often affects performance in quite different ways. We analyze the impact of kernel fusion on those hardware platforms and show how it can be used to improve performance. Based on our study we also introduce a basic transformation method for generating fused kernels, which has good potential to be automatized.
Jiri Filipovic, Siegfried Benkner
SBAC-PAD1
2015 Optimizing CUDA code by kernel fusion: application on BLAS
Jiri Filipovic, Matus Madzin, Jan Fousek, Ludek Matyska
J. Supercomput.1
2008 Multiple Ligand Trajectory Docking Study - Semiautomatic Analysis of Molecular Dynamics Simulations using EGEE gLite Services
abstract
Interactions between large biomolecules and smaller bio-active ligands are usually studied through a process called docking. Its aim is to find an energetically favorable orientation of a ligand within an active site of abiomolecule. Chemical reactions take place in active siteand the role of the ligand is either to speed up, slow down or change the reaction (e.g., an enzyme catalyzed hydrolysis), which is why it can have huge pharmaceutical or other commercial impact. We present a tool that supports effective management and control of a typical workflow of docking parametric study. Selected subsets of ligands and protein trajectory snapshots can be displayed in three different views and further analyzed. Finally, the application supports spawning and steering underlying computations running on the Grid.
Ales Krenek, Martin Petrek, Jan Kmunícek, Jiri Filipovic, Zdenek Sustr, Frantisek Dvorák, Jirí Sitera, Jiri Wiesner, Ludek Matyska
PDP4