EDBT 2026 Demo / reviewers in the wild / expert
Jirí Jaros
dblp:13/2290
· DBLP profile ↗
24ranked-venue papers
11as first author
9since 2021 · last 2025
0000-0002-0087-8804ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KuBench: A Kubernetes-Based Environment for Standardized REST API Framework Performance Evaluation
Ondrej Olsak, Matej Sauer, Marta Jaros, Jirí Jaros |
ICWE | 4 |
| 2024 | Prediction of Distributed Ultrasound Simulation Execution Time U sing Machine LearningabstractThis study introduces a comprehensive system de-signed to predict the execution time of k- Wave ultrasound simulations, factoring in the domain size and allocated computing resources. The predictive models, developed using symbolic regression and neural networks, were trained on historical performance data acquired from the Barbora supercomputer. For domain sizes with optimal parameters, the symbolic regression model outperformed, achieving an average error of 5.64 %. Conversely, the neural network showed commendable efficacy in general domain scenarios, with an average error of 8.25 %. Notably, in both instances, the average error remained below the 10% threshold, aligning closely with the uncertainty inherent in the measured data and the execution of real large-scale jobs. Consequently, this predictive system is well-suited for deployment in resource optimization frameworks, significantly enhancing the efficiency of Iarge-scale simulation executions. Jirí Jaros, Marta Jaros, Martin Buchta |
CEC | 1 |
| 2024 | Accelerating Ultrasound Wave Propagation Simulations using Pruned FFTabstractThe use of ultrasound in non-invasive medical procedures is a rapidly expanding area of medicine. The success of these treatments often depends on complex ultrasound simulations that require significant computing power, time, and associated calculation costs. To solve the differential equations associated with these simulations, the pseudo-spectral method with Fourier basis functions is employed. Thus, a significant part of the simulation time is spent computing the Fast Fourier Transform. This paper presents an approach that has the potential to reduce computation time and, consequently, the calculation costs of ultrasound wave propagation simulations used in the pre-planning phase of non-invasive treatments by involving the pruned Fast Fourier Transform algorithm (pruned FFT). The paper employs spectrum filtration using a binary map to emulate the behaviour of the pruned FFT. This allows for the evaluation of the impact of the pruned FFT on the number of computed elements in the spectral domain and the accuracy of the simulation. Results on real data have shown that it is possible to replace the Fast Fourier Transform (FFT) applied to acoustic pressure and velocity with the pruned version of the algorithm while obtaining results that are suitable for pre-planning purposes, thereby reducing computation time of the treatment planning. Involving the pruned FFT can also enable the execution of simulations in higher resolution domains with much faster execution times. In some cases, we were able to achieve around 90% accuracy on the single edge of the 2D domain. Ondrej Olsak, Jirí Jaros |
HPCC | 2 |
| 2024 | Acceleration of Ultrasound Neurostimulation Using Mixed-Precision ArithmeticabstractUltrasound neurostimulation, a technique that modulates the brain's electrical activity, has emerged as a significant secondary treatment option for cases resistant to pharmacological interventions. The therapy is achievable through the application of a three-dimensional steerable ultrasound, directed by patient-specific stimulation plans. These plans are meticulously crafted through full-wave ultrasound propagation simulations. Nonetheless, the computational intensity required for calculating these plans poses a significant challenge, often reaching the memory capacities of contemporary graphics processing units (GPUs). By representing material properties and k-space operators more efficiently, we achieved up to 22% reduction in GPU memory usage, while accelerating calculations by 8.5% on an Nvidia Volta V100. This optimization introduced an error that reduced focal pressure by 0.5% without any focus movement, values that are clinically acceptable. Jirí Jaros, Radek Duchon |
HPDC | 1 |
| 2024 | k-Dispatch: Enabling Cost-Optimized Biomedical Workflow OffloadingabstractAutomated execution of computational workflows has become a critical issue in achieving high productivity in various research and development fields. Over the last few years, workflows have emerged as a significant abstraction of numerous real-world processes and phenomena, including digital twins, personalized medicine, and simulation-based science in general. k-Dispatch is a novel tool designed for the efficient offloading of biomedical workflows to remote high-performance computing clusters or cloud. In addition to data transfers, reporting, error handling and remote computations monitoring, k-Dispatch leverages a set of optimizations to dynamically determine suitable execution parameters for individual tasks within workflows, aiming to meet predefined constraints and optimization criteria. k-Dispatch has been successfully deployed within k-Plan, an advanced modelling tool for planning transcranial ultrasound stimulation (TUS) procedures. Marta Jaros, Jirí Jaros |
HPDC | 2 |
| 2024 | Techniques for Efficient Fourier Transform Computation in Ultrasound SimulationsabstractNoninvasive ultrasound surgeries represent a rapidly growing field in medical applications. Preoperative planning often relies on computationally expensive ultrasound simulations. This paper explores methods to accelerate these simulations by reducing the computation time of the Fourier transform, which is an integral part of the simulation in the k-Wave toolbox. Two experiments and their results will be presented. The first investigates substituting the standard Fast Fourier Transform (FFT) with a Sparse Fourier Transform (SFT). The second approach utilises filtering of the frequency spectrum, inspired by image compression algorithms. The aim of both experiments is to find a suitable method for accelerating the Fourier transform while utilising the sparsity of the spectrum in acoustic pressure. Our findings show that filtering offers significantly better results in terms of computation error, leading to a substantial reduction in overall simulation runtime. Ondrej Olsak, Jirí Jaros |
HPDC | 2 |
| 2022 | Optimization of Execution Parameters of Moldable Ultrasound Workflows Under Incomplete Performance Data
Marta Jaros, Jirí Jaros |
JSSPP | 2 |
| 2021 | Performance-Cost Optimization of Moldable Scientific Workflows
Marta Jaros, Jirí Jaros |
JSSPP | 2 |
| 2021 | Distributed Evolutionary Design of High Intensity Focused Ultrasound Treatment PlansabstractHigh-Intensity Focused Ultrasound (HIFU) is a modern and still evolving technique used to treat a variety of solid malignant cells in a well-defined volume including breast, liver, pancreas, prostate or uterine broids or other general soft-tissue sarcomas. HIFU treatments allow a noninvasive and non-ionising approach when compared to more conventional cancer treatments, such as radio and chemo-therapy or open surgery, which can lead to a multitude of complications after the treatment. In recent years, a realistic thermal model accounting for patient-specific materials to design HIFU treatment plans was introduced, along with an evolutionary strategy to optimize them. However, the execution times of this model is too prohibitive to allow for a routine use. This paper presents a comparison of two distinct distributed evolutionary models employing a further optimized fitness model. The experiments show up to 6 times decrease in the evolution time. These improvements allowed to investigate a new real-life based benchmark and use-case. Jakub Chlebik, Jirí Jaros |
SMC | 2 |
| 2020 | Accelerated Design of HIFU Treatment Plans Using Island-Based Evolutionary Strategy
Filip Kuklis, Marta Jaros, Jirí Jaros |
EvoApplications | 3 |
| 2020 | Optimizing Biomedical Ultrasound Workflow Scheduling Using Cluster Simulations
Marta Jaros, Dalibor Klusácek, Jirí Jaros |
JSSPP | 3 |
| 2014 | Solving the Multidimensional Knapsack Problem using a CUDA accelerated PSOabstractThe Multidimensional Knapsack Problem (MKP) represents an important model having numerous applications in combinatorial optimisation, decision-making and scheduling processes, cryptography, etc. Although the MKP is easy to define and implement, the time complexity of finding a good solution grows exponentially with the problem size. Therefore, novel software techniques and hardware platforms are being developed and employed to reduce the computation time. This paper addresses the possibility of solving the MKP using a GPU accelerated Particle Swarm Optimisation (PSO). The goal is to evaluate the attainable performance benefit when using a highly optimised GPU code instead of an efficient multi-core CPU implementation, while preserving the quality of the search process. The paper shows that a single Nvidia GTX 580 graphics card can outperform a quad-core CPU by a factor of 3.5 to 9.6, depending on the problem size. As both implementations are memory bound, these speed-ups directly correspond to the memory bandwidth ratio between the investigated GPU and CPU. Drahoslav Zan, Jirí Jaros |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | GPU-accelerated evolutionary design of the complete exchange communication on wormhole networksabstractThe communication overhead is one of the main challenges in the exascale era, where millions of compute cores are expected to collaborate on solving complex jobs. However, many algorithms will not scale since they require complex global communication and synchronisation. In order to perform the communication as fast as possible, contentions, blocking and deadlock must be avoided. Recently, we have developed an evolutionary tool producing fast and safe communication schedules reaching the lower bound of the theoretical time complexity. Unfortunately, the execution time associated with the evolution process raises up to tens of hours, even when being run on a multi-core processor. In this paper, we propose a revised implementation accelerated by a single Graphic Processing Unit (GPU) delivering speed-up of 5 compared to a quad-core CPU. Subsequently, we introduce an extended version employing up to 8 GPUs in a shared memory environment offering a speed-up of almost 30. This significantly extends the range of interconnection topologies we can cover. Jirí Jaros, Radek Tyrala |
GECCO | 1 |
| 2012 | Multi-GPU island-based genetic algorithm for solving the knapsack problemabstractThis paper introduces a novel implementation of the genetic algorithm exploiting a multi-GPU cluster. The proposed implementation employs an island-based genetic algorithm where every GPU evolves a single island. The individuals are processed by CUDA warps, which enables the solution of large knapsack instances and eliminates undesirable thread divergence. The MPI interface is used to exchange genetic material among isolated islands and collect statistical data. The characteristics of the proposed GAs are investigated on a two-node cluster composed of 14 Fermi GPUs and 4 six-core Intel Xeon processors. The overall GPU performance of the proposed GA reaches 5.67 TFLOPS. Jirí Jaros |
IEEE Congress on Evolutionary Computation | 1 |
| 2012 | A Fair Comparison of Modern CPUs and GPUs Running the Genetic Algorithm under the Knapsack Benchmark
Jirí Jaros, Petr Pospichal |
EvoApplications | 1 |
| 2012 | Implementation of 3D FFTs Across Multiple GPUs in Shared Memory EnvironmentsabstractIn this paper, a novel implementation of the distributed 3D Fast Fourier Transform (FFT) on a multi-GPU platform using CUDA is presented. The 3D FFT is the core of many simulation methods, thus its fast calculation is critical. The main bottleneck of the distributed 3D FFT is the global data exchange which must be performed. The latest version of CUDA introduces direct GPU-to-GPU transfers using a Unified Virtual Address space (UVA) that provides new possibilities for optimising the communication part of the FFT. Here, we propose different implementations of the distributed 3D FFT, investigate their behaviour, and compare their performance with the single GPU CUFFT and CPU-based FFTW libraries. In particular, we demonstrate the advantage of direct GPU-to-GPU transfers over data exchanges via host main memory. Our preliminary results show that running the distributed 3D FFT with four GPUs can bring a 12% speedup over the single node (CUFFT) while also enabling the calculation of 3D FFTs of larger datasets. Replacing the global data exchange via shared memory with direct GPU-to-GPU transfers reduces the execution time by up to 49%. This clearly shows that direct GPU-to-GPU transfers are the key factor in obtaining good performance on multi-GPU systems. Nimalan Nandapalan, Jirí Jaros, Alistair P. Rendell, Bradley E. Treeby |
PDCAT | 2 |
| 2010 | Parallel Genetic Algorithm on the CUDA Architecture
Petr Pospichal, Jirí Jaros, Josef Schwarz |
EvoApplications (1) | 2 |
| 2010 | Evolutionary-based conflict-free scheduling of collective communications on spidergon NoCsabstractThe Spidergon interconnection network has become popular recently in multiprocessor systems on chips. To the best of our knowledge, algorithms for collective communications (CC) have not been discussed in the literature as yet, contrary to pair-wise routing algorithms. The paper investigates complexity of CCs in terms of lower bounds on the number of communication steps at conflict-free scheduling. The considered networks on chip make use of wormhole switching, full duplex links and all-port non-combining nodes. A search for conflict-free scheduling of CCs has been done by means of evolutionary algorithms and the resulting numbers of communication steps have been summarized and compared to lower bounds. Time performance of CCs can be evaluated from the obtained number of steps, the given start-up time and link bandwidth. Performance prediction of applications with CCs among computing nodes of the Spidergon network is thus possible. Jirí Jaros, Václav Dvorák |
GECCO | 1 |
| 2009 | Parallel BMDA with an aggregation of probability modelsabstractThe paper is focused on the problem of aggregation of probability distribution applicable for parallel bivariate marginal distribution algorithm (pBMDA). A new approach based on quantitative combination of probabilistic models is presented. Using this concept, the traditional migration of individuals is replaced with a newly proposed technique of probability parameter migration. In the proposed strategy, the adaptive learning of the resident probability model is used. The short theoretical study is completed by an experimental works for the implemented parallel BMDA algorithm (pBMDA). The performance of pBMDA algorithm is evaluated for various problem size (scalability) and interconnection topology. In addition, the comparison with the previously published aBMDA [24] is presented. Jirí Jaros, Josef Schwarz |
IEEE Congress on Evolutionary Computation | 1 |
| 2009 | Evolutionary optimization of multistage interconnection networks performanceabstractThe paper deals with optimization of collective communications on multistage interconnection networks (MINs). In the experimental work, unidirectional MINs like Omega, Butterfly and Clos are investigated. The study is completed by bidirectional binary, fat and full binary tree. To avoid link contentions and associated delays, collective communications are processed in synchronized steps. Minimum number of steps is sought for the given network topology, wormhole switching, minimum routing and given sets of sender and/or receiver nodes. Evolutionary algorithm proposed in this paper is able to design optimal schedules for broadcast and scatter collective communications. Acquired optimum schedules can simplify the consecutive writing high-performance communication routines for application-specific networks on chip, or for development of communication libraries in case of general-purpose multistage interconnection networks. Jirí Jaros |
GECCO | 1 |
| 2008 | An evolutionary design technique for collective communications on optimal diameter-degree networksabstractScheduling collective communications (CC) in networks based on optimal graphs and digraphs has been done with the use of the evolutionary techniques. Inter-node communication patterns scheduled in the minimum number of time slots have been obtained. Numerical values of communication times derived for illustration can be used to estimate speedup of typical applications that use CC frequently. The results show that evolutionary techniques often lead to ultimate scheduling of CC that reaches theoretical bounds on the number of steps. Analysis of fault tolerance by the same techniques revealed graceful CC performance degradation for a single link fault. Once the faulty link is located, CC can be re-scheduled during a recovery period. Jirí Jaros, Václav Dvorák |
GECCO | 1 |
| 2007 | Parallel BMDA with probability model migrationabstractThe paper presents a new concept of parallel bivariate marginal distribution algorithm using the stepping stone based model of communication with the unidirectional ring topology. The traditional migration of individuals is compared with a newly proposed technique of probability model migration. The idea of the new xBMDA algorithms is to modify the learning of classical probability model (applied in the sequential BMDA). In the first strategy, the adaptive learning of the resident probability model is used. The evaluation of pair dependency, using Pearson's chi-square statistics is influenced by the relevant immigrant pair dependency according to the quality of resident and immigrant subpopulation. In the second proposed strategy, the evaluation metric is applied for the diploid mode of the aggregated resident and immigrant subpopulation. Experimental results show that the proposed adaptive BMDA outperforms the traditional concept of individual migration. Jirí Jaros, Josef Schwarz |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | An evolutionary approach to collective communication schedulingabstractIn this paper, we describe two evolutionary algorithms aimed at scheduling collective communications on interconnection networks of parallel computers. To avoid contention for links and associated delays, collective communications proceed in synchronized steps. Minimum number of steps is sought for the given network topology, wormhole (pipelined) switching, minimum routing and given sets of sender and/or receiver nodes. Used algorithms are able not only re-invent optimum schedules for known symmetric topologies like hyper-cubes, but they can find schedules even for any asymmetric or irregular topologies in case of general many-to-many collective communications. In most cases does the number of steps reach the theoretical lower bound for the given type of collective communication; if it does not, non-minimum routing can provide further improvement. Optimum schedules may serve for writing high-performance communication routines for application-specific networks on chip or for development of communication libraries in case of general-purpose interconnection networ. Jirí Jaros, Milos Ohlídal, Václav Dvorák |
GECCO | 1 |
| 2007 | Migration of probabilistic models for island-based bivariate EDA algorithmabstractThe paper presents a new concept of parallel bivariate EDA algorithm using the island-based model with the ring topology. The traditional migration of individuals is compared with a newly proposed technique for the migration of probabilistic models. Josef Schwarz, Jirí Jaros, Jiri Ocenasek |
GECCO | 2 |