VLDB 2026 Research / reviewers in the wild / expert
Martin Krulis
dblp:49/8460
· DBLP profile ↗
28ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0002-0985-8949ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tutoring LLM into a Better CUDA Optimizer
Matyás Brabec, Jirí Klepl, Michal Töpfer, Martin Krulis |
Euro-Par (2) | 4 |
| 2025 | Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
Jirí Klepl, Martin Krulis, Matyás Brabec |
EuroMPI | 2 |
| 2025 | Efficient GPU-accelerated parallel cross-correlation
Karel Madera, Adam Smelko, Martin Krulis |
J. Parallel Distributed Comput. | 3 |
| 2024 | Abstractions for C++ code optimizations in parallel high-performance applicationsabstractMany computational problems consider memory throughput a performance bottleneck, especially in the domain of parallel computing. Software needs to be attuned to hardware features like cache architectures or concurrent memory banks to reach a decent level of performance efficiency. This can be achieved by selecting the right memory layouts for data structures or changing the order of data structure traversal. In this work, we present an abstraction for traversing a set of regular data structures (e.g., multidimensional arrays) that allows the design of traversal-agnostic algorithms. Such algorithms can easily optimize for memory performance and employ semi-automated parallelization or autotuning without altering their internal code. We also add an abstraction for autotuning that allows defining tuning parameters in one place and removes boilerplate code. The proposed solution was implemented as an extension of the Noarr library that simplifies a layout-agnostic design of regular data structures. It is implemented entirely using C++ template meta-programming without any nonstandard dependencies, so it is fully compatible with existing compilers, including CUDA NVCC or Intel DPC++. We evaluate the performance and expressiveness of our approach on the Polybench-C benchmarks. Jirí Klepl, Adam Smelko, Lukás Rozsypal, Martin Krulis |
Parallel Comput. | 4 |
| 2023 | Modeling Machine Learning Concerns in Collective Adaptive Systems
Petr Hnetynka, Martin Krulis, Michal Töpfer, Tomás Bures |
MODELSWARD | 2 |
| 2023 | Generating adaptation rule-specific neural networks
Tomás Bures, Petr Hnetynka, Martin Krulis, Frantisek Plásil, Danylo Khalyeyev, Sebastian Hahner, Stephan Seifermann, Maximilian Walter, Robert Heinrich |
Int. J. Softw. Tools Technol. Transf. | 3 |
| 2023 | Machine-learning abstractions for component-based self-optimizing systems
Michal Töpfer, Milad Abdullah, Tomás Bures, Petr Hnetynka, Martin Krulis |
Int. J. Softw. Tools Technol. Transf. | 5 |
| 2022 | Astute Approach to Handling Memory Layouts of Regular Data Structures
Adam Smelko, Martin Krulis, Miroslav Kratochvíl, Jirí Klepl, Jirí Mayer, Petr Simunek |
ICA3PP | 2 |
| 2022 | Attuning Adaptation Rules via a Rule-Specific Neural Network
Tomás Bures, Petr Hnetynka, Martin Krulis, Frantisek Plásil, Danylo Khalyeyev, Sebastian Hahner, Stephan Seifermann, Maximilian Walter, Robert Heinrich |
ISoLA (3) | 3 |
| 2022 | Ensemble-Based Modeling Abstractions for Modern Self-optimizing Systems
Michal Töpfer, Milad Abdullah, Tomás Bures, Petr Hnetynka, Martin Krulis |
ISoLA (3) | 5 |
| 2022 | Towards Model-driven Fuzzification of Adaptive Systems Specification
Tomás Bures, Petr Hnetynka, Martin Krulis, Jan Pacovsky |
MODELSWARD | 3 |
| 2022 | Simdex: A Simulator of a Real Self-adaptive job-dispatching System BackendabstractSelf-adaptive systems comprise a complex domain of computing systems that are intensively studied but sparsely employed in real applications. Furthermore, recent trends in computer science are steering towards machine learning which has yet to fully penetrate this domain. We would like to present Simdex --- a realistic simulator of the self-adaptive backend that dispatches computing jobs among multiple workers. It is based on ReCodEx, a system for semi-automated evaluation of coding assignments that have been used for the past 5 years at our School of Computer Science. The simulator replays the workload logs recorded from ReCodEx over that period which provides a quite thorough evaluation and near-to-real feedback for the simulated scenarios. Furthermore, the design of the simulator is highly modular and allows the implementation of different self-adaptive controllers, including ones based on machine learning, as we demonstrate in our examples. Martin Krulis, Tomás Bures, Petr Hnetynka |
SEAMS | 1 |
| 2021 | GPU-Accelerated Mahalanobis-Average Hierarchical Clustering Analysis
Adam Smelko, Miroslav Kratochvíl, Martin Krulis, Tomás Sieger |
Euro-Par | 3 |
| 2021 | Letting future programmers experience performance-related tasks
David Bednárek, Martin Krulis, Jakub Yaghob |
J. Parallel Distributed Comput. | 2 |
| 2020 | Detailed Analysis and Optimization of CUDA K-means AlgorithmabstractK-means is one of the most frequently used algorithms for unsupervised clustering data analysis. Individual steps of the k-means algorithm include nearest neighbor finding, efficient distance computation, and cluster-wise reduction, which may be generalized to many other purposes in data analysis, visualization, and machine learning. Efficiency of the available implementations of k-means computation steps therefore directly affect many other applications. In this work, we examine the performance limits in the context of modern massively parallel GPU accelerators. Despite the existence of many published papers on this topic, we have found that crucial performance aspects of the GPU implementations remain unaddressed, including the optimizations for memory bandwidth, cache limits, and workload dispatching on problem instances of varying cluster count, dataset size, and dimensionality. We present a detailed analysis of individual computation steps and propose several optimizations that improve the overall performance on contemporary GPU architectures. Our open-source prototype exhibits significant speedup over the current state-of-the-art implementations in virtually all practical scenarios. Martin Krulis, Miroslav Kratochvíl |
ICPP | 1 |
| 2017 | Data Preprocessing of eSport Game Records - Counter-Strike: Global Offensive
David Bednárek, Martin Krulis, Jakub Yaghob, Filip Zavoral |
DATA | 2 |
| 2017 | Improving matrix-based dynamic programming on massively parallel accelerators
David Bednárek, Michal Brabec, Martin Krulis |
Inf. Syst. | 3 |
| 2017 | Employing GPU architectures for permutation-based indexing
Martin Krulis, Hasmik Osipyan, Stéphane Marchand-Maillet |
Multim. Tools Appl. | 1 |
| 2016 | Creating Distributed Execution Plans with BobolangNG
David Bednárek, Martin Krulis, Jakub Yaghob, Filip Zavoral |
ICA3PP | 2 |
| 2016 | Efficient extraction of clustering-based feature signatures using GPU architectures
Martin Krulis, Jakub Lokoc, Tomás Skopal |
Multim. Tools Appl. | 1 |
| 2015 | Is There a Free Lunch for Image Feature Extraction in Web Applications
Martin Krulis |
SISAP | 1 |
| 2015 | Improving Parallel Processing of Matrix-Based Similarity Measures on Modern GPUs
Martin Krulis, David Bednárek, Michal Brabec |
SISAP | 1 |
| 2014 | Bobolang: a language for parallel streaming applicationsabstractAt present time, the programmers may choose from a number of streaming languages. They cover various aspects of the development process of streaming applications; however, specification of complex or runtime-dependent parts of the applications still remains a great challenge. We have analysed a large amount of requirements raised by the development of multiple data streaming parallel applications and proposed a novel language called Bobolang. It contains syntactic and semantic features which allow the programmer to naturally solve most of the problems, which we met in the design of streaming applications. The language is used to specify the structure of the whole application as well as the inner structure of each operator. Thanks to the properties of the language, Bobolang can create an optimized evaluation plan which is capable of making the best use of the available hardware resources. The language has been employed in several practical problems and it has proven itself to be a very powerful tool for the development of data-intensive parallel applications. Zbynek Falt, David Bednárek, Martin Krulis, Jakub Yaghob, Filip Zavoral |
HPDC | 3 |
| 2014 | Employing Similarity Methods for Stellar Spectra Classification in Astroinformatics
Martin Krulis, David Bednárek, Jakub Yaghob, Filip Zavoral |
SISAP | 1 |
| 2014 | Perils of Combining Parallel Distance Computations with Metric and Ptolemaic Indexing in kNN Queries
Martin Krulis, Steffen Kirchhoff, Jakub Yaghob |
SISAP | 1 |
| 2013 | Efficient Extraction of Feature Signatures Using Multi-GPU Architecture
Martin Krulis, Jakub Lokoc, Tomás Skopal |
MMM (2) | 1 |
| 2012 | Combining CPU and GPU architectures for fast similarity search
Martin Krulis, Tomás Skopal, Jakub Lokoc, Christian Beecks |
Distributed Parallel Databases | 1 |
| 2011 | Processing the signature quadratic form distance on many-core GPU architecturesabstractThe Signature Quadratic Form Distance on feature signatures represents a flexible distance-based similarity model for effective content-based multimedia retrieval. Although metric indexing approaches are able to speed up query processing by two orders of magnitude, their applicability to large-scale multimedia databases containing billions of images is still a challenging issue. In this paper, we propose the utilization of GPUs for efficient query processing with the Signature Quadratic Form Distance. We show how to process multiple distance computations in parallel and demonstrate efficient query processing by comparing many-core GPU with multi-core CPU implementations. Martin Krulis, Jakub Lokoc, Christian Beecks, Tomás Skopal, Thomas Seidl 0001 |
CIKM | 1 |