Martin Krulis

dblp:49/8460 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0002-0985-8949ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-author
YearPublicationVenuePosition
2025 Tutoring LLM into a Better CUDA Optimizer
Matyás Brabec, Jirí Klepl, Michal Töpfer, Martin Krulis
Euro-Par (2)4
2025 Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
Jirí Klepl, Martin Krulis, Matyás Brabec
EuroMPI2
2025 Efficient GPU-accelerated parallel cross-correlation
Karel Madera, Adam Smelko, Martin Krulis
J. Parallel Distributed Comput.3
2024 Abstractions for C++ code optimizations in parallel high-performance applications
abstract
Many computational problems consider memory throughput a performance bottleneck, especially in the domain of parallel computing. Software needs to be attuned to hardware features like cache architectures or concurrent memory banks to reach a decent level of performance efficiency. This can be achieved by selecting the right memory layouts for data structures or changing the order of data structure traversal. In this work, we present an abstraction for traversing a set of regular data structures (e.g., multidimensional arrays) that allows the design of traversal-agnostic algorithms. Such algorithms can easily optimize for memory performance and employ semi-automated parallelization or autotuning without altering their internal code. We also add an abstraction for autotuning that allows defining tuning parameters in one place and removes boilerplate code. The proposed solution was implemented as an extension of the Noarr library that simplifies a layout-agnostic design of regular data structures. It is implemented entirely using C++ template meta-programming without any nonstandard dependencies, so it is fully compatible with existing compilers, including CUDA NVCC or Intel DPC++. We evaluate the performance and expressiveness of our approach on the Polybench-C benchmarks.
Jirí Klepl, Adam Smelko, Lukás Rozsypal, Martin Krulis
Parallel Comput.4
2023 Modeling Machine Learning Concerns in Collective Adaptive Systems
Petr Hnetynka, Martin Krulis, Michal Töpfer, Tomás Bures
MODELSWARD2
2023 Generating adaptation rule-specific neural networks
Tomás Bures, Petr Hnetynka, Martin Krulis, Frantisek Plásil, Danylo Khalyeyev, Sebastian Hahner, Stephan Seifermann, Maximilian Walter, Robert Heinrich
Int. J. Softw. Tools Technol. Transf.3
2023 Machine-learning abstractions for component-based self-optimizing systems
Michal Töpfer, Milad Abdullah, Tomás Bures, Petr Hnetynka, Martin Krulis
Int. J. Softw. Tools Technol. Transf.5
2022 Astute Approach to Handling Memory Layouts of Regular Data Structures
Adam Smelko, Martin Krulis, Miroslav Kratochvíl, Jirí Klepl, Jirí Mayer, Petr Simunek
ICA3PP2
2022 Attuning Adaptation Rules via a Rule-Specific Neural Network
Tomás Bures, Petr Hnetynka, Martin Krulis, Frantisek Plásil, Danylo Khalyeyev, Sebastian Hahner, Stephan Seifermann, Maximilian Walter, Robert Heinrich
ISoLA (3)3
2022 Ensemble-Based Modeling Abstractions for Modern Self-optimizing Systems
Michal Töpfer, Milad Abdullah, Tomás Bures, Petr Hnetynka, Martin Krulis
ISoLA (3)5
2022 Towards Model-driven Fuzzification of Adaptive Systems Specification
Tomás Bures, Petr Hnetynka, Martin Krulis, Jan Pacovsky
MODELSWARD3
2022 Simdex: A Simulator of a Real Self-adaptive job-dispatching System Backend
abstract
Self-adaptive systems comprise a complex domain of computing systems that are intensively studied but sparsely employed in real applications. Furthermore, recent trends in computer science are steering towards machine learning which has yet to fully penetrate this domain. We would like to present Simdex --- a realistic simulator of the self-adaptive backend that dispatches computing jobs among multiple workers. It is based on ReCodEx, a system for semi-automated evaluation of coding assignments that have been used for the past 5 years at our School of Computer Science. The simulator replays the workload logs recorded from ReCodEx over that period which provides a quite thorough evaluation and near-to-real feedback for the simulated scenarios. Furthermore, the design of the simulator is highly modular and allows the implementation of different self-adaptive controllers, including ones based on machine learning, as we demonstrate in our examples.
Martin Krulis, Tomás Bures, Petr Hnetynka
SEAMS1
2021 GPU-Accelerated Mahalanobis-Average Hierarchical Clustering Analysis
Adam Smelko, Miroslav Kratochvíl, Martin Krulis, Tomás Sieger
Euro-Par3
2021 Letting future programmers experience performance-related tasks
David Bednárek, Martin Krulis, Jakub Yaghob
J. Parallel Distributed Comput.2
2020 Detailed Analysis and Optimization of CUDA K-means Algorithm
abstract
K-means is one of the most frequently used algorithms for unsupervised clustering data analysis. Individual steps of the k-means algorithm include nearest neighbor finding, efficient distance computation, and cluster-wise reduction, which may be generalized to many other purposes in data analysis, visualization, and machine learning. Efficiency of the available implementations of k-means computation steps therefore directly affect many other applications. In this work, we examine the performance limits in the context of modern massively parallel GPU accelerators. Despite the existence of many published papers on this topic, we have found that crucial performance aspects of the GPU implementations remain unaddressed, including the optimizations for memory bandwidth, cache limits, and workload dispatching on problem instances of varying cluster count, dataset size, and dimensionality. We present a detailed analysis of individual computation steps and propose several optimizations that improve the overall performance on contemporary GPU architectures. Our open-source prototype exhibits significant speedup over the current state-of-the-art implementations in virtually all practical scenarios.
Martin Krulis, Miroslav Kratochvíl
ICPP1
2017 Data Preprocessing of eSport Game Records - Counter-Strike: Global Offensive
David Bednárek, Martin Krulis, Jakub Yaghob, Filip Zavoral
DATA2
2017 Improving matrix-based dynamic programming on massively parallel accelerators
David Bednárek, Michal Brabec, Martin Krulis
Inf. Syst.3
2017 Employing GPU architectures for permutation-based indexing
Martin Krulis, Hasmik Osipyan, Stéphane Marchand-Maillet
Multim. Tools Appl.1
2016 Creating Distributed Execution Plans with BobolangNG
David Bednárek, Martin Krulis, Jakub Yaghob, Filip Zavoral
ICA3PP2
2016 Efficient extraction of clustering-based feature signatures using GPU architectures
Martin Krulis, Jakub Lokoc, Tomás Skopal
Multim. Tools Appl.1
2015 Is There a Free Lunch for Image Feature Extraction in Web Applications
Martin Krulis
SISAP1
2015 Improving Parallel Processing of Matrix-Based Similarity Measures on Modern GPUs
Martin Krulis, David Bednárek, Michal Brabec
SISAP1
2014 Bobolang: a language for parallel streaming applications
abstract
At present time, the programmers may choose from a number of streaming languages. They cover various aspects of the development process of streaming applications; however, specification of complex or runtime-dependent parts of the applications still remains a great challenge. We have analysed a large amount of requirements raised by the development of multiple data streaming parallel applications and proposed a novel language called Bobolang. It contains syntactic and semantic features which allow the programmer to naturally solve most of the problems, which we met in the design of streaming applications. The language is used to specify the structure of the whole application as well as the inner structure of each operator. Thanks to the properties of the language, Bobolang can create an optimized evaluation plan which is capable of making the best use of the available hardware resources. The language has been employed in several practical problems and it has proven itself to be a very powerful tool for the development of data-intensive parallel applications.
Zbynek Falt, David Bednárek, Martin Krulis, Jakub Yaghob, Filip Zavoral
HPDC3
2014 Employing Similarity Methods for Stellar Spectra Classification in Astroinformatics
Martin Krulis, David Bednárek, Jakub Yaghob, Filip Zavoral
SISAP1
2014 Perils of Combining Parallel Distance Computations with Metric and Ptolemaic Indexing in kNN Queries
Martin Krulis, Steffen Kirchhoff, Jakub Yaghob
SISAP1
2013 Efficient Extraction of Feature Signatures Using Multi-GPU Architecture
Martin Krulis, Jakub Lokoc, Tomás Skopal
MMM (2)1
2012 Combining CPU and GPU architectures for fast similarity search
Martin Krulis, Tomás Skopal, Jakub Lokoc, Christian Beecks
Distributed Parallel Databases1
2011 Processing the signature quadratic form distance on many-core GPU architectures
abstract
The Signature Quadratic Form Distance on feature signatures represents a flexible distance-based similarity model for effective content-based multimedia retrieval. Although metric indexing approaches are able to speed up query processing by two orders of magnitude, their applicability to large-scale multimedia databases containing billions of images is still a challenging issue. In this paper, we propose the utilization of GPUs for efficient query processing with the Signature Quadratic Form Distance. We show how to process multiple distance computations in parallel and demonstrate efficient query processing by comparing many-core GPU with multi-core CPU implementations.
Martin Krulis, Jakub Lokoc, Christian Beecks, Tomás Skopal, Thomas Seidl 0001
CIKM1