EDBT 2026 Demo / reviewers in the wild / expert
Daniel Langr
dblp:21/6449
· DBLP profile ↗
11ranked-venue papers
8as first author
1since 2021 · last 2022
0000-0001-9760-7068ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-authorSystems, architecture and hardware · 5 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 30% High-performance computing · 22% Hardware accelerators and domain-specific architectures · 14% | |
| Theoretical computer science
1 paper |
Computational geometry · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory management › memory allocation
dynamic memory allocation |
0.4 | 1 | 2020 | Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems › storage management
memory-efficient storage |
0.4 | 1 | 2020 | Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020 |
Parallel and multicore computing
parallel programming models and runtimes |
0.4 | 1 | 2020 | Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs · IEEE Trans. Parallel Distributed Syst. 2020 |
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation |
0.4 | 1 | 2020 | Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020 |
Storage systems › file systems › file organization
storage formats |
0.4 | 1 | 2020 | Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020 |
High-performance computing
performance optimization at scale |
0.4 | 2 | 2020 | Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016 Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs · IEEE Trans. Parallel Distributed Syst. 2020 |
Performance modeling and evaluation
benchmarking |
0.2 | 1 | 2016 | Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.2 | 1 | 2016 | Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016 |
Computational geometry › spatial data structures
kd-tree |
0.1 | 1 | 2020 | Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020 |
Storage systems
data layout |
0.1 | 1 | 2016 | Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016 |
High-performance computing › sparse linear algebra
sparse matrix storage format |
0.1 | 1 | 2016 | Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016 |
Methods — techniques the papers use, named apart from their topics
small buffer optimization · 0.4scalable heap · 0.4memory pooling · 0.4kd-tree · 0.4k-d tree · 0.4evaluation criteria · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | CPP11sort: A parallel quicksort based on C++ threadingabstractSummary A new efficient implementation of the multithreaded quicksort algorithm called CPP11sort is presented. This implementation is built exclusively upon the threading primitives of the C++ programming language itself. The performance of CPP11sort is evaluated and compared with its mainstream competitors provided by GNU, Intel, and Microsoft. It is shown that out of the considered implementations, CPP11sort mostly yields the shortest sorting times and is the only one that is portable to any conforming C++ implementation without a need of external libraries or nonstandard compiler extensions. The experimental evaluation with various input data distributions resulted in parallel speedup between 16.1 and 44.2 on a 56‐core server and between 6.8 and 14.5 on a 10‐core workstation with enabled hyperthreading. Daniel Langr, Klára Schovánková |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Space-Efficient k-d Tree-Based Storage Format for Sparse TensorsabstractComputations with tensors are widespread in many scientific areas. Usually, the used tensors are very large but sparse, i.e., the vast majority of their elements are zero. The space complexity of sparse tensor storage formats varies significantly. For overall efficiency, it is important to reduce the execution time and additional space requirements of the initial preprocessing (i.e., converting tensors from common storage formats to the given internal format). Ivan Simecek, Claudio Kozický, Daniel Langr, Pavel Tvrdík |
HPDC | 3 |
| 2020 | Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded ProgramsabstractFrequent dynamic memory allocations (DyMAs) can significantly hinder the scalability of parallel multi-threaded programs. As the number of threads grows, DyMAs can even become the main performance bottleneck. We introduce modern tools and methods for evaluating the impact of DyMAs and present techniques for its reduction, which include scalable heap implementations, small buffer optimization, and memory pooling. Additionally, we provide a survey of state-of-the-art implementations of these techniques and study them experimentally by using a benchmark program, server simulator software, and a real-world high-performance computing application. As a result, we show that relatively small modifications in parallel program's source code or a way of its execution may substantially reduce the runtime overhead associated with the use of dynamic data structures. Daniel Langr, Martin Kocicka |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | On Memory Footprints of Partitioned Sparse MatricesabstractThe presented study analyses 563 representative benchmark sparse matrices with respect to their partitioning into uniformly-sized blocks.The aim is to minimize memory footprints of matrices.Different block sizes and different ways of storing blocks in memory are considered and statistically evaluated.Memory footprints of partitioned matrices are additionally compared with lower bounds and the CSR storage format.The average measured memory savings against CSR in case of single and double precision are 42.3 and 28.7 percents, respectively.The corresponding worst-case savings are 25.5 and 17.1 percents.Moreover, memory footprints of partitioned matrices were in average 5 times closer to their lower bounds than CSR.Based on the obtained results, we provide generic suggestions for efficient partitioning and storage of sparse matrices in a computer memory. Daniel Langr, Ivan Simecek |
FedCSIS | 1 |
| 2016 | Block Iterators for Sparse MatricesabstractFinding an optimal block size for a given sparse matrix forms an important problem for storage formats that partition matrices into uniformly-sized blocks.Finding a solution to this problem can take a significant amount of time, which, effectively, may negate the benefits that such a format brings into sparse-matrix computations.A key for an efficient solution is the ability to quickly iterate, for a particular block size, over matrix nonzero blocks.This work proposes an efficient parallel algorithm for this task and evaluate it experimentally on modern multi-core and many-core high performance computing (HPC) architectures. Daniel Langr, Ivan Simecek, Tomás Dytrych |
FedCSIS | 1 |
| 2016 | Efficient parallel evaluation of block properties of sparse matricesabstractMany storage formats for sparse matrices have been developed.Majority of these formats can be parametrized, so the algorithm for finding optimal parameters is crucial.For overall efficiency, it is important to reduce the execution time of this preprocessing.In this paper, we propose a new algorithm for the determination of the number of nonzero blocks of the given size in a sparse matrix.The proposed algorithm requires relatively a small amount of auxiliary memory.Our approach is based on the Morton reordering and bitwise manipulations.We also present a parallel (multithreaded) version and evaluate its performance and space complexity. Ivan Simecek, Daniel Langr |
FedCSIS | 2 |
| 2016 | Evaluation Criteria for Sparse Matrix Storage FormatsabstractWhen authors present new storage formats for sparse matrices, they usually focus mainly on a single evaluation criterion, which is the performance of sparse matrix-vector multiplication (SpMV) in FLOPS. Though such an evaluation is essential, it does not allow to directly compare the presented format with its competitors. Moreover, in case that matrices are within an HPC application constructed in different formats, this criterion alone is not sufficient for the key decision whether or not to convert them into the presented format for the SpMV-based application phase. We establish ten evaluation criteria for sparse matrix storage formats, discuss their advantages and disadvantages, and provide general suggestions for format authors/evaluators to make their work more valuable for the HPC community. Daniel Langr, Pavel Tvrdík |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Algorithm 947: Paraperm - Parallel Generation of Random Permutations with MPIabstractAn algorithm for parallel generation of a random permutation of a large set of distinct integers is presented. This algorithm is designed for massively parallel systems with distributed memory architectures and the MPI-based runtime environments. Scalability of the algorithm is analyzed according to the memory and communication requirements. An implementation of the algorithm in a form of a software library based on the C++ programming language and the MPI application programming interface is further provided. Finally, performed experiments are described and their results discussed. The biggest of these experiments resulted in a generation of a random permutation of 2 41 integers in slightly more than four minutes using 131072 CPU cores. Daniel Langr, Pavel Tvrdík, Tomás Dytrych, Jerry P. Draayer |
ACM Trans. Math. Softw. | 1 |
| 2013 | Storing Sparse Matrices to Files in the Adaptive-Blocking Hierarchical Storage Format
Daniel Langr, Ivan Simecek, Pavel Tvrdík |
FedCSIS | 1 |
| 2012 | Adaptive-Blocking Hierarchical Storage Format for Sparse Matrices
Daniel Langr, Ivan Simecek, Pavel Tvrdík, Tomás Dytrych, Jerry P. Draayer |
FedCSIS | 1 |
| 2005 | Clondike: Linux cluster of non-dedicated workstationsabstractClusters of workstations are a promising platform for high-performance computing. Most of such clusters are built of dedicated workstations that cannot be used for other purposes. Even though these clusters offer good price/performance ratio, their costs are not negligible. On the other hand, there exist many idle workstations connected via computer networks. The idea of exploiting such idle resources is obvious. In this paper, we describe an architecture of clusters made of non-dedicated idle workstations. These clusters aim at providing a single-system-image Linux environment. The main innovative feature is that a cluster administration is separate from administration of individual workstations. We discuss issues of security, availability, performance, and administration of such clusters. We also describe a pilot implementation and give promising experimental results. Martin Kacer, Daniel Langr, Pavel Tvrdík |
CCGRID | 2 |