Daniel Langr

dblp:21/6449 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
1since 2021 · last 2022
0000-0001-9760-7068ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-authorSystems, architecture and hardware · 5 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 30% High-performance computing · 22% Hardware accelerators and domain-specific architectures · 14%
Theoretical computer science
1 paper
Computational geometry · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management › memory allocation
dynamic memory allocation
0.412020
Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs · IEEE Trans. Parallel Distributed Syst. 2020
Storage systems › storage management
memory-efficient storage
0.412020
Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020
Parallel and multicore computing
parallel programming models and runtimes
0.412020
Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs · IEEE Trans. Parallel Distributed Syst. 2020
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation
0.412020
Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020
Storage systems › file systems › file organization
storage formats
0.412020
Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020
High-performance computing
performance optimization at scale
0.422020
Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016
Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs · IEEE Trans. Parallel Distributed Syst. 2020
Performance modeling and evaluation
benchmarking
0.212016
Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication
0.212016
Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016
Computational geometry › spatial data structures
kd-tree
0.112020
Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors · HPDC 2020
Storage systems
data layout
0.112016
Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016
High-performance computing › sparse linear algebra
sparse matrix storage format
0.112016
Evaluation Criteria for Sparse Matrix Storage Formats · IEEE Trans. Parallel Distributed Syst. 2016

Methods — techniques the papers use, named apart from their topics

small buffer optimization · 0.4scalable heap · 0.4memory pooling · 0.4kd-tree · 0.4k-d tree · 0.4evaluation criteria · 0.2
YearPublicationVenuePosition
2022 CPP11sort: A parallel quicksort based on C++ threading
abstract
Summary A new efficient implementation of the multithreaded quicksort algorithm called CPP11sort is presented. This implementation is built exclusively upon the threading primitives of the C++ programming language itself. The performance of CPP11sort is evaluated and compared with its mainstream competitors provided by GNU, Intel, and Microsoft. It is shown that out of the considered implementations, CPP11sort mostly yields the shortest sorting times and is the only one that is portable to any conforming C++ implementation without a need of external libraries or nonstandard compiler extensions. The experimental evaluation with various input data distributions resulted in parallel speedup between 16.1 and 44.2 on a 56‐core server and between 6.8 and 14.5 on a 10‐core workstation with enabled hyperthreading.
Daniel Langr, Klára Schovánková
Concurr. Comput. Pract. Exp.1
2020 Space-Efficient k-d Tree-Based Storage Format for Sparse Tensors
abstract
Computations with tensors are widespread in many scientific areas. Usually, the used tensors are very large but sparse, i.e., the vast majority of their elements are zero. The space complexity of sparse tensor storage formats varies significantly. For overall efficiency, it is important to reduce the execution time and additional space requirements of the initial preprocessing (i.e., converting tensors from common storage formats to the given internal format).
Ivan Simecek, Claudio Kozický, Daniel Langr, Pavel Tvrdík
HPDC3
2020 Reducing the Impact of Intensive Dynamic Memory Allocations in Parallel Multi-Threaded Programs
abstract
Frequent dynamic memory allocations (DyMAs) can significantly hinder the scalability of parallel multi-threaded programs. As the number of threads grows, DyMAs can even become the main performance bottleneck. We introduce modern tools and methods for evaluating the impact of DyMAs and present techniques for its reduction, which include scalable heap implementations, small buffer optimization, and memory pooling. Additionally, we provide a survey of state-of-the-art implementations of these techniques and study them experimentally by using a benchmark program, server simulator software, and a real-world high-performance computing application. As a result, we show that relatively small modifications in parallel program's source code or a way of its execution may substantially reduce the runtime overhead associated with the use of dynamic data structures.
Daniel Langr, Martin Kocicka
IEEE Trans. Parallel Distributed Syst.1
2017 On Memory Footprints of Partitioned Sparse Matrices
abstract
The presented study analyses 563 representative benchmark sparse matrices with respect to their partitioning into uniformly-sized blocks.The aim is to minimize memory footprints of matrices.Different block sizes and different ways of storing blocks in memory are considered and statistically evaluated.Memory footprints of partitioned matrices are additionally compared with lower bounds and the CSR storage format.The average measured memory savings against CSR in case of single and double precision are 42.3 and 28.7 percents, respectively.The corresponding worst-case savings are 25.5 and 17.1 percents.Moreover, memory footprints of partitioned matrices were in average 5 times closer to their lower bounds than CSR.Based on the obtained results, we provide generic suggestions for efficient partitioning and storage of sparse matrices in a computer memory.
Daniel Langr, Ivan Simecek
FedCSIS1
2016 Block Iterators for Sparse Matrices
abstract
Finding an optimal block size for a given sparse matrix forms an important problem for storage formats that partition matrices into uniformly-sized blocks.Finding a solution to this problem can take a significant amount of time, which, effectively, may negate the benefits that such a format brings into sparse-matrix computations.A key for an efficient solution is the ability to quickly iterate, for a particular block size, over matrix nonzero blocks.This work proposes an efficient parallel algorithm for this task and evaluate it experimentally on modern multi-core and many-core high performance computing (HPC) architectures.
Daniel Langr, Ivan Simecek, Tomás Dytrych
FedCSIS1
2016 Efficient parallel evaluation of block properties of sparse matrices
abstract
Many storage formats for sparse matrices have been developed.Majority of these formats can be parametrized, so the algorithm for finding optimal parameters is crucial.For overall efficiency, it is important to reduce the execution time of this preprocessing.In this paper, we propose a new algorithm for the determination of the number of nonzero blocks of the given size in a sparse matrix.The proposed algorithm requires relatively a small amount of auxiliary memory.Our approach is based on the Morton reordering and bitwise manipulations.We also present a parallel (multithreaded) version and evaluate its performance and space complexity.
Ivan Simecek, Daniel Langr
FedCSIS2
2016 Evaluation Criteria for Sparse Matrix Storage Formats
abstract
When authors present new storage formats for sparse matrices, they usually focus mainly on a single evaluation criterion, which is the performance of sparse matrix-vector multiplication (SpMV) in FLOPS. Though such an evaluation is essential, it does not allow to directly compare the presented format with its competitors. Moreover, in case that matrices are within an HPC application constructed in different formats, this criterion alone is not sufficient for the key decision whether or not to convert them into the presented format for the SpMV-based application phase. We establish ten evaluation criteria for sparse matrix storage formats, discuss their advantages and disadvantages, and provide general suggestions for format authors/evaluators to make their work more valuable for the HPC community.
Daniel Langr, Pavel Tvrdík
IEEE Trans. Parallel Distributed Syst.1
2014 Algorithm 947: Paraperm - Parallel Generation of Random Permutations with MPI
abstract
An algorithm for parallel generation of a random permutation of a large set of distinct integers is presented. This algorithm is designed for massively parallel systems with distributed memory architectures and the MPI-based runtime environments. Scalability of the algorithm is analyzed according to the memory and communication requirements. An implementation of the algorithm in a form of a software library based on the C++ programming language and the MPI application programming interface is further provided. Finally, performed experiments are described and their results discussed. The biggest of these experiments resulted in a generation of a random permutation of 2 41 integers in slightly more than four minutes using 131072 CPU cores.
Daniel Langr, Pavel Tvrdík, Tomás Dytrych, Jerry P. Draayer
ACM Trans. Math. Softw.1
2013 Storing Sparse Matrices to Files in the Adaptive-Blocking Hierarchical Storage Format
Daniel Langr, Ivan Simecek, Pavel Tvrdík
FedCSIS1
2012 Adaptive-Blocking Hierarchical Storage Format for Sparse Matrices
Daniel Langr, Ivan Simecek, Pavel Tvrdík, Tomás Dytrych, Jerry P. Draayer
FedCSIS1
2005 Clondike: Linux cluster of non-dedicated workstations
abstract
Clusters of workstations are a promising platform for high-performance computing. Most of such clusters are built of dedicated workstations that cannot be used for other purposes. Even though these clusters offer good price/performance ratio, their costs are not negligible. On the other hand, there exist many idle workstations connected via computer networks. The idea of exploiting such idle resources is obvious. In this paper, we describe an architecture of clusters made of non-dedicated idle workstations. These clusters aim at providing a single-system-image Linux environment. The main innovative feature is that a cluster administration is separate from administration of individual workstations. We discuss issues of security, availability, performance, and administration of such clusters. We also describe a pilot implementation and give promising experimental results.
Martin Kacer, Daniel Langr, Pavel Tvrdík
CCGRID2