EDBT 2026 Demo / reviewers in the wild / expert
Daniel Mlakar
dblp:218/5849
· DBLP profile ↗
11ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-4500-0325ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 30% High-performance computing · 24% Memory systems · 19% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU memory management |
0.8 | 2 | 2021 | Are dynamic memory managers on GPUs slow?: a survey and benchmarks · PPoPP 2021 faimGraph: high performance management of fully-dynamic graphs under tight memory constraints on the GPU · SC 2018 |
Performance modeling and evaluation
benchmarking |
0.5 | 1 | 2021 | Are dynamic memory managers on GPUs slow?: a survey and benchmarks · PPoPP 2021 |
Memory systems › memory management › memory allocation
dynamic memory allocation |
0.5 | 1 | 2021 | Are dynamic memory managers on GPUs slow?: a survey and benchmarks · PPoPP 2021 |
Processor architecture and microarchitecture › dynamic optimization
dynamic parameter tuning |
0.4 | 1 | 2020 | spECK: accelerating GPU sparse matrix-matrix multiplication through lightweight analysis · PPoPP 2020 |
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
sparse general matrix-matrix multiplication |
0.4 | 1 | 2020 | spECK: accelerating GPU sparse matrix-matrix multiplication through lightweight analysis · PPoPP 2020 |
GPUs and heterogeneous computing › GPU computing › GPU sparse computation
GPU sparse linear algebra |
0.4 | 1 | 2019 | Adaptive sparse matrix-matrix multiplication on the GPU · PPoPP 2019 |
High-performance computing
sparse linear algebra |
0.4 | 1 | 2019 | Adaptive sparse matrix-matrix multiplication on the GPU · PPoPP 2019 |
High-performance computing › sparse linear algebra
sparse matrix multiplication |
0.4 | 1 | 2019 | Adaptive sparse matrix-matrix multiplication on the GPU · PPoPP 2019 |
High-performance computing › sparse linear algebra › sparse matrix multiplication
SpGEMM |
0.4 | 1 | 2019 | Adaptive sparse matrix-matrix multiplication on the GPU · PPoPP 2019 |
Geometric modeling and processing › mesh generation
surface meshing |
0.3 | 1 | 2018 | Layered fields for natural tessellations on surfaces · ACM Trans. Graph. 2018 |
GPUs and heterogeneous computing
GPU graph processing |
0.3 | 1 | 2018 | faimGraph: high performance management of fully-dynamic graphs under tight memory constraints on the GPU · SC 2018 |
Memory systems
memory management |
0.3 | 1 | 2018 | faimGraph: high performance management of fully-dynamic graphs under tight memory constraints on the GPU · SC 2018 |
High-performance computing › scientific computing
scientific computing kernels |
0.1 | 1 | 2019 | Adaptive sparse matrix-matrix multiplication on the GPU · PPoPP 2019 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.5benchmarking · 0.5lightweight matrix analysis · 0.4dynamic algorithm selection · 0.4GPU vectorization · 0.4graph data structure design · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | End-to-End Compressed Meshlet RenderingabstractAbstract In this paper, we study rendering of end‐to‐end compressed triangle meshes using modern GPU techniques, in particular, mesh shaders. Our approach allows us to keep unstructured triangle meshes in GPU memory in compressed form and decompress them in shader code just in time for rasterization. Typical previous approaches use a compressed mesh format only for persistent storage and streaming, but must decompress it into GPU memory before submitting it to rendering. In contrast, our approach uses an identical compressed format in both storage and GPU memory. Hence, our compression method effectively reduces the in‐memory requirements of huge triangular meshes and avoids any waiting times on streaming geometry induced by the need for a decompression stage on the CPU. End‐to‐end compression also means that scenes with more geometric detail than previously possible can be made fully resident in GPU memory. Our approach is based on a novel decomposition of meshes into meshlets,i.e. disjoint primitive groups that are compressed individually. Decompression using a mesh shader allows de facto random access on the primitive level, which is important for applications such as selective streaming and fine‐grained visibility computation. We compare our approach to multiple commonly used compressed meshlet formats in terms of required memory and rendering times. The results imply that our approach reduces the required CPU–GPU memory bandwidth, a frequent bottleneck in out‐of‐core rendering. Daniel Mlakar, Markus Steinberger, Dieter Schmalstieg |
Comput. Graph. Forum | 1 |
| 2023 | A Variational Loop Shrinking Analogy for Handle and Tunnel Detection and Reeb Graph Construction on SurfacesabstractAbstract The humble loop shrinking property played a central role in the inception of modern topology but it has been eclipsed by more abstract algebraic formalisms. This is particularly true in the context of detecting relevant non‐contractible loops on surfaces where elaborate homological and/or graph theoretical constructs are favored in algorithmic solutions. In this work, we devise a variational analogy to the loop shrinking property and show that it yields a simple, intuitive, yet powerful solution allowing a streamlined treatment of the problem of handle and tunnel loop detection. Our formalization tracks the evolution of a diffusion front randomly initiated on a single location on the surface. Capitalizing on a diffuse interface representation combined with a set of rules for concurrent front interactions, we develop a dynamic data structure for tracking the evolution on the surface encoded as a sparse matrix which serves for performing both diffusion numerics and loop detection and acts as the workhorse of our fully parallel implementation. The substantiated results suggest our approach outperforms state of the art and robustly copes with highly detailed geometric models. As a byproduct, our approach can be used to construct Reeb graphs by diffusion thus avoiding commonly encountered issues when using Morse functions. Alexander Weinrauch, Daniel Mlakar, Hans-Peter Seidel, Markus Steinberger, Rhaleb Zayer |
Comput. Graph. Forum | 2 |
| 2021 | Speculative Parallel Reverse Cuthill-McKee Reordering on Multi- and Many-core ArchitecturesabstractBandwidth reduction of sparse matrices is used to reduce fill-in of linear solvers and to increase performance of other sparse matrix operations, e.g., sparse matrix vector multiplication in iterative solvers. To compute a bandwidth reducing permutation, Reverse Cuthill-McKee (RCM) reordering is often applied, which is challenging to parallelize, as its core is inherently serial. As many-core architectures, like the GPU, offer subpar single-threading performance and are typically only connected to high-performance CPU cores via a slow memory bus, neither computing RCM on the GPU nor moving the data to the CPU are viable options. Nevertheless, reordering matrices, potentially multiple times in-between operations, might be essential for high throughput. Still, to the best of our knowledge, we are the first to propose an RCM implementation that can execute on multicore CPUs and many-core GPUs alike, moving the computation to the data rather than vice versa.Our algorithm parallelizes RCM into mostly independent batches of nodes. For every batch, a single CPU-thread/a GPU thread-block speculatively discovers child nodes and sorts them according to the RCM algorithm. Before writing their permutation, we re-evaluate the discovery and build new batches. To increase parallelism and reduce dependencies, we create a signaling chain along successive batches and introduce early signaling conditions. In combination with a parallel work queue, new batches are started in order and the resulting RCM permutation is identical to the ground-truth single-threaded algorithm.We propose the first RCM implementation that runs on the GPU. It achieves several orders of magnitude speed-up over NVIDIA's single-threaded cuSolver RCM implementation and is significantly faster than previous parallel CPU approaches. Our results are especially significant for many-core architectures, as it is now possible to include RCM reordering into sequences of sparse matrix operations without major performance loss. Daniel Mlakar, Mathias Parger, Markus Steinberger |
IPDPS | 1 |
| 2021 | Are dynamic memory managers on GPUs slow?: a survey and benchmarksabstractDynamic memory management on GPUs is generally understood to be a challenging topic. On current GPUs, hundreds of thousands of threads might concurrently allocate new memory or free previously allocated memory. This leads to problems with thread contention, synchronization overhead and fragmentation. Various approaches have been proposed in the last ten years and we set out to evaluate them on a level playing field on modern hardware to answer the question, if dynamic memory managers are as slow as commonly thought of. In this survey paper, we provide a consistent framework to evaluate all publicly available memory managers in a large set of scenarios. We summarize each approach and thoroughly evaluate allocation performance (thread-based as well as warp-based), and look at performance scaling, fragmentation and real-world performance considering a synthetic workload as well as updating dynamic graphs. We discuss the strengths and weaknesses of each approach and provide guidelines for the respective best usage scenario. We provide a unified interface to integrate any of the tested memory managers into an application and switch between them for benchmarking purposes. Given our results, we can dispel some of the dread associated with dynamic memory managers on the GPU. Mathias Parger, Daniel Mlakar, Markus Steinberger |
PPoPP | 3 |
| 2020 | Ouroboros: virtualized queues for dynamic memory management on GPUsabstractDynamic memory allocation on a single instruction, multiple threads architecture, like the Graphics Processing Unit (GPU), is challenging and implementation guidelines caution against it. Data structures must rise to the challenge of thousands of concurrently active threads trying to allocate memory. Efficient queueing structures have been used in the past to allow for simple allocation and reuse of memory directly on the GPU but do not scale well to different allocation sizes, as each requires its own queue. Daniel Mlakar, Mathias Parger, Markus Steinberger |
ICS | 2 |
| 2020 | spECK: accelerating GPU sparse matrix-matrix multiplication through lightweight analysisabstractSparse general matrix-matrix multiplication on GPUs is challenging due to the varying sparsity patterns of sparse matrices. Existing solutions achieve good performance for certain types of matrices, but fail to accelerate all kinds of matrices in the same manner. Our approach combines multiple strategies with dynamic parameter selection to dynamically choose and tune the best fitting algorithm for each row of the matrix. This choice is supported by a lightweight, multi-level matrix analysis, which carefully balances analysis cost and expected performance gains. Our evaluation on thousands of matrices with various characteristics shows that we outperform all currently available solutions in 79% over all matrices with >15k products and that we achieve the second best performance in 15%. For these matrices, our solution is on average 83% faster than the second best approach and up to 25X faster than other state-of-the-art GPU implementations. Using our approach, applications can expect great performance independent of the matrices they work on. Mathias Parger, Daniel Mlakar, Markus Steinberger |
PPoPP | 3 |
| 2020 | Subdivision-Specialized Linear Algebra Kernels for Static and Dynamic Mesh Connectivity on the GPUabstractAbstract Subdivision surfaces have become an invaluable asset in production environments. While progress over the last years has allowed the use of graphics hardware to meet performance demands during animation and rendering, high‐performance is limited to immutable mesh connectivity scenarios. Motivated by recent progress in mesh data structures, we show how the complete Catmull‐Clark subdivision scheme can be abstracted in the language of linear algebra. While this high‐level formulation allows for a fully parallel implementation with significant performance gains, the underlying algebraic operations require further specialization for modern parallel hardware. Integrating domain knowledge about the mesh matrix data structure, we replace costly general linear algebra operations like matrix‐matrix multiplication by specialized kernels. By further considering innate properties of Catmull‐Clark subdivision, like the quad‐only structure after refinement, we achieve an additional order of magnitude in performance and significantly reduce memory footprints. Our approach can be adapted seamlessly for different use cases, such as regular subdivision of dynamic meshes, fast evaluation for immutable topology and feature‐adaptive subdivision for efficient rendering of animated models. In this way, patchwork solutions are avoided in favor of a streamlined solution with consistent performance gains throughout the production pipeline. The versatility of the sparse matrix linear algebra abstraction underlying our work is further demonstrated by extension to other schemes such as and Loop subdivision. Daniel Mlakar, Pascal Stadlbauer, Hans-Peter Seidel, Markus Steinberger, Rhaleb Zayer |
Comput. Graph. Forum | 1 |
| 2020 | Interactive Modeling of Cellular Structures on Surfaces with Application to Additive ManufacturingabstractAbstract The rich and evocative patterns of natural tessellations endow them with an unmistakable artistic appeal and structural properties which are echoed across design, production, and manufacturing. Unfortunately, interactive control of such patterns‐as modeled by Voronoi diagrams, is limited to the simple two dimensional case and does not extend well tofreeform surfaces. We present an approach for direct modeling and editing of such cellular structures on surface meshes. The overall modeling experience is driven by a set of editing primitives which are efficiently implemented on graphics hardware. We feature a novel application for 3D printing on modern support‐free additive manufacturing platforms. Our method decomposes the input surface into a cellular skeletal structure which hosts a set of overlay shells. In this way, material saving can be channeled to the shells while structural stability is channeled to the skeleton. To accommodate the available printer build volume, the cellular structure can be further split into moderately sized parts. Together with shells, they can be conveniently packed to save on production time. The assembly of the printed parts is streamlined by a part numbering scheme which respects the geometric layout of the input model. Pascal Stadlbauer, Daniel Mlakar, Hans-Peter Seidel, Markus Steinberger, Rhaleb Zayer |
Comput. Graph. Forum | 2 |
| 2019 | Adaptive sparse matrix-matrix multiplication on the GPUabstractIn the ongoing efforts targeting the vectorization of linear algebra primitives, sparse matrix-matrix multiplication (SpGEMM) has received considerably less attention than sparse Matrix-Vector multiplication (SpMV). While both are equally important, this disparity can be attributed mainly to the additional formidable challenges raised by SpGEMM. Daniel Mlakar, Rhaleb Zayer, Hans-Peter Seidel, Markus Steinberger |
PPoPP | 2 |
| 2018 | faimGraph: high performance management of fully-dynamic graphs under tight memory constraints on the GPU
Daniel Mlakar, Rhaleb Zayer, Hans-Peter Seidel, Markus Steinberger |
SC | 2 |
| 2018 | Layered fields for natural tessellations on surfaces
Rhaleb Zayer, Daniel Mlakar, Markus Steinberger, Hans-Peter Seidel |
ACM Trans. Graph. | 2 |