VLDB 2026 Research / reviewers in the wild / expert
Lukas Hübner
dblp:246/0570
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-9213-7597ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 78% Distributed systems · 15% High-performance computing · 7% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel computing |
1.0 | 1 | 2026 | Bit-reproducible parallel phylogenetic tree inference · Bioinform. 2026 |
Bioinformatics and computational biology › phylogenetics › phylogenetic inference
maximum likelihood estimation |
0.8 | 2 | 2026 | Exploring parallel MPI fault tolerance mechanisms for phylogenetic inference with RAxML-NG · Bioinform. 2021 Bit-reproducible parallel phylogenetic tree inference · Bioinform. 2026 |
Parallel and multicore computing
MPI |
0.8 | 1 | 2024 | KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI · SC 2024 |
Parallel and multicore computing
parallel programming models |
0.8 | 1 | 2024 | KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI · SC 2024 |
Bioinformatics and computational biology › phylogenetics
phylogenetic inference |
0.5 | 1 | 2021 | Exploring parallel MPI fault tolerance mechanisms for phylogenetic inference with RAxML-NG · Bioinform. 2021 |
Distributed systems
fault tolerance |
0.5 | 1 | 2021 | Exploring parallel MPI fault tolerance mechanisms for phylogenetic inference with RAxML-NG · Bioinform. 2021 |
Bioinformatics and computational biology
phylogenetics |
0.3 | 1 | 2026 | Bit-reproducible parallel phylogenetic tree inference · Bioinform. 2026 |
High-performance computing
scientific computing systems |
0.2 | 1 | 2024 | KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI · SC 2024 |
Methods — techniques the papers use, named apart from their topics
reprored reduction algorithm · 2.0MPI · 2.0message passing interface · 1.0failure simulation · 1.0checkpointing · 1.0type system · 0.8template metaprogramming · 0.8named parameters · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bit-reproducible parallel phylogenetic tree inferenceabstractMOTIVATION: Phylogenetic trees describe the evolutionary history among biological species based on their genomic data. Maximum likelihood (ML) based phylogenetic inference tools search for the tree and evolutionary model that best explain the observed genomic data. Given the independence of likelihood score calculations between different genomic sites, parallel computation is commonly deployed. This is followed by a parallel summation over the per-site scores to obtain the overall likelihood score of the tree. However, basic arithmetic operations on IEEE 754 floating-point numbers, such as addition and multiplication, inherently introduce rounding errors. Consequently, the order by which floating-point operations are executed affects the exact resulting likelihood value since these operations are not associative. Moreover, parallel reduction algorithms in numerical codes re-associate operations as a function of the core count and cluster network topology, inducing different round-off errors. These low-level deviations can cause heuristic searches to diverge and induce high-level result discrepancies (e.g. yield topologically distinct phylogenies). This effect has also been observed in multiple scientific fields beyond phylogenetics. RESULTS: We observe that varying the degree of parallelism results in diverging phylogenetic tree searches (high-level results) for over 31% out of 10 179 empirical datasets. More importantly, 8% of these diverging datasets yield trees that are statistically significantly worse than the best-known ML tree for the dataset (AU-test, P < .05). To alleviate this, we develop a variant of the widely used phylogenetic inference tool RAxML-NG, which does yield bit-reproducible results under varying core-counts, with a slowdown of only 0%-12.7% (median 0.8%) on up to 768 cores. For this, we introduce the ReproRed reduction algorithm, which yields bit-identical results under varying core-counts, by maintaining a fixed operation order that is independent of the communication pattern. ReproRed is thus applicable to all associative reduction operations-in contrast to competitors, which are confined to summation. Our ReproRed reduction algorithm only exchanges the theoretical minimum number of messages, overlaps communication with computation, and utilizes fast base-cases for local reductions. ReproRed is able to all-reduce (via a subsequent broadcast) 4.1×106 operands across 48-768 cores in 19.7-48.61 μs, thereby exhibiting a slowdown of 13%-93% over a non-reproducible all-reduce algorithm. ReproRed outperforms the state-of-the-art reproducible all-reduction algorithm ReproBLAS (offers summation only) beyond 10 000 elements per core. In summary, we re-assess non-reproducibility in parallel phylogenetic inference, present the first bit-reproducible parallel phylogenetic inference tool, as well as introduce a general algorithm and open-source code for conducting reproducible associative parallel reduction operations. AVAILABILITY AND IMPLEMENTATION: ReproRed: https://doi.org/10.5281/zenodo.15004918 (LGPL)-Reproducible RAxML-NG version https://doi.org/10.5281/zenodo.15017407 (GPL). Christoph Stelz, Lukas Hübner, Alexandros Stamatakis |
Bioinform. | 2 |
| 2024 | KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPIabstractThe Message-Passing Interface (MPI) and C++ form the backbone of high-performance computing, but MPI only provides $\mathbf{C}$ and Fortran bindings. While this offers great language interoperability, high-level programming languages like C++ make software development quicker and less error-prone.We propose novel $\mathrm{C}_{++}$language bindings that cover all abstraction levels from low-level MPI calls to convenient STL-style bindings, where most parameters are inferred from a small subset of parameters, by bringing named parameters to C++. This enables rapid prototyping and fine-tuning runtime behavior and memory management. A flexible type system and additional safety guarantees help to prevent programming errors.By exploiting C++’s template metaprogramming capabilities, this has (near) zero overhead, as only required code paths are generated at compile time.We demonstrate that our library is a strong foundation for a future distributed standard library using multiple application benchmarks, ranging from text-book sorting algorithms to phylogenetic interference. Tim Niklas Uhl, Matthias Schimek, Lukas Hübner, Demian Hespe, Florian Kurpicz, Daniel Seemaier, Christoph Stelz, Peter Sanders 0001 |
SC | 3 |
| 2024 | Brief Announcement: (Near) Zero-Overhead C++ Bindings for MPIabstractThe Message-Passing Interface (MPI) and C++ form the backbone of high-performance computing and algorithmic research in the field of distributed-memory computing, but MPI only provides C and Fortran bindings.This provides good language interoperability, but higher-level programming languages make development quicker and less error-prone.We propose novel C++ language bindings designed to cover the whole range of abstraction levels from low-level MPI calls to convenient STL-style bindings, where most parameters are inferred from a small subset of the full parameter set.This allows for both rapid prototyping and fine-tuning of distributed code with predictable runtime behavior and memory management.Using template-metaprogramming, only code paths required for computing missing parameters are generated at compile time, which results in (near) zero-overhead bindings. Demian Hespe, Lukas Hübner, Florian Kurpicz, Peter Sanders 0001, Matthias Schimek, Daniel Seemaier, Tim Niklas Uhl |
SPAA | 2 |
| 2024 | Memoization on Shared Subtrees Accelerates Computations on Genealogical Forests
Lukas Hübner, Alexandros Stamatakis |
WABI | 1 |
| 2021 | Exploring parallel MPI fault tolerance mechanisms for phylogenetic inference with RAxML-NGabstractMOTIVATION: Phylogenetic trees are now routinely inferred on large scale high performance computing systems with thousands of cores as the parallel scalability of phylogenetic inference tools has improved over the past years to cope with the molecular data avalanche. Thus, the parallel fault tolerance of phylogenetic inference tools has become a relevant challenge. To this end, we explore parallel fault tolerance mechanisms and algorithms, the software modifications required and the performance penalties induced via enabling parallel fault tolerance by example of RAxML-NG, the successor of the widely used RAxML tool for maximum likelihood-based phylogenetic tree inference. RESULTS: We find that the slowdown induced by the necessary additional recovery mechanisms in RAxML-NG is on average 1.00 ± 0.04. The overall slowdown by using these recovery mechanisms in conjunction with a fault-tolerant Message Passing Interface implementation amounts to on average 1.7 ± 0.6 for large empirical datasets. Via failure simulations, we show that RAxML-NG can successfully recover from multiple simultaneous failures, subsequent failures, failures during recovery and failures during checkpointing. Recoveries are automatic and transparent to the user. AVAILABILITY AND IMPLEMENTATION: The modified fault-tolerant RAxML-NG code is available under GNU GPL at https://github.com/lukashuebner/ft-raxml-ng. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lukas Hübner, Alexey M. Kozlov, Demian Hespe, Peter Sanders 0001, Alexandros Stamatakis |
Bioinform. | 1 |