EDBT 2026 Demo / reviewers in the wild / expert
Christoph Niethammer
dblp:94/10669
· DBLP profile ↗
8ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-3840-1016ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenMPI: Cluster Scalable Variant Calling for Short/Long Reads Sequencing DataabstractRapid technological advancements in sequencing technologies allow producing cost effective and high volume sequencing data. Processing this data for real-time clinical diagnosis is potentially time-consuming if done on a single computing node. This work presents a complete variant calling workflow, implemented using the Message Passing Interface (MPI) to leverage the benefits of high bandwidth interconnects. This solution (GenMPI) is portable and flexible, meaning it can be deployed to any private or public cluster/cloud infrastructure. Any alignment or variant calling application can be used with minimal adaptation. To achieve high performance, compressed input data can be streamed in parallel to alignment applications while uncompressed data can use internal file seek functionality to eliminate the bottleneck of streaming input data from a single node. Alignment output can be directly stored in multiple chromosome-specific SAM files or a single SAM file. After alignment, a distributed queue using MPI RMA (Remote Memory Access) atomic operations is created for sorting, indexing, marking of duplicates (if necessary) and variant calling applications. We ensure the accuracy of variants as compared to the original single node methods. We also show that for 300x coverage data, alignment scales almost linearly up to 64 nodes (8192 CPU cores). Overall, this work outperforms existing Big Data based workflows by a factor of two and is almost 20% faster than other MPI-based implementations for alignment without any extra memory overheads. Sorting, indexing, duplicate removal and variant calling is also scalable up to 8 nodes cluster. For pair-end short-reads (Illumina) data, we integrated the BWA-MEM aligner and three variant callers (GATK HaplotypeCaller, DeepVariant and Octopus), while for long-reads data, we integrated the Minimap2 aligner and three different variant callers (DeepVariant, DeepVariant with WhatsHap for phasing (PacBio) and Clair3 (ONT)). Joseph Schuchart, Zaid Al-Ars, Christoph Niethammer, José Gracia, H. Peter Hofstee |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Leveraging the Compute Power of Two HPC Systems for Higher-Dimensional Grid-Based Simulations with the Widely-Distributed Sparse Grid Combination TechniqueabstractGrid-based simulations of hot fusion plasmas are often severely limited by computational and memory resources; the grids live in four- to six-dimensional space and thus suffer the curse of dimensionality. However, high resolutions are required to fully capture the physics of interest. The sparse grid combination technique is a multi-scale method in which many anisotropically coarse resolved grids are used to approximate a fine-scale solution---and it alleviates the curse of dimensionality. Theresa Pollinger, Alexander Van Craen, Christoph Niethammer, Marcel Breyer, Dirk Pflüger |
SC | 3 |
| 2023 | The EU Center of Excellence for Exascale in Solid Earth (ChEESE): Implementation, results, and roadmap for the second phaseabstractThe EU Center of Excellence for Exascale in Solid Earth (ChEESE) develops exascale transition capabilities in the domain of Solid Earth, an area of geophysics rich in computational challenges embracing different approaches to exascale (capability, capacity, and urgent computing). The first implementation phase of the project (ChEESE-1P; 2018–2022) addressed scientific and technical computational challenges in seismology, tsunami science, volcanology, and magnetohydrodynamics, in order to understand the phenomena, anticipate the impact of natural disasters, and contribute to risk management. The project initiated the optimisation of 10 community flagship codes for the upcoming exascale systems and implemented 12 Pilot Demonstrators that combine the flagship codes with dedicated workflows in order to address the underlying capability and capacity computational challenges. Pilot Demonstrators reaching more mature Technology Readiness Levels (TRLs) were further enabled in operational service environments on critical aspects of geohazards such as long-term and short-term probabilistic hazard assessment, urgent computing, and early warning and probabilistic forecasting. Partnership and service co-design with members of the project Industry and User Board (IUB) leveraged the uptake of results across multiple research institutions, academia, industry, and public governance bodies (e.g. civil protection agencies). This article summarises the implementation strategy and the results from ChEESE-1P, outlining also the underpinning concepts and the roadmap for the on-going second project implementation phase (ChEESE-2P; 2023–2026). Arnau Folch, Claudia Abril, Michael Afanasiev, Giorgio Amati, Michael Bader, Rosa M. Badia, Hafize B. Bayraktar, Sara Barsotti, Roberto Basili 0002, Fabrizio Bernardi, Christian Boehm, Beatriz Brizuela, Federico Brogi, Eduardo Cabrera, Emanuele Casarotti, Manuel Jesús Castro Díaz, Matteo Cerminara, Antonella Cirella, Alexey Cheptsov, Javier Conejero, Antonio Costa 0002, Marc de la Asunción, Josep de la Puente, Marco Djuric, Ravil Dorozhinskii, Gabriela Espinosa, Tomaso Esposti Ongaro, Joan Farnós, Nathalie Favretto-Cristini, Andreas Fichtner, Alexandre Fournier, Alice-Agnes Gabriel, Jean-Matthieu Gallard, Steven J. Gibbons, Sylfest Glimsdal, José Manuel González-Vida, José Gracia, Rose Gregorio, Natalia Gutiérrez, Benedikt Halldorsson, Okba Hamitou, Guillaume Houzeaux, Stephan Jaure, Mouloud Kessar, Lukas Krenz, Lion Krischer, Soline Laforet, Piero Lanucara, Bo Li 0147, Maria Concetta Lorenzino, Stefano Lorito, Finn Løvholt, Giovanni Macedonio, Jorge Macías Sánchez, Guillermo Marin, Beatriz Martínez Montesinos, Leonardo Mingari, Geneviève Moguilny, Vadim Montellier, Marisol Monterrubio Velasco, Georges-Emmanuel Moulard, Masaru Nagaso, Massimo Nazaria, Christoph Niethammer, Federica Pardini, Marta Pienkowska, Luca Pizzimenti, Natalia Poiata, Leonhard Rannabauer, Otilio Rojas, Juan Esteban Rodriguez, Fabrizio Romano, Oleksandr Rudyy, Vittorio Ruggiero, Philipp Samfass, Carlos Sánchez-Linares, Sabrina Sanchez, Laura Sandri, Antonio Scala, Nathanaël Schaeffer, Joseph Schuchart, Jacopo Selva, Amadine Sergeant, Angela Stallone, Matteo Taroni, Solvi Thrastarson, Manuel Titos, Nadia Tonelllo, Roberto Tonini, Thomas Ulrich, Jean-Pierre Vilotte, Malte Vöge, Manuela Volpe, Sara Aniko Wirp, Uwe Wössner |
Future Gener. Comput. Syst. | 64 |
| 2021 | Callback-based completion notification using MPI Continuations
Joseph Schuchart, Philipp Samfass, Christoph Niethammer, José Gracia, George Bosilca |
Parallel Comput. | 3 |
| 2020 | Fibers are not (P)Threads: The Case for Loose Coupling of Asynchronous Programming Models and MPI Through ContinuationsabstractAsynchronous programming models (APM) are gaining more and more traction, allowing applications to expose the available concurrency to a runtime system tasked with coordinating the execution. While MPI has long provided support for multi-threaded communication and non-blocking operations, it falls short of adequately supporting APMs as correctly and efficiently handling MPI communication in different models is still a challenge. Meanwhile, new low-level implementations of light-weight, cooperatively scheduled execution contexts (fibers, aka user-level threads (ULT)) are meant to serve as a basis for higher-level APMs and their integration in MPI implementations has been proposed as a replacement for traditional POSIX thread support to alleviate these challenges. Joseph Schuchart, Christoph Niethammer, José Gracia |
EuroMPI | 2 |
| 2019 | An MPI interface for application and hardware aware cartesian topology optimizationabstractMany scientific applications perform computations on a Cartesian grid. The common approach for the parallelization of these applications with MPI is domain decomposition. To help developers with the mapping of MPI processes to subdomains, the MPI standard provides the concept of process topologies. However, the current interface causes problems and requires too much care in its usage: MPI_Dims_create does not take into account the application topology and most implementations of MPI_Cart_create do not consider the underlying network topology and node architecture. To overcome these shortcomings, we defined a new interface that includes application-aware weights to address the communication needs of grid-based applications. The new interface provides a hardware-aware factorization of the processes together with an optimized process mapping onto the underlying hardware resources. The paper describes the underlying implementation, which uses a new multi-level factorization and decomposition approach minimizing slow inter-node communication. Benchmark results show the significant performance gains on multi node NUMA systems. Christoph Niethammer, Rolf Rabenseifner |
EuroMPI | 1 |
| 2012 | Hybrid MPI/StarSs - A Case StudyabstractHybrid parallel programming models combining distributed and shared memory paradigms are well established in high-performance computing. The classical prototype of hybrid programming in HPC is MPI/OpenMP, but many other combinations are being investigated. Recently, the data-dependency driven, task parallel model for shared memory parallelisation named StarSs has been suggested for usage in combination with MPI. In this paper we apply hybrid MPI/StarSs to a Lattice-Boltzmann code. In particular, we present the hybrid programming model, the benefits we expect, the challenges in porting, and finally a comparison of the performance of MPI/StarSs hybrid, MPI/OpenMP hybrid and the original MPI-only versions of the same code. José Gracia, Christoph Niethammer, Manuel Hasert, Steffen Brinkmann, Rainer Keller, Colin W. Glass |
ISPA | 2 |
| 2012 | Avoiding Serialization Effects in Data / Dependency Aware Task Parallel Algorithms for Spatial DecompositionabstractSpatial decomposition is a popular basis for parallelising code. Cast in the frame of task parallelism, calculations on a spatial domain can be treated as a task. If neighbouring domains interact and share results, access to the specific data needs to be synchronized to avoid race conditions. This is the case for a variety of applications, like most molecular dynamics and many computational fluid dynamics codes. Here we present an unexpected problem which can occur in dependency-driven task parallelization models like StarSs: the tasks accessing a specific spatial domain are treated as interdependent, as dependencies are detected automatically via memory addresses. Thus, the order in which tasks are generated will have a severe impact on the dependency tree. In the worst case, a complete serialization is reached and no two tasks can be calculated in parallel. We present the problem in detail based on an example from molecular dynamics, and introduce a theoretical framework to calculate the degree of serialization. Furthermore, we present strategies to avoid this unnecessary problem. We recommend treating these strategies as best practice when using dependency-driven task parallel programming models like StarSs on such scenarios. Christoph Niethammer, Colin W. Glass, José Gracia |
ISPA | 1 |