EDBT 2026 Demo / reviewers in the wild / expert
Michael Kuhn 0003
dblp:k/MichaelKuhn3
· DBLP profile ↗
16ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-8167-8574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Making MPI Collective Operations Visible: Understanding Their Utility and Algorithmic Insights
Anna-Lena Roth, David James, Michael Kuhn 0003, Dustin Frisch |
Euro-Par (1) | 3 |
| 2025 | NVM in Data Storage: A Post-Optane FutureabstractThe dynamic evolution of non-volatile memory (NVM) technologies from Read-Only Memory (ROM) to flash memory, and recent innovations in Magnetoresistive RAM (MRAM), Phase Change Memory (PCM), and Resistive RAM (ReRAM) signify a pivotal shift in data storage capabilities and applications. This progression, marked by enhancements in performance and density, fills the gap between Dynamic RAM (DRAM) and flash storage, meeting the demands of data-rich domains such as high-performance computing and databases. The integration of NVM like the formerly commercially available Optane Persistent Memory (PMem) has introduced a paradigm shift in storage system design, addressing challenges related to latency, data proximity, memory constraints, and energy efficiency which are critical in computing. Despite facing hurdles, such as performance asymmetry in Optane PMem, the potential of NVM in data storage systems has been shown with research focusing on speeding up indexing, small data accesses, and cache-less I/O paths. We review existing and emerging NVM technologies based on their physical properties and application in data storage. We underscore the significance of NVM in offering unique access characteristics that address the limitations of traditional storage devices, thereby extending the data storage landscape. We discuss the programming challenges and hardware-software strategies crucial for NVM’s seamless integration and widespread industry adoption. Additionally, we discuss the discontinuation of Optane, the ongoing challenges and the need for storage system designs to adapt to non-block based storage to leverage NVM’s performance benefits. We conclude that while advancements in NVM promise improved performance and endurance, practical availability depends on addressing challenges such as density, heat management, and reliability. Support in the existing software ecosystem is also crucial for its widespread adoption. Furthermore, Compute Express Link (CXL) is highlighted as a significant development that can streamline the memory hierarchy and support the adoption of NVM, enabling more flexible designs and effective usage of storage media in next-generation data storage systems. Overall, CXL offers a promising method to include NVM in future systems. However, in the meantime, data storage systems will rely on NVMe SSDs due to their performance and availability. But, we encourage further research into integrating NVM devices due to data management constraints with conventional block-based data management and advocate for the availability of NVM-ready software to users. Sajad Karim, Johannes Wünsche, Michael Kuhn 0003, Gunter Saake, David Broneske |
ACM Trans. Storage | 3 |
| 2024 | Towards End-to-End Compression in LustreabstractScientific applications generate massive amounts of data, posing storage limitations and network traffic challenges. While scientists struggle with the usage of application-side compression and parallel I/O, we design a transparent feature. Our Lustre-based prototype automatically applies lossless compression, offering flexibility in compression-related decisions to minimize computational costs and optimize application performance. We outline the challenges posed by our prototype and illustrate through a comprehensive assessment how integrating Lustre and ZFS as backend solutions provides the essential elements for performance and scalability: specifically, asynchronous operations and parallel processing of compression. In our evaluation, we illustrate the interaction of different buffer levels within a distributed system. Additionally, we showcase how the I/O pattern, hardware setup, and various system software optimizations can impact overall performance and influence the choice of compression strategy. Anna Fuchs, Jannek Squar, Michael Kuhn 0003 |
ISPDC | 3 |
| 2024 | Ensemble-Based System Benchmarking for HPCabstractIn HPC supercomputers, the CPU, memory, network and storage play a critical role in application performance. Established benchmarks measure the theoretical peak performance of these components, but storage benchmarks often focus solely on I/O and lack realism. To address this, we present numio, a benchmark that simulates and evaluates overlapping compute, communication and I/O phases. Furthermore, we present a novel ensemble-driven system benchmarking strategy. This approach involves running multiple benchmarks in parallel to analyse their interactions and assess the system’s ability to handle the workload. Using a real-world example, we demonstrate how this approach reveals performance issues in complex HPC systems that remain hidden when using traditional methods using isolated benchmarks on empty systems. Anna Fuchs, Jannek Squar, Michael Kuhn 0003 |
ISPDC | 3 |
| 2022 | Automated performance analysis tools framework for HPC programsabstractIn the high performance computing (HPC) field, computer and natural scientists work together to solve complex problems. While computer scientists try their best to meet the needs of domain experts, the latter ones still require deeper know-how in informatics to bring their science onto machines in an optimal way. Applications typically require significant optimisation and tuning efforts to harness the performance capabilities of HPC systems. A wide range of very different analyse tools are available for this purpose, but their use is time-consuming and thus also a financial burden. This work is the first to introduce an extensible framework which does not only provide a convenient graphical interface for HPC performance analysis tools, but also simplifies their usage extremely. On this path the foundation for automated scientific software analysis and optimisation workflows is layed, which is believed to become increasingly necessary. Maximilian Keiff, Frederic Voigt, Anna Fuchs, Michael Kuhn 0003, Jannek Squar, Thomas Ludwig 0002 |
KES | 4 |
| 2022 | Content queries and in-depth analysis on version-controlled softwareabstractWriting scientific code usually implies the need to coordinate and conflate the contributions of several scientific programmers. Using Git hosting services eases this process, because the hosting services offer many features, which assist in collaborated work on code. The well-established hosting service GitHub has seen continuous growth in terms of number of users, repositories and commits over the last few years; therefore it offers a large data source of scientific codes as well as social interaction of associated scientific programmers. We present a tool, which allows to easily search through relevant GitHub repositories and perform more advanced analyses, which cannot be conducted solely with the GitHub API. Our tool combines benefits from online as well as offline approaches to retrieve and analyse data to optimise time of execution and consumption of storage. We discuss possible use cases and demonstrate the tool's capabilities by investigating the popularity of OpenMP directives in the scientific community. Jannek Squar, Niclas Schroeter, Anna Fuchs, Michael Kuhn 0003, Thomas Ludwig 0002 |
KES | 4 |
| 2021 | Dissecting self-describing data formats to enable advanced querying of file metadataabstractIn times of continuously growing data sizes, performing insightful analysis is increasingly difficult. I/O libraries such as NetCDF and ADIOS2 offer options to manage additional metadata to make the data retrieval more efficient. However, queries on this metadata are difficult as it is currently stored inside the corresponding self-describing data formats. Kira Duwe, Michael Kuhn 0003 |
SYSTOR | 2 |
| 2020 | Compiler Assisted Source Transformation of OpenMP KernelsabstractMany scientific applications use OpenMP as a relatively easy and fast approach to utilise symmetric multiprocessor systems at their full capacity. However, scalability on shared memory systems is limited and thus distributed parallel computing is inevitable if the full potential through horizontal scaling shall be achieved. Additional software layers like MPI must be used, which require further knowledge on the scientific developers' side. This paper presents CATO, a tool prototype using LLVM and Clang, to transform existing OpenMP code to MPI; this enables distributed code execution while keeping OpenMP's relatively low barrier of entry. The main focus lies on increasing the maximum problem size, which a scientific application can work on; converting an intra-node problem into an inter-node problem makes it possible to overcome the limitation of memory of a single node. Our tool does not focus on improving the absolute runtime, even though it might improve it by e.g. introducing concurrency during the I/O phase; but we rather focus on increasing the maximal problem size and our benchmark of a stencil code shows promising results: The transformation preserves the speedup trend of the code to some extent. Another example demonstrates the capability to increase the maximum problem size while using additional compute nodes. Jannek Squar, Tim Jammer, Michael Blesel, Michael Kuhn 0003, Thomas Ludwig 0002 |
ISPDC | 4 |
| 2020 | Mission possible: Unify HPC and Big Data stacks towards application-defined blobs at the storage layer
Pierre Matri, Yevhen Alforov, Álvaro Brandón, María S. Pérez 0001, Alexandru Costan, Gabriel Antoniu, Michael Kuhn 0003, Philip H. Carns, Thomas Ludwig 0002 |
Future Gener. Comput. Syst. | 7 |
| 2018 | Towards Green Scientific Data Compression Through High-Level I/O InterfacesabstractEvery HPC system today has to cope with a deluge of data generated by scientific applications, simulations or large-scale experiments. The upscaling of supercomputer systems and infrastructures, generally results in a dramatic increase of their energy consumption. In this paper, we argue that techniques like data compression can lead to significant gains in terms of power efficiency by reducing both network and storage requirements. However, any data reduction is highly data specific and should comply with established requirements. Therefore, unsuitable or inappropriate compression strategy can utilize more resources and energy than necessary. To that end, we propose a novel methodology for achieving on-the-fly intelligent determination of energy efficient data reduction for a given data set by leveraging state-of-the-art compression algorithms and meta data at application-level I/O. We motivate our work by analyzing the energy and storage saving needs of data sets from real-life scientific HPC applications, and review the various lossless compression techniques that can be applied. We find that the resulting data reduction can decrease the data volume transferred and stored by as much as 80 % in some cases, consequently leading to significant savings in storage and networking costs. Yevhen Alforov, Thomas Ludwig 0002, Anastasiia Novikova, Michael Kuhn 0003, Julian M. Kunkel |
SBAC-PAD | 4 |
| 2017 | Could Blobs Fuel Storage-Based Convergence Between HPC and Big Data?abstractThe increasingly growing data sets processed on HPC platforms raise major challenges for the underlying storage layer. A promising alternative to POSIX-IO-compliant file systems are simpler blobs (binary large objects), or object storage systems. They offer lower overhead and better performance at the cost of largely unused features such as file hierarchies or permissions. Similarly, blobs are increasingly considered for replacing distributed file systems for big data analytics or as a base for storage abstractions like key-value stores or time-series databases. This growing interest in such object storage on HPC and big data platforms raises the question: Are blobs the right level of abstraction to enable storage-based convergence between HPC and Big Data? In this paper we take a first step towards answering the question by analyzing the applicability of blobs for both platforms. Pierre Matri, Yevhen Alforov, Álvaro Brandón, Michael Kuhn 0003, Philip H. Carns, Thomas Ludwig 0002 |
CLUSTER | 4 |
| 2016 | Analyzing the energy consumption of the storage data path
Pablo Llopis, Manuel F. Dolz, Francisco Javier García Blas, Florin Isaila, Mohammad Reza Heidari, Michael Kuhn 0003 |
J. Supercomput. | 6 |
| 2012 | Simulation-Aided Performance Evaluation of Server-Side Input/Output OptimizationsabstractThe performance of parallel distributed file systems suffers from many clients executing a large number of operations in parallel, because the I/O subsystem can be easily overwhelmed by the sheer amount of incoming I/O operations. Many optimizations exist that try to alleviate this problem. Client-side optimizations perform preprocessing to minimize the amount of work the file servers have to do. Server-side optimizations use server-internal knowledge to improve performance. The HD Trace framework contains components to simulate, trace and visualize applications. It is used as a test bed to evaluate optimizations that could later be implemented in real-life projects. This paper compares existing client-side optimizations and newly implemented server-side optimizations and evaluates their usefulness for I/O patterns commonly found in HPC. Server-directed I/O chooses the order of non-contiguous I/O operations and tries to aggregate as many operations as possible to decrease the load on the I/O subsystem and improve overall performance. The results show that server-side optimizations beat client-side optimizations in terms of performance for many use cases. Integrating such optimizations into parallel distributed file systems could alleviate the need for sophisticated client-side optimizations. Due to their additional knowledge of internal workflows server-side optimizations may be better suited to provide high performance in general. Michael Kuhn 0003, Julian M. Kunkel, Thomas Ludwig 0002 |
PDP | 1 |
| 2012 | A study on data deduplication in HPC storage systemsabstractDeduplication is a storage saving technique that is highly successful in enterprise backup environments. On a file system, a single data block might be stored multiple times across different files, for example, multiple versions of a file might exist that are mostly identical. With deduplication, this data replication is localized and redundancy is removed -- by storing data just once, all files that use identical regions refer to the same unique data. The most common approach splits file data into chunks and calculates a cryptographic fingerprint for each chunk. By checking if the fingerprint has already been stored, a chunk is classified as redundant or unique. Only unique chunks are stored. This paper presents the first study on the potential of data deduplication in HPC centers, which belong to the most demanding storage producers. We have quantitatively assessed this potential for capacity reduction for 4 data centers (BSC, DKRZ, RENCI, RWTH). In contrast to previous deduplication studies focusing mostly on backup data, we have analyzed over one PB (1212 TB) of online file system data. The evaluation shows that typically 20% to 30% of this online data can be removed by applying data deduplication techniques, peaking up to 70% for some data sets. This reduction can only be achieved by a subfile deduplication approach, while approaches based on whole-file comparisons only lead to small capacity savings. Dirk Meister, Jürgen Kaiser, André Brinkmann, Toni Cortes, Michael Kuhn 0003, Julian M. Kunkel |
SC | 5 |
| 2009 | Dynamic file system semantics to enable metadata optimizations in PVFSabstractAbstract Modern file systems maintain extensive metadata about stored files. While metadata typically is useful, there are situations when the additional overhead of such a design becomes a problem in terms of performance. This is especially true for parallel and cluster file systems, where every metadata operation is even more expensive due to their architecture. In this paper several changes made to the parallel cluster file system Parallel Virtual File System (PVFS) are presented. The changes target at the optimization of workloads with large numbers of small files. To improve the metadata performance, PVFS was modified such that unnecessary metadata is not managed anymore. Several tests with a large quantity of files were performed to measure the benefits of these changes. The tests have shown that common file system operations can be sped up by a factor of two even with relatively few changes. Copyright © 2009 John Wiley & Sons, Ltd. Michael Kuhn 0003, Julian M. Kunkel, Thomas Ludwig 0002 |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | Directory-Based Metadata Optimizations for Small Files in PVFS
Michael Kuhn 0003, Julian M. Kunkel, Thomas Ludwig 0002 |
Euro-Par | 1 |