Markus Büttner

dblp:98/5479 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
1since 2021 · last 2025
0009-0004-7205-1357ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%
Theoretical computer science
1 paper
Approximation and online algorithms · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.112005
Integrated prefetching and caching in single and parallel disk systems · Inf. Comput. 2005
Memory systems › cache management
prefetching and caching
0.112005
Integrated prefetching and caching in single and parallel disk systems · Inf. Comput. 2005
Approximation and online algorithms
online algorithms
0.012005
Integrated prefetching and caching in single and parallel disk systems · Inf. Comput. 2005
YearPublicationVenuePosition
2025 Analyzing performance portability for a SYCL implementation of the 2D shallow water equations
abstract
Abstract SYCL is an open standard for targeting heterogeneous hardware from C++. In this work, we evaluate a SYCL implementation for a discontinuous Galerkin discretization of the 2D shallow water equations targeting CPUs, GPUs, and also FPGAs. The discretization uses polynomial orders zero to two on unstructured triangular meshes. Separating memory accesses from the numerical code allow us to optimize data accesses for the target architecture. A performance analysis shows good portability across x86 and ARM CPUs, GPUs from different vendors, and even two variants of Intel Stratix 10 FPGAs. Measuring the energy to solution shows that GPUs yield an up to 10x higher energy efficiency in terms of degrees of freedom per joule compared to CPUs. With custom designed caches, FPGAs offer a meaningful complement to the other architectures with particularly good computational performance on smaller meshes. FPGAs with High Bandwidth Memory are less affected by bandwidth issues and have similar energy efficiency as latest generation CPUs.
Markus Büttner, Christoph Alt, Tobias Kenter, Harald Köstler, Christian Plessl, Vadym Aizinger
J. Supercomput.1
2019 Deterministic Fuzzy Checkpoints
abstract
Replicated systems tolerating arbitrary (Byzantine) faults require periodic and deterministic application-state checkpoints to perform essential tasks such as initializing new replicas, enabling faulty replicas to recover, and garbage-collecting old agreement-protocol messages. Existing techniques to create checkpoints in these systems make it necessary to temporarily suspend request execution in order to capture a consistent checkpoint, causing significant service disruptions for applications with large states. Unfortunately, state-of-the-art approaches from the domain of crash-tolerant systems also are not directly applicable, because the checkpoints they produce are not comparable across replicas and therefore cannot be validated in an environment in which replicas may fail arbitrarily and do not trust each other. In this paper, we address these problems by proposing deterministic fuzzy checkpoints (DFC), a novel technique that enables all correct replicas in a system to create consistent and matching checkpoints in parallel to processing requests. As a consequence, DFC increases service availability while still allowing replicas to verify the correctness of a checkpoint before applying it to their local states. In addition to our general approach, we present different alternatives to implement DFC within a replication library and furthermore discuss support for the creation of differential checkpoints. Experiments with a key-value store show that DFC is able to snapshot states of 3 GB while sustaining high performance throughout the entire checkpointing process.
Michael Eischer, Markus Büttner, Tobias Distler
SRDS2
2005 Enhanced prefetching and caching strategies for single- and multi-disk systems
Markus Büttner
Acta Informatica1
2005 Integrated prefetching and caching in single and parallel disk systems
Susanne Albers, Markus Büttner
Inf. Comput.2
2003 Integrated prefetching and caching in single and parallel disk systems
abstract
We study integrated prefetching and caching in single and parallel disk systems. There exist two very popular approximation algorithms called Aggressive and Conservative for minimizing the total elapsed time in the single disk problem. For D parallel disks, approximation algorithms are known for both the elapsed time and stall time performance measures. In particular, there exists a D-approximation algorithm for the stall time measure that uses D-1 additional memory locations in cache.In the first part of the paper we investigate approximation algorithms for the single disk problem. We give a refined analysis of the Aggressive algorithm, showing that the original analysis was too pessimistic. We prove that our new bound is tight. Additionally we present a new family of prefetching and caching strategies and give algorithms that perform better than Aggressive and Conservative.In the second part of the paper we investigate the problem of minimizing stall time in parallel disk systems. We present a polynomial time algorithm for computing a prefetching/caching schedule whose stall time is bounded by that of an optimal solution. The schedule uses at most 3(D-1) extra memory locations in cache. This is the first polynomial time algorithm for computing schedules with a minimum stall time. Our algorithm is based on the linear programming approach of [1]. However, in order to achieve minimum stall times, we introduce the new concept of synchronized schedules in which fetches on the D disks are performed completely in parallel.
Susanne Albers, Markus Büttner
SPAA2
2003 Integrated Prefetching and Caching with Read and Write Requests
Susanne Albers, Markus Büttner
WADS2