Thomas Zeiser

dblp:32/2936 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
0since 2021 · last 2020
0009-0002-2916-911XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 79% High-performance computing · 11% Performance modeling and evaluation · 10%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › fault tolerance
checkpointing
0.412019
CRAFT: A Library for Easier Application-Level Checkpoint/Restart and Automatic Fault Tolerance · IEEE Trans. Parallel Distributed Syst. 2019
Distributed systems
fault tolerance
0.412019
CRAFT: A Library for Easier Application-Level Checkpoint/Restart and Automatic Fault Tolerance · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
benchmarking
0.012004
Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures · SC 2004
High-performance computing › scientific computing systems › computational fluid dynamics
lattice boltzmann method
0.012004
Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures · SC 2004
Performance modeling and evaluation
parallel performance evaluation
0.012004
Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures · SC 2004
High-performance computing
performance optimization at scale
0.012004
Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures · SC 2004
High-performance computing
supercomputer architecture
0.012004
Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures · SC 2004

Methods — techniques the papers use, named apart from their topics

asynchronous checkpointing · 0.4ULFM · 0.4SCR · 0.4
YearPublicationVenuePosition
2020 What does Power Consumption Behavior of HPC Jobs Reveal? : Demystifying, Quantifying, and Predicting Power Consumption Characteristics
abstract
As we approach exascale computing, large-scale HPC systems are becoming increasingly power-constrained, requiring them to run HPC workloads in an energy-efficient manner. The first step toward achieving this goal is to better understand, analyze, and quantify the power consumption characteristics of HPC jobs. However, there is a lack of understanding of the power consumption characteristics of HPC jobs which run on production HPC systems. Such characterization is required to guide the design of the next generation of power-aware resource management. To the best of our knowledge, we are the first study to open-source the data and analysis of power-consumption characteristics of HPC jobs and users from two medium-scale production HPC clusters.
Tirthak Patel, Adam Wagenhäuser, Christopher Eibel, Timo Hönig, Thomas Zeiser, Devesh Tiwari
IPDPS5
2019 ClusterCockpit - A web application for job-specific performance monitoring
abstract
Monitoring is a common component of HPC system software. Up to now, monitoring focused mainly on health checking and system level performance as well as on job scheduler information and was targeted towards system administrators. Recently job-specific performance monitoring based on hardware performance counter metrics has gained attention at academic HPC computing centers. HPC is becoming a mainstream tool that is also used by non-HPC experts, and HPC centers see a demand to check for pathological jobs and jobs with large optimization potential. The possibility to measure hardware performance counter data with negligible overhead allows assessment of efficient resource utilization and detection of pathological jobs. Pathological jobs are, e.g. jobs with errors in the batch script, jobs which do not terminate, jobs with severe load imbalance, or jobs that do not use any resources. This paper introduces ClusterCockpit, a web front-end tailor-made tool for job-specific performance monitoring. While many recent job-specific performance monitoring efforts concentrate on the measurement and data collection layers, ClusterCockpit provides a modern user interface targeted towards performance analysts as well as application users.
Jan Eitzinger, Thomas Gruber 0007, Ayesha Afzal, Thomas Zeiser, Gerhard Wellein
CLUSTER4
2019 CRAFT: A Library for Easier Application-Level Checkpoint/Restart and Automatic Fault Tolerance
abstract
In order to efficiently use the future generations of supercomputers, fault tolerance and power consumption are two of the prime challenges anticipated by the High Performance Computing (HPC) community. Checkpoint/Restart (CR) has been and still is the most widely used technique to deal with hard failures. Application-level CR is the most effective CR technique in terms of overhead efficiency but it takes a lot of implementation effort. This work presents the implementation of our C++ based library CRAFT (Checkpoint-Restart and Automatic Fault Tolerance), which serves two purposes. First, it provides an extendable library that significantly eases the implementation of application-level checkpointing. The most basic and frequently used checkpoint data-types are already part of CRAFT and can be directly used out of the box. The library can be easily extended to add more data-types. As means of overhead reduction, the library offers a built-in asynchronous checkpointing mechanism and also supports the Scalable Checkpoint/Restart (SCR) library for node level checkpointing. Second, CRAFT provides an easier interface for User-Level Failure Mitigation (ULFM) based dynamic process recovery, which significantly reduces the complexity and effort of failure detection and communication recovery mechanism. By utilizing both functionalities together, applications can write application-level checkpoints and recover dynamically from process failures with very limited programming effort. This work presents the challenges addressed by the library, its design, and its use. The associated overheads are analyzed using benchmarks.
Faisal Shahzad 0001, Jonas Thies, Moritz Kreutzer, Thomas Zeiser, Georg Hager, Gerhard Wellein
IEEE Trans. Parallel Distributed Syst.4
2016 Chip-level and multi-node analysis of energy-optimized lattice Boltzmann CFD simulations
abstract
Summary Memory‐bound algorithms show complex performance and energy consumption behavior on multicore processors. We choose the lattice Boltzmann method on an Intel Sandy Bridge cluster as a prototype scenario to investigate if and how single‐chip performance and power characteristics can be generalized to the highly parallel case. First, we perform an analysis of a sparse‐lattice lattice Boltzmann method implementation for complex geometries. Using a single‐core performance model, we predict the intra‐chip saturation characteristics and the optimal operating point in terms of energy‐to‐solution as a function of implementation details, clock frequency, vectorization, and number of active cores per chip. We show that high single‐core performance and a correct choice of the number of active cores per chip are the essential optimizations for the lowest energy‐to‐solution at minimal performance degradation. Then we extrapolate to the Message Passing Interface (MPI)‐parallel level and quantify the energy‐saving potential of various optimizations and execution modes, where we find these guidelines to be even more important, especially when communication overhead is non‐negligible. In our setup, we could achieve energy savings of 35% in this case, compared with a naive approach. We also demonstrate that a simple non‐reflective reduction of the clock speed leaves most of the energy‐saving potential unused. Copyright © 2015 John Wiley & Sons, Ltd.
Markus Wittmann, Georg Hager, Thomas Zeiser, Jan Eitzinger, Gerhard Wellein
Concurr. Comput. Pract. Exp.3
2015 Building a Fault Tolerant Application Using the GASPI Communication Layer
abstract
It is commonly agreed that highly parallel software on Exascale computers will suffer from many more runtime failures due to the decreasing trend in the mean time to failures (MTTF). Therefore, it is not surprising that a lot of research is going on in the area of fault tolerance and fault mitigation. Applications should survive a failure and/or be able to recover with minimal cost. MPI is not yet very mature in handling failures, the User-Level Failure Mitigation (ULFM) proposal being currently the most promising approach is still in its prototype phase. In our work we use GASPI, which is a relatively new communication library based on the PGAS model. It provides the missing features to allow the design of fault-tolerant applications. Instead of introducing algorithm-based fault tolerance in its true sense, we demonstrate how we can build on (existing) clever checkpointing and extend applications to allow integrate a low cost fault detection mechanism and, if necessary, recover the application on the fly. The aspects of process management, the restoration of groups and the recovery mechanism is presented in detail. We use a sparse matrix vector multiplication based application to perform the analysis of the overhead introduced by such modifications. Our fault detection mechanism causes no overhead in failure-free cases, whereas in case of failure(s), the failure detection and recovery cost is of reasonably acceptable order and shows good scalability.
Faisal Shahzad 0001, Moritz Kreutzer, Thomas Zeiser, Andreas Pieper, Georg Hager, Gerhard Wellein
CLUSTER3
2012 Asynchronous Checkpointing by Dedicated Checkpoint Threads
Faisal Shahzad 0001, Markus Wittmann, Thomas Zeiser, Gerhard Wellein
EuroMPI3
2009 The world's fastest CPU and SMP node: Some performance results from the NEC SX-9
abstract
Classic vector systems have all but vanished from recent TOP500 lists. Looking at the newly introduced NEC SX-9 series, we benchmark its memory subsystem using the low level vector triad and employ an advanced lattice Boltzmann flow solver kernel to demonstrate that classic vectors still combine excellent performance with a well-established optimization approach. Results for commodity x86-based systems are provided for reference.
Thomas Zeiser, Georg Hager, Gerhard Wellein
IPDPS1
2008 Data access optimizations for highly threaded multi-core CPUs with multiple memory controllers
abstract
Processor and system architectures that feature multiple memory controllers are prone to show bottlenecks and erratic performance numbers on codes with regular access patterns. Although such effects are well known in the form of cache thrashing and aliasing conflicts, they become more severe when memory access is involved. Using the new Sun UltraSPARC T2 processor as a prototypical multi-core design, we analyze performance patterns in low-level and application benchmarks and show ways to circumvent bottlenecks by careful data layout and padding.
Georg Hager, Thomas Zeiser, Gerhard Wellein
IPDPS2
2004 Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures
abstract
Computationally intensive programs with moderate communication requirements such as CFD codes suffer from the standard slow interconnects of commodity "off the shelf" (COTS) hardware. We will introduce different large-scale applications of the Lattice Boltzmann Method (LBM) in fluid dynamics, material science, and chemical engineering and present results of the parallel performance on different architectures. It will be shown that a high speed communication network in combination with an efficient CPU is mandatory in order to achieve the required performance. An estimation of the necessary CPU count to meet the performance of 1 TFlop/s will be given as well as a prediction as to which architecture is the most suitable for LBM. Finally, ratios of costs to application performance for tailored HPC systems and COTS architectures will be presented.
Thomas Pohl, Frank Deserno, Nils Thürey, Ulrich Rüde, Peter Lammers, Gerhard Wellein, Thomas Zeiser
SC7