VLDB 2026 Research / reviewers in the wild / expert
Georg Hager
dblp:83/6202
· DBLP profile ↗
31ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-8723-2781ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring metrics for analyzing dynamic behavior in MPI programs via a coupled-oscillator modelabstractWe propose a novel, lightweight, and physically inspired approach to modeling the dynamics of parallel distributed-memory programs. Inspired by the Kuramoto model, we represent MPI processes as coupled oscillators with topology-aware interactions, custom coupling potentials, and stochastic noise. The resulting system of nonlinear ordinary differential equations opens a path to modeling key performance phenomena of parallel programs, including synchronization, delay propagation and decay, bottlenecks, and self-desynchronization. This paper introduces interaction potentials to describe memory- and compute-bound workloads and employs multiple quantitative metrics – such as an order parameter, synchronization entropy, phase gradients, and phase differences – to evaluate phase coherence and disruption. We also investigate the role of local noise and show that moderate noise can accelerate resynchronization in scalable applications. Our simulations align qualitatively with MPI trace data, showing the potential of physics-informed abstractions to predict performance patterns, which offers a new perspective for performance modeling and software-hardware co-design in parallel computing. Ayesha Afzal, Georg Hager, Gerhard Wellein |
Parallel Comput. | 2 |
| 2026 | Microarchitectural comparison, in-core modeling, and memory hierarchy analysis of state-of-the-art CPUs: Grace, Sapphire Rapids, and GenoaabstractThree big semiconductor companies in HPC are currently competing in the race for the best CPU: AMD, Intel, and NVIDIA. There are significant differences among their state-of-the-art CPU designs, spanning the entire range from instruction execution to cache behavior and main memory bandwidth. In this work, we analyze the performance of CPUs based on the Zen 4, Golden Cove, and Neoverse V2 microarchitectures. We create accurate in-core performance models for use with the Open Source Architecture Code Analyzer (OSACA) tool and compare its prediction accuracy with llvm-mca. Beyond the tool aspect, this reveals interesting differences in in-core design points but also some commonalities. Beyond the single core, we extend our comparison by measuring data-transfer behavior through the memory hierarchy using a variety of microbenchmarks. We thoroughly investigate the “write-allocate (WA) evasion” feature, which can automatically reduce the memory traffic caused by write misses. We show that the Grace Superchip has a next-to-optimal implementation of WA evasion while the Sapphire Rapids CPU can avoid write allocates completely only in specific scenarios. The only way to eliminate WAs on AMD Genoa is the explicit use of non-temporal stores. Finally, we study the cache hierarchy of the CPUs in view of the Execution-Cache-Memory (ECM) performance model, revealing overlapping cache hierarchies on Genoa and Grace in contrast to Sapphire Rapids. Jan Laukemann, Georg Hager, Gerhard Wellein |
Parallel Comput. | 2 |
| 2024 | CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate EvasionabstractIn this paper we analyze the MPI-only version of the CloverLeaf code from the SPEChpc 2021 benchmark suite on recent Intel Xeon "Ice Lake" and "Sapphire Rapids" server CPUs. We observe peculiar breakdowns in performance when the number of processes is prime. Investigating this effect, we create first-principles data traffic models for each of the stencil-like hotspot loops. With application measurements and microbenchmarks to study memory data traffic behavior, we can connect the breakdowns to SpecI2M, a new write-allocate evasion feature in current Intel CPUs. For serial and full-node cases we are able to predict the memory data volume analytically with an error of a few percent. We find that if the number of processes is prime, SpecI2M fails to work properly, which we can attribute to short inner loops emerging from the one-dimensional domain decomposition in this case. We can also rule out other possible causes of the prime number effect, such as breaking layer conditions, MPI communication overhead, and load imbalance. Jan Laukemann, Thomas Gruber 0007, Georg Hager, Dossay Oryspayev, Gerhard Wellein |
IPDPS | 3 |
| 2023 | Application Knowledge Required: Performance Modeling for Fun and ProfitabstractIn High Performance Computing, resource efficiency is paramount. Expensive systems need to be utilized to the maximum of their capabilities, but deep insight into the bottlenecks of a particular hardware-software combination is often lacking on the users' side. Analytic, first-principles performance models can provide such insight. They are built on simplified descriptions of the machine, the software, and how they interact. This goes, to some extent, against the general trend towards automation in computer science; the individual conducting the analysis does require some knowledge of the application and the hardware in order to make performance engineering a scientific process instead of blindly generating data with tools that are poorly understood. This talk uses examples from parallel high-performance computing to demonstrate how analytic performance models can support scientific thinking in performance engineering: Sparse matrix-vector multiplication, the HPCG benchmark, the CloverLeaf proxy app, and a lattice-Boltzmann solver. Interestingly, the most intriguing insights emerge from the failure of analytic models to accurately predict performance measurements. Georg Hager |
ICPE | 1 |
| 2023 | Making applications faster by asynchronous execution: Slowing down processes or relaxing MPI collectives
Ayesha Afzal, Georg Hager, Stefano Markidis, Gerhard Wellein |
Future Gener. Comput. Syst. | 2 |
| 2023 | MD-Bench: A performance-focused prototyping harness for state-of-the-art short-range molecular dynamics algorithms
Rafael Ravedutti L. Machado, Jan Eitzinger, Jan Laukemann, Georg Hager, Harald Köstler, Gerhard Wellein |
Future Gener. Comput. Syst. | 4 |
| 2023 | Analytical performance estimation during code generation on modern GPUs
Dominik Ernst, Markus Holzer 0005, Georg Hager, Matthias Knorr 0002, Gerhard Wellein |
J. Parallel Distributed Comput. | 3 |
| 2023 | The Role of Idle Waves, Desynchronization, and Bottleneck Evasion in the Performance of Parallel ProgramsabstractThe performance of highly parallel applications on distributed-memory systems is influenced by many factors. Analytic performance modeling techniques aim to provide insight into performance limitations and are often the starting point of optimization efforts. However, coupling analytic models across the system hierarchy (socket, node, network) fails to encompass the intricate interplay between the program code and the hardware, especially when execution and communication bottlenecks are involved. In this paper we investigate the effect ofbottleneck evasionand how it can lead to automatic overlap of communication overhead with computation. Bottleneck evasion leads to a gradual loss of the initial bulk-synchronous behavior of a parallel code so that its processes become desynchronized. This occurs most prominently in memory-bound programs, which is why we choose memory-bound benchmark and application codes, specifically an MPI-augmented STREAM Triad, sparse matrix-vector multiplication, and a collective-avoiding Chebyshev filter diagonalization code to demonstrate the consequences of desynchronization on two different supercomputing platforms. We investigate the role of idle waves as possible triggers for desynchronization and show the impact of automatic asynchronous communication for a spectrum of code properties and parameters, such as saturation point, matrix structures, domain decomposition, and communication concurrency. Our findings reveal how eliminating synchronization points (such as collective communication or barriers) precipitates performance improvements that go beyond what can be expected by simply subtracting the overhead of the collective from the overall runtime. Ayesha Afzal, Georg Hager, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Level-Based Blocking for Sparse Matrices: Sparse Matrix-Power-Vector MultiplicationabstractThe multiplication of a sparse matrix with a dense vector (SpMV) is a key component in many numerical schemes and its performance is known to be severely limited by main memory access. Several numerical schemes require the multiplication of a sparse matrix polynomial with a dense vector which is typically implemented as a sequence of SpMVs. This results in low performance and ignores the potential to increase the arithmetic intensity by reusing the matrix data from cache. In this work we use the recursive algebraic coloring engine (RACE) to enable blocking of sparse matrix data across the polynomial computations. In the graph representing the sparse matrix we form levels using a breadth-first search. Locality relations of these levels are then used to improve spatial and temporal locality when accessing the matrix data and to implement an efficient multithreaded parallelization. Our approach is independent of the matrix structure and avoids shortcomings of existing “blocking” strategies in terms of hardware efficiency and parallelization overhead. We quantify the quality of our implementation using performance modelling and demonstrate speedups of up to 3× and 5× compared to an optimal SpMV-based baseline on a single multicore chip of recent Intel and AMD architectures. Various numerical schemes like$s$-step Krylov solvers, polynomial preconditioners and power clustering algorithms will benefit from our development. Christie L. Alappat, Georg Hager, Olaf Schenk, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Addressing White-box Modeling and Simulation Challenges in Parallel ComputingabstractNo abstract available. Ayesha Afzal, Gerhard Wellein, Georg Hager |
SIGSIM-PADS | 3 |
| 2022 | Analytic performance model for parallel overlapping memory-bound kernelsabstractAbstract Complex applications running on multicore processors show a rich performance phenomenology. The growing number of cores per ccNUMA domain complicates performance analysis of memory‐bound code since system noise, load imbalance, or task‐based programming models can lead to thread desynchronization. Hence, the simplifying assumption that all cores execute the same loop can not be upheld. Motivated by observations on plain and modified versions of the HPCG benchmark, we construct a performance model of execution of memory‐bound loop kernels. It can predict the memory bandwidth share per kernel on a memory contention domain depending on the number of active cores and which other workload the kernel is paired with. The only code features required are the single‐thread memory request fraction per kernel, which is directly related to the single‐thread memory bandwidth, and its saturated bandwidth. The former can either be measured directly or predicted using the Execution‐Cache‐Memory performance model. The computational intensity of the kernels and the detailed structure of the code is of no significance. We validate our model on Intel Broadwell, Intel Cascade Lake, and AMD Rome processors pairing various streaming and stencil kernels. The error in predicting the bandwidth share per kernel is less than 8%. Ayesha Afzal, Georg Hager, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Execution-Cache-Memory modeling and performance tuning of sparse matrix-vector multiplication and Lattice quantum chromodynamics on A64FXabstractAbstract The A64FX CPU is arguably the most powerful Arm‐based processor design to date. Although it is a traditional cache‐based multicore processor, its peak performance and memory bandwidth rival accelerator devices. A good understanding of its performance features is of paramount importance for developers who wish to leverage its full potential. We present an architectural analysis of the A64FX used in the Fujitsu FX1000 supercomputer at a level of detail that allows for the construction of Execution‐Cache‐Memory performance models for steady‐state loops. In the process we identify architectural peculiarities that point to viable generic optimization strategies. After validating the model using simple streaming loops we apply the insight gained to sparse matrix‐vector multiplication (SpMV) and the domain wall (DW) kernel from quantum chromodynamics. For SpMV we show why the compressed row storage (CRS) matrix storage format is not a good practical choice on this architecture and how the SELL‐C‐ format can achieve bandwidth saturation. For the DW kernel we provide a cache‐reuse analysis and show how an appropriate choice of data layout for complex arrays can realize memory‐bandwidth saturation in this case as well. A comparison with state‐of‐the‐art high‐end Intel Cascade Lake AP and Nvidia V100 systems puts the capabilities of the A64FX into perspective. We also explore the potential for power optimizations using the tuning knobs provided by the Fugaku system, achieving energy savings of about 31% for SpMV and 18% for DW. Christie L. Alappat, Nils Meyer, Jan Laukemann, Thomas Gruber 0007, Georg Hager, Gerhard Wellein, Tilo Wettig |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | YaskSite: Stencil Optimization Techniques Applied to Explicit ODE Methods on Modern ArchitecturesabstractThe landscape of multi-core architectures is growing more complex and diverse. Optimal application performance tuning parameters can vary widely across CPUs, and finding them in a possibly multidimensional parameter search space can be time consuming, expensive and potentially infeasible. In this work, we introduce YaskSite, a tool capable of tackling these challenges for stencil computations. YaskSite is built upon Intel's YASK framework. It combines YASK's flexibility to deal with different target architectures with the Execution-Cache-Memory performance model, which enables identifying optimal performance parameters analytically without the need to run the code. Further we show that YaskSite's features can be exploited by external tuning frameworks to reliably select the most efficient kernel(s) for the application at hand. To demonstrate this, we integrate YaskSite into Offsite, an offline tuner for explicit ordinary differential equation methods, and show that the generated performance predictions are reliable and accurate, leading to considerable performance gains at minimal code generation time and autotuning costs on the latest Intel Cascade Lake and AMD Rome CPUs. Christie L. Alappat, Johannes Seiferth, Georg Hager, Matthias Korch, Thomas Rauber, Gerhard Wellein |
CGO | 3 |
| 2021 | Opening the Black Box: Performance Estimation during Code Generation for GPUsabstractAutomatic code generation is frequently used to create implementations of algorithms specifically tuned to particular hardware and application parameters. The code generation process involves the selection of adequate code transformations, tuning parameters, and parallelization strategies. To cover the huge search space, code generation frameworks may apply time-intensive autotuning, exploit scenario-specific performance models, or treat performance as an intangible black box that must be described via machine learning. This paper addresses the selection problem by identifying the relevant performance-defining mechanisms through a performance model coupled with an analytic hardware metric estimator. This enables a quick exploration of large configuration spaces to identify highly efficient candidates with high accuracy. Our current approach targets memory-intensive GPGPU applications and focuses on the correct modeling of data transfer volumes to all levels of the memory hierarchy. We show how our method can be coupled to the “pystencils” stencil code generator, which is used to generate kernels for a range four 3D25pt stencil and a complex two phase fluid solver based on the Lattice Boltzmann Method. For both, it delivers a ranking that can be used to select the best performing candidate. The method is not limited to stencil kernels, but can be integrated into any code generator that can generate the required address expressions. Dominik Ernst, Georg Hager, Matthias Knorr 0002, Gerhard Wellein, Markus Holzer 0005 |
SBAC-PAD | 2 |
| 2020 | PHIST: A Pipelined, Hybrid-Parallel Iterative Solver ToolkitabstractThe increasing complexity of hardware and software environments in high-performance computing poses big challenges on the development of sustainable and hardware-efficient numerical software. This article addresses these challenges in the context of sparse solvers. Existing solutions typically target sustainability, flexibility, or performance, but rarely all of them. Our new library PHIST provides implementations of solvers for sparse linear systems and eigenvalue problems. It is a productivity platform for performance-aware developers of algorithms and application software with abstractions that do not obscure the view on hardware-software interaction. The PHIST software architecture and the PHIST development process were designed to overcome shortcomings of existing packages. An interface layer for basic sparse linear algebra functionality that can be provided by multiple backends ensures sustainability, and PHIST supports common techniques for improving scalability and performance of algorithms such as blocking and kernel fusion. We showcase these concepts using the PHIST implementation of a block Jacobi-Davidson solver for non-Hermitian and generalized eigenproblems. We study its performance on a multi-core CPU, a GPU, and a large-scale many-core system. Furthermore, we show how an existing implementation of a block Krylov-Schur method in the Trilinos package Anasazi can benefit from the performance engineering techniques used in PHIST. Jonas Thies, Melven Röhrig-Zöllner, Nigel Overmars, Achim Basermann, Dominik Ernst, Georg Hager, Gerhard Wellein |
ACM Trans. Math. Softw. | 6 |
| 2019 | Propagation and Decay of Injected One-Off Delays on Clusters: A Case StudyabstractAnalytic, first-principles performance modeling of distributed-memory applications is difficult due to a wide spectrum of random disturbances caused by the application and the system. These disturbances (commonly called “noise”) run contrary to the assumptions about regularity that one usually employs when constructing simple analytic models. Despite numerous efforts to quantify, categorize, and reduce such effects, a comprehensive quantitative understanding of their performance impact is not available, especially for long, one-off delays of execution periods that have global consequences for the parallel application. In this work, we investigate various traces collected from synthetic benchmarks that mimic real applications on simulated and real message-passing systems in order to pin-point the mechanisms behind delay propagation. We analyze the dependence of the propagation speed of “idle waves,” i.e., propagating phases of inactivity, emanating from injected delays with respect to the execution and communication properties of the application, study how such delays decay under increased noise levels, and how they interact with each other. We also show how fine-grained noise can make a system immune against the adverse effects of propagating idle waves. Our results contribute to a better understanding of the collective phenomena that manifest themselves in distributed-memory parallel applications. Ayesha Afzal, Georg Hager, Gerhard Wellein |
CLUSTER | 2 |
| 2019 | CRAFT: A Library for Easier Application-Level Checkpoint/Restart and Automatic Fault ToleranceabstractIn order to efficiently use the future generations of supercomputers, fault tolerance and power consumption are two of the prime challenges anticipated by the High Performance Computing (HPC) community. Checkpoint/Restart (CR) has been and still is the most widely used technique to deal with hard failures. Application-level CR is the most effective CR technique in terms of overhead efficiency but it takes a lot of implementation effort. This work presents the implementation of our C++ based library CRAFT (Checkpoint-Restart and Automatic Fault Tolerance), which serves two purposes. First, it provides an extendable library that significantly eases the implementation of application-level checkpointing. The most basic and frequently used checkpoint data-types are already part of CRAFT and can be directly used out of the box. The library can be easily extended to add more data-types. As means of overhead reduction, the library offers a built-in asynchronous checkpointing mechanism and also supports the Scalable Checkpoint/Restart (SCR) library for node level checkpointing. Second, CRAFT provides an easier interface for User-Level Failure Mitigation (ULFM) based dynamic process recovery, which significantly reduces the complexity and effort of failure detection and communication recovery mechanism. By utilizing both functionalities together, applications can write application-level checkpoints and recover dynamically from process failures with very limited programming effort. This work presents the challenges addressed by the library, its design, and its use. The associated overheads are analyzed using benchmarks. Faisal Shahzad 0001, Jonas Thies, Moritz Kreutzer, Thomas Zeiser, Georg Hager, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Multicore Performance Engineering of Sparse Triangular Solves Using a Modified Roofline ModelabstractThe Roofline model is widely used to visualize the performance of executed code together with the upper performance bounds given by the memory bandwidth and the processor peak performance. The model can thus provide an insightful visualization of bottlenecks. In this paper, we try to establish realistic bandwidth ceilings for the sparse triangular solve step of PARDISO, a leading sparse direct solver package, which is also part of the Intel MKL library. The performance of the forward and backward substitution process is analyzed and benchmarked for a representative set of sparse matrices on seven modern x86-type multicore architectures and the Knights Landing manycore architecture. It is shown how to accurately measure the necessary quantities also for threaded code, and the measurement approach, its validation, as well as limitations are discussed. Our modeling approach covers the serial and parallel execution phases, allowing for in-socket performance predictions. Markus Wittmann, Georg Hager, Radim Janalík, Martin Lanser, Axel Klawonn, Oliver Rheinbach, Olaf Schenk, Gerhard Wellein |
SBAC-PAD | 2 |
| 2017 | LIKWID Monitoring Stack: A Flexible Framework Enabling Job Specific Performance monitoring for the massesabstractSystem monitoring is an established tool to measure the utilization and health of HPC systems. Usually system monitoring infrastructures make no connection to job information and do not utilize hardware performance monitoring (HPM) data. To increase the efficient use of HPC systems automatic and continuous performance monitoring of jobs is an essential component. It can help to identify pathological cases, provides instant performance feedback to the users, offers initial data to judge on the optimization potential of applications and helps to build a statistical foundation about application specific system usage. The LIKWID monitoring stack is a modular framework build on top of the LIKWID tools library. It aims on enabling job specific performance monitoring using HPM data, system metrics and applicationlevel data for small to medium sized commodity clusters. Moreover, it is designed to integrate in existing monitoring infrastructures to speed up the change from pure system monitoring to job-aware monitoring. Thomas Röhl, Jan Eitzinger, Georg Hager, Gerhard Wellein |
CLUSTER | 3 |
| 2017 | Performance analysis of the Kahan-enhanced scalar product on current multi-core and many-core processorsabstractSummary We investigate the performance characteristics of a numerically enhanced scalar product (dot) kernel loop that uses the Kahan algorithm to compensate for numerical errors, and describe efficient single instruction multiple data‐vectorized implementations on recent multi‐core and many‐core processors. Using low‐level instruction analysis and the execution‐cache‐memory performance model, we pinpoint the relevant performance bottlenecks for single‐core and thread‐parallel execution and predict performance and saturation behavior. We show that the Kahan‐enhanced scalar product comes at almost no additional cost compared with the naive (non‐Kahan) scalar product if appropriate low‐level optimizations, notably single instruction multiple data vectorization and unrolling, are applied. The execution‐cache‐memory model is extended appropriately to accommodate not only modern Intel multicore chips but also the Intel Xeon Phi ‘Knights Corner’ coprocessor and an IBM POWER8 CPU. This allows us to discuss the impact of processor features on the performance across four modern architectures that are relevant for high performance computing. Copyright © 2016 John Wiley & Sons, Ltd. Johannes Hofmann 0001, Dietmar Fey, Michael Riedmann, Jan Eitzinger, Georg Hager, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | Optimization of an Electromagnetics Code with Multicore Wavefront Diamond Blocking and Multi-dimensional Intra-Tile ParallelizationabstractUnderstanding and optimizing the properties of solar cells is becoming a key issue in the search for alternatives to nuclear and fossil energy sources. A theoretical analysis via numerical simulations involves solving Maxwell's Equations in discretized form and typically requires substantial computing effort. We start from a hybrid-parallel (MPI+OpenMP) production code that implements the Time Harmonic Inverse Iteration Method (THIIM) with Finite-Difference Frequency Domain (FDFD) discretization. Although this algorithm has the characteristics of a strongly bandwidth-bound stencil update scheme, it is significantly different from the popular stencil types that have been exhaustively studied in the high performance computing literature to date. We apply a recently developed stencil optimization technique, multicore wavefront diamond tiling with multi-dimensional cache block sharing, and describe in detail the peculiarities that need to be considered due to the special stencil structure. Concurrency in updating the components of the electric and magnetic fields provides an additional level of parallelism. The dependence of the cache size requirement of the optimized code on the blocking parameters is modeled accurately, and an auto-tuner searches for optimal configurations in the remaining parameter space. We were able to completely decouple the execution from the memory bandwidth bottleneck, accelerating the implementation by a factor of three to four compared to an optimal implementation with pure spatial blocking on an 18-core Intel Haswell CPU. Tareq M. Malas, Julian Hornich, Georg Hager, Hatem Ltaief, Christoph Pflaum, David E. Keyes |
IPDPS | 3 |
| 2016 | Performance and power for highly parallel systemsabstractThis special issue is the result of an open call for papers initiated after the minisymposium ‘Analysis and Modeling: Techniques and Tools’, conducted at the Society for Industrial and Applied Mathematics (SIAM) Conference on Parallel Processing in Scientific Computing in Savannah, GA, in February 2012. The minisymposium brought together tool developers, performance and power modeling experts, and application analysts to present the state of the art on performance analysis and modeling techniques. This unique combination of expertise is highly needed in a time where complex, hierarchical architectures are the standard for all highly parallel computer systems. Without a good grip on relevant performance limitations, any ptimization attempt is just a shot in the dark. Hence, it is crucial to fully understand the performance properties and bottlenecks that come about with clustered multicore/many-core, multisocket nodes. Another aspect of modern systems is the complicated interplay between power constraints and the need for compute performance, which leads to complicated trade-offs. The challenges ahead are many-fold as systems scale in size. While parallelism is increasing, memory systems, interconnection networks, storage, and uncertainties in programming models all add to the complexities. More rapid realization of energy savings will require significant increases in measurement resolution and optimization techniques. This special issue is focused on how performance and power properties of modern highly parallel systems can be analyzed using state-of-the-art modeling and analysis techniques and real-world applications and tools. G. Hager, J. Treibig, J. Habich, and G. Wellein 1 introduce simple but insightful analytic models for execution performance and energy consumption of multicore CPUs. Automatic dynamic voltage and frequency scaling is leveraged by the “Green Queue” framework presented by J. Peraza, A. Tiwari, M. Laurenzano, L. Carrington, and A. E. Snavely 2 in their article. They show that significant energy savings at low performance loss are in reach if dynamic voltage and frequency scaling is used in an application-aware manner. A. D. Breslow, L. Porter, A. Tiwari, M. Laurenzano, L. Carrington, D. M. Tullsen, and A. E. Snavely 3 investigate the potential of job striping, a technique for co-locating HPC workloads with different characteristics on the same CPU chip, and demonstrate increased throughput and energy efficiency for a mix of typical simulation codes on a production cluster. The problem of how to deal with coarse-grained power measurements is tackled in the paper by H. Servat, G. Llort, J. Giménez, and J. Labarta 4. They present a tool that can derive fine-grained power and performance data for code with quickly alternating phases. The power usage and power variability of workloads on production supercomputers at Los Alamos National Laboratory are studied by S. Pakin, C. Storlie, M. Lang, R. E. Fields, E. E. Romero, C. Idler, S. Michalak, H. Greenberg, J. Loncaric, R. Rheinheimer, G. Grider, and J. Wendelberger in their paper 5. One of their central findings is that real power dissipation under real-world workloads is significantly lower than what the power infrastructure can handle, which opens interesting possibilities for saving cost via power capping. We think that this selection of papers is unique in providing several very different views on the problem of performance and power efficiency on present-day parallel machines from the core to the computing center level. Georg Hager, Darren J. Kerbyson, Abhinav Vishnu, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Exploring performance and power properties of modern multi-core chips via simple machine modelsabstractSummary Modern multi‐core chips show complex behavior with respect to performance and power. Starting with the Intel Sandy Bridge processor, it has become possible to directly measure the power dissipation of a CPU chip and correlate this data with the performance properties of the running code. Going beyond a simple bottleneck analysis, we employ the recently published Execution‐Cache‐Memory (ECM) model to describe the single‐core and multi‐core performance of streaming kernels. The model refines the well‐known roofline model, because it can predict the scaling and the saturation behavior of bandwidth‐limited loop kernels on a multi‐core chip. The saturation point is especially relevant for considerations of energy consumption. From power dissipation measurements of benchmark programs with vastly different requirements to the hardware, we derive a simple, phenomenological power model for the Sandy Bridge processor. Together with the ECM model, we are able to explain many peculiarities in the performance and power behavior of multi‐core processors and derive guidelines for energy‐efficient execution of parallel programs. Finally, we show that the ECM and power models can be successfully used to describe the scaling and power behavior of a lattice Boltzmann flow solver code. Copyright © 2013 John Wiley & Sons, Ltd. Georg Hager, Jan Eitzinger, Johannes Habich, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Chip-level and multi-node analysis of energy-optimized lattice Boltzmann CFD simulationsabstractSummary Memory‐bound algorithms show complex performance and energy consumption behavior on multicore processors. We choose the lattice Boltzmann method on an Intel Sandy Bridge cluster as a prototype scenario to investigate if and how single‐chip performance and power characteristics can be generalized to the highly parallel case. First, we perform an analysis of a sparse‐lattice lattice Boltzmann method implementation for complex geometries. Using a single‐core performance model, we predict the intra‐chip saturation characteristics and the optimal operating point in terms of energy‐to‐solution as a function of implementation details, clock frequency, vectorization, and number of active cores per chip. We show that high single‐core performance and a correct choice of the number of active cores per chip are the essential optimizations for the lowest energy‐to‐solution at minimal performance degradation. Then we extrapolate to the Message Passing Interface (MPI)‐parallel level and quantify the energy‐saving potential of various optimizations and execution modes, where we find these guidelines to be even more important, especially when communication overhead is non‐negligible. In our setup, we could achieve energy savings of 35% in this case, compared with a naive approach. We also demonstrate that a simple non‐reflective reduction of the clock speed leaves most of the energy‐saving potential unused. Copyright © 2015 John Wiley & Sons, Ltd. Markus Wittmann, Georg Hager, Thomas Zeiser, Jan Eitzinger, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Building a Fault Tolerant Application Using the GASPI Communication LayerabstractIt is commonly agreed that highly parallel software on Exascale computers will suffer from many more runtime failures due to the decreasing trend in the mean time to failures (MTTF). Therefore, it is not surprising that a lot of research is going on in the area of fault tolerance and fault mitigation. Applications should survive a failure and/or be able to recover with minimal cost. MPI is not yet very mature in handling failures, the User-Level Failure Mitigation (ULFM) proposal being currently the most promising approach is still in its prototype phase. In our work we use GASPI, which is a relatively new communication library based on the PGAS model. It provides the missing features to allow the design of fault-tolerant applications. Instead of introducing algorithm-based fault tolerance in its true sense, we demonstrate how we can build on (existing) clever checkpointing and extend applications to allow integrate a low cost fault detection mechanism and, if necessary, recover the application on the fly. The aspects of process management, the restoration of groups and the recovery mechanism is presented in detail. We use a sparse matrix vector multiplication based application to perform the analysis of the overhead introduced by such modifications. Our fault detection mechanism causes no overhead in failure-free cases, whereas in case of failure(s), the failure detection and recovery cost is of reasonably acceptable order and shows good scalability. Faisal Shahzad 0001, Moritz Kreutzer, Thomas Zeiser, Andreas Pieper, Georg Hager, Gerhard Wellein |
CLUSTER | 6 |
| 2015 | Quantifying Performance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory ModelabstractStencil algorithms on regular lattices appear in many fields of computational science, and much effort has been put into optimized implementations. Such activities are usually not guided by performance models that provide estimates of expected speedup. Understanding the performance properties and bottlenecks by performance modeling enables a clear view on promising optimization opportunities. In this work we refine the recently developed Execution-Cache-Memory (ECM) model and use it to quantify the performance bottlenecks of stencil algorithms on a contemporary Intel processor. This includes applying the model to arrive at single-core performance and scalability predictions for typical "corner case" stencil loop kernels. Guided by the ECM model we accurately quantify the significance of "layer conditions," which are required to estimate the data traffic through the memory hierarchy, and study the impact of typical optimization approaches such as spatial blocking, strength reduction, and temporal blocking for their expected benefits. We also compare the ECM model to the widely known Roofline model. Holger Stengel, Jan Eitzinger, Georg Hager, Gerhard Wellein |
ICS | 3 |
| 2015 | Performance Engineering of the Kernel Polynomal Method on Large-Scale CPU-GPU SystemsabstractThe Kernel Polynomial Method (KPM) is a well-established scheme in quantum physics and quantum chemistry to determine the Eigen value density and spectral properties of large sparse matrices. In this work we demonstrate the high optimization potential and feasibility of peta-scale heterogeneous CPU-GPU implementations of the KPM. At the node level we show that it is possible to decouple the sparse matrix problem posed by KPM from main memory bandwidth both on CPU and GPU. To alleviate the effects of scattered data access we combine loosely coupled outer iterations with tightly coupled block sparse matrix multiple vector operations, which enables pure data streaming. All optimizations are guided by a performance analysis and modelling process that indicates how the computational bottlenecks change with each optimization step. Finally we use the optimized node-level KPM with a hybrid-parallel framework to perform large-scale heterogeneous electronic structure calculations for novel topological materials on a pet scale-class Cray XC30 system. Moritz Kreutzer, Andreas Pieper, Georg Hager, Gerhard Wellein, Andreas Alvermann, Holger Fehske |
IPDPS | 3 |
| 2011 | A flexible Patch-based lattice Boltzmann parallelization approach for heterogeneous GPU-CPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Georg Hager, Ulrich Rüde, Gerhard Wellein |
Parallel Comput. | 4 |
| 2009 | The world's fastest CPU and SMP node: Some performance results from the NEC SX-9abstractClassic vector systems have all but vanished from recent TOP500 lists. Looking at the newly introduced NEC SX-9 series, we benchmark its memory subsystem using the low level vector triad and employ an advanced lattice Boltzmann flow solver kernel to demonstrate that classic vectors still combine excellent performance with a well-established optimization approach. Results for commodity x86-based systems are provided for reference. Thomas Zeiser, Georg Hager, Gerhard Wellein |
IPDPS | 2 |
| 2009 | Hybrid MPI/OpenMP Parallel Programming on Clusters of Multi-Core SMP NodesabstractToday most systems in high-performance computing (HPC) feature a hierarchical hardware design: Shared memory nodes with several multi-core CPUs are connected via a network infrastructure. Parallel programming must combine distributed memory parallelization on the node interconnect with shared memory parallelization inside each node. We describe potentials and challenges of the dominant programming models on hierarchically structured hardware: Pure MPI (Message Passing Interface), pure OpenMP (with distributed shared memory extensions) and hybrid MPI+OpenMP in several flavors. We pinpoint cases where a hybrid programming model can indeed be the superior solution because of reduced communication needs and memory consumption, or improved load balance. Furthermore we show that machine topology has a significant impact on performance for all parallelization strategies and that topology awareness should be built into all applications in the future. Finally we give an outlook on possible standardization goals and extensions that could make hybrid programming easier to do with performance in mind. Rolf Rabenseifner, Georg Hager, Gabriele Jost |
PDP | 2 |
| 2008 | Data access optimizations for highly threaded multi-core CPUs with multiple memory controllersabstractProcessor and system architectures that feature multiple memory controllers are prone to show bottlenecks and erratic performance numbers on codes with regular access patterns. Although such effects are well known in the form of cache thrashing and aliasing conflicts, they become more severe when memory access is involved. Using the new Sun UltraSPARC T2 processor as a prototypical multi-core design, we analyze performance patterns in low-level and application benchmarks and show ways to circumvent bottlenecks by careful data layout and padding. Georg Hager, Thomas Zeiser, Gerhard Wellein |
IPDPS | 1 |