EDBT 2026 Demo / reviewers in the wild / expert
Gerhard Wellein
dblp:00/3710
· DBLP profile ↗
35ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0001-7371-3026ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 12 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring metrics for analyzing dynamic behavior in MPI programs via a coupled-oscillator modelabstractWe propose a novel, lightweight, and physically inspired approach to modeling the dynamics of parallel distributed-memory programs. Inspired by the Kuramoto model, we represent MPI processes as coupled oscillators with topology-aware interactions, custom coupling potentials, and stochastic noise. The resulting system of nonlinear ordinary differential equations opens a path to modeling key performance phenomena of parallel programs, including synchronization, delay propagation and decay, bottlenecks, and self-desynchronization. This paper introduces interaction potentials to describe memory- and compute-bound workloads and employs multiple quantitative metrics – such as an order parameter, synchronization entropy, phase gradients, and phase differences – to evaluate phase coherence and disruption. We also investigate the role of local noise and show that moderate noise can accelerate resynchronization in scalable applications. Our simulations align qualitatively with MPI trace data, showing the potential of physics-informed abstractions to predict performance patterns, which offers a new perspective for performance modeling and software-hardware co-design in parallel computing. Ayesha Afzal, Georg Hager, Gerhard Wellein |
Parallel Comput. | 3 |
| 2026 | Microarchitectural comparison, in-core modeling, and memory hierarchy analysis of state-of-the-art CPUs: Grace, Sapphire Rapids, and GenoaabstractThree big semiconductor companies in HPC are currently competing in the race for the best CPU: AMD, Intel, and NVIDIA. There are significant differences among their state-of-the-art CPU designs, spanning the entire range from instruction execution to cache behavior and main memory bandwidth. In this work, we analyze the performance of CPUs based on the Zen 4, Golden Cove, and Neoverse V2 microarchitectures. We create accurate in-core performance models for use with the Open Source Architecture Code Analyzer (OSACA) tool and compare its prediction accuracy with llvm-mca. Beyond the tool aspect, this reveals interesting differences in in-core design points but also some commonalities. Beyond the single core, we extend our comparison by measuring data-transfer behavior through the memory hierarchy using a variety of microbenchmarks. We thoroughly investigate the “write-allocate (WA) evasion” feature, which can automatically reduce the memory traffic caused by write misses. We show that the Grace Superchip has a next-to-optimal implementation of WA evasion while the Sapphire Rapids CPU can avoid write allocates completely only in specific scenarios. The only way to eliminate WAs on AMD Genoa is the explicit use of non-temporal stores. Finally, we study the cache hierarchy of the CPUs in view of the Execution-Cache-Memory (ECM) performance model, revealing overlapping cache hierarchies on Genoa and Grace in contrast to Sapphire Rapids. Jan Laukemann, Georg Hager, Gerhard Wellein |
Parallel Comput. | 3 |
| 2024 | CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate EvasionabstractIn this paper we analyze the MPI-only version of the CloverLeaf code from the SPEChpc 2021 benchmark suite on recent Intel Xeon "Ice Lake" and "Sapphire Rapids" server CPUs. We observe peculiar breakdowns in performance when the number of processes is prime. Investigating this effect, we create first-principles data traffic models for each of the stencil-like hotspot loops. With application measurements and microbenchmarks to study memory data traffic behavior, we can connect the breakdowns to SpecI2M, a new write-allocate evasion feature in current Intel CPUs. For serial and full-node cases we are able to predict the memory data volume analytically with an error of a few percent. We find that if the number of processes is prime, SpecI2M fails to work properly, which we can attribute to short inner loops emerging from the one-dimensional domain decomposition in this case. We can also rule out other possible causes of the prime number effect, such as breaking layer conditions, MPI communication overhead, and load imbalance. Jan Laukemann, Thomas Gruber 0007, Georg Hager, Dossay Oryspayev, Gerhard Wellein |
IPDPS | 5 |
| 2024 | Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUsabstractThis paper addresses the challenge of providing portable and highly efficient code structures for CPU and GPU architectures. We choose the assembly of the right-hand term in the incompressible flow module of the High-Performance Computational Mechanics code Alya, which is one of the two CFD codes in the Unified European Benchmark Suite. Starting from an efficient CPU-code and a related OpenACC-port for GPUs we successively investigate performance potentials arising from code specialization, algorithmic restructuring and low-level optimizations.We demonstrate that only the combination of these different dimensions of runtime optimization unveils the full performance potential on the GPU and CPU. Roofline-based performance modelling is applied in this process and we demonstrate the need to investigate new optimization strategies if a classical roofline limit such as memory bandwidth utilization is achieved, rather than stopping the process. The final unified OpenACC-based implementation boosts performance by more than 50x on an NVIDIA A100 GPU (achieving approximately 2.5 TF/s FP64) and a further factor of 5x for an Intel Icelake based CPU-node (achieving approximately 1.0 TF/s FP64).The insights gained in our manual approach lays ground implementing unified but still highly efficient code structures for related kernels in Alya and other applications. These can be realized by manual coding or automatic code generation frameworks. Herbert Owen, Dominik Ernst, Thomas Gruber 0007, Oriol Lehmkuhl, Guillaume Houzeaux, Lucas Gasparino, Gerhard Wellein |
IPDPS | 7 |
| 2023 | Making applications faster by asynchronous execution: Slowing down processes or relaxing MPI collectives
Ayesha Afzal, Georg Hager, Stefano Markidis, Gerhard Wellein |
Future Gener. Comput. Syst. | 4 |
| 2023 | MD-Bench: A performance-focused prototyping harness for state-of-the-art short-range molecular dynamics algorithms
Rafael Ravedutti L. Machado, Jan Eitzinger, Jan Laukemann, Georg Hager, Harald Köstler, Gerhard Wellein |
Future Gener. Comput. Syst. | 6 |
| 2023 | Analytical performance estimation during code generation on modern GPUs
Dominik Ernst, Markus Holzer 0005, Georg Hager, Matthias Knorr 0002, Gerhard Wellein |
J. Parallel Distributed Comput. | 5 |
| 2023 | The Role of Idle Waves, Desynchronization, and Bottleneck Evasion in the Performance of Parallel ProgramsabstractThe performance of highly parallel applications on distributed-memory systems is influenced by many factors. Analytic performance modeling techniques aim to provide insight into performance limitations and are often the starting point of optimization efforts. However, coupling analytic models across the system hierarchy (socket, node, network) fails to encompass the intricate interplay between the program code and the hardware, especially when execution and communication bottlenecks are involved. In this paper we investigate the effect ofbottleneck evasionand how it can lead to automatic overlap of communication overhead with computation. Bottleneck evasion leads to a gradual loss of the initial bulk-synchronous behavior of a parallel code so that its processes become desynchronized. This occurs most prominently in memory-bound programs, which is why we choose memory-bound benchmark and application codes, specifically an MPI-augmented STREAM Triad, sparse matrix-vector multiplication, and a collective-avoiding Chebyshev filter diagonalization code to demonstrate the consequences of desynchronization on two different supercomputing platforms. We investigate the role of idle waves as possible triggers for desynchronization and show the impact of automatic asynchronous communication for a spectrum of code properties and parameters, such as saturation point, matrix structures, domain decomposition, and communication concurrency. Our findings reveal how eliminating synchronization points (such as collective communication or barriers) precipitates performance improvements that go beyond what can be expected by simply subtracting the overhead of the collective from the overall runtime. Ayesha Afzal, Georg Hager, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Level-Based Blocking for Sparse Matrices: Sparse Matrix-Power-Vector MultiplicationabstractThe multiplication of a sparse matrix with a dense vector (SpMV) is a key component in many numerical schemes and its performance is known to be severely limited by main memory access. Several numerical schemes require the multiplication of a sparse matrix polynomial with a dense vector which is typically implemented as a sequence of SpMVs. This results in low performance and ignores the potential to increase the arithmetic intensity by reusing the matrix data from cache. In this work we use the recursive algebraic coloring engine (RACE) to enable blocking of sparse matrix data across the polynomial computations. In the graph representing the sparse matrix we form levels using a breadth-first search. Locality relations of these levels are then used to improve spatial and temporal locality when accessing the matrix data and to implement an efficient multithreaded parallelization. Our approach is independent of the matrix structure and avoids shortcomings of existing “blocking” strategies in terms of hardware efficiency and parallelization overhead. We quantify the quality of our implementation using performance modelling and demonstrate speedups of up to 3× and 5× compared to an optimal SpMV-based baseline on a single multicore chip of recent Intel and AMD architectures. Various numerical schemes like$s$-step Krylov solvers, polynomial preconditioners and power clustering algorithms will benefit from our development. Christie L. Alappat, Georg Hager, Olaf Schenk, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Addressing White-box Modeling and Simulation Challenges in Parallel ComputingabstractNo abstract available. Ayesha Afzal, Gerhard Wellein, Georg Hager |
SIGSIM-PADS | 2 |
| 2022 | Analytic performance model for parallel overlapping memory-bound kernelsabstractAbstract Complex applications running on multicore processors show a rich performance phenomenology. The growing number of cores per ccNUMA domain complicates performance analysis of memory‐bound code since system noise, load imbalance, or task‐based programming models can lead to thread desynchronization. Hence, the simplifying assumption that all cores execute the same loop can not be upheld. Motivated by observations on plain and modified versions of the HPCG benchmark, we construct a performance model of execution of memory‐bound loop kernels. It can predict the memory bandwidth share per kernel on a memory contention domain depending on the number of active cores and which other workload the kernel is paired with. The only code features required are the single‐thread memory request fraction per kernel, which is directly related to the single‐thread memory bandwidth, and its saturated bandwidth. The former can either be measured directly or predicted using the Execution‐Cache‐Memory performance model. The computational intensity of the kernels and the detailed structure of the code is of no significance. We validate our model on Intel Broadwell, Intel Cascade Lake, and AMD Rome processors pairing various streaming and stencil kernels. The error in predicting the bandwidth share per kernel is less than 8%. Ayesha Afzal, Georg Hager, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Execution-Cache-Memory modeling and performance tuning of sparse matrix-vector multiplication and Lattice quantum chromodynamics on A64FXabstractAbstract The A64FX CPU is arguably the most powerful Arm‐based processor design to date. Although it is a traditional cache‐based multicore processor, its peak performance and memory bandwidth rival accelerator devices. A good understanding of its performance features is of paramount importance for developers who wish to leverage its full potential. We present an architectural analysis of the A64FX used in the Fujitsu FX1000 supercomputer at a level of detail that allows for the construction of Execution‐Cache‐Memory performance models for steady‐state loops. In the process we identify architectural peculiarities that point to viable generic optimization strategies. After validating the model using simple streaming loops we apply the insight gained to sparse matrix‐vector multiplication (SpMV) and the domain wall (DW) kernel from quantum chromodynamics. For SpMV we show why the compressed row storage (CRS) matrix storage format is not a good practical choice on this architecture and how the SELL‐C‐ format can achieve bandwidth saturation. For the DW kernel we provide a cache‐reuse analysis and show how an appropriate choice of data layout for complex arrays can realize memory‐bandwidth saturation in this case as well. A comparison with state‐of‐the‐art high‐end Intel Cascade Lake AP and Nvidia V100 systems puts the capabilities of the A64FX into perspective. We also explore the potential for power optimizations using the tuning knobs provided by the Fugaku system, achieving energy savings of about 31% for SpMV and 18% for DW. Christie L. Alappat, Nils Meyer, Jan Laukemann, Thomas Gruber 0007, Georg Hager, Gerhard Wellein, Tilo Wettig |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Multiway p-spectral graph cuts on Grassmann manifoldsabstractAbstract Nonlinear reformulations of the spectral clustering method have gained a lot of recent attention due to their increased numerical benefits and their solid mathematical background. We present a novel direct multiway spectral clustering algorithm in thep-norm, for $$p\in (1,2]$$ p∈(1,2] . The problem of computing multiple eigenvectors of the graphp-Laplacian, a nonlinear generalization of the standard graph Laplacian, is recasted as an unconstrained minimization problem on a Grassmann manifold. The value ofpis reduced in a pseudocontinuous manner, promoting sparser solution vectors that correspond to optimal graph cuts aspapproaches one. Monitoring the monotonic decrease of the balanced graph cuts guarantees that we obtain the best available solution from thep-levels considered. We demonstrate the effectiveness and accuracy of our algorithm in various artificial test-cases. Our numerical examples and comparative results with various state-of-the-art clustering methods indicate that the proposed method obtains high quality clusters both in terms of balanced graph cut metrics and in terms of the accuracy of the labelling assignment. Furthermore, we conduct studies for the classification of facial images and handwritten characters to demonstrate the applicability in real-world datasets. Dimosthenis Pasadakis, Christie L. Alappat, Olaf Schenk, Gerhard Wellein |
Mach. Learn. | 4 |
| 2021 | YaskSite: Stencil Optimization Techniques Applied to Explicit ODE Methods on Modern ArchitecturesabstractThe landscape of multi-core architectures is growing more complex and diverse. Optimal application performance tuning parameters can vary widely across CPUs, and finding them in a possibly multidimensional parameter search space can be time consuming, expensive and potentially infeasible. In this work, we introduce YaskSite, a tool capable of tackling these challenges for stencil computations. YaskSite is built upon Intel's YASK framework. It combines YASK's flexibility to deal with different target architectures with the Execution-Cache-Memory performance model, which enables identifying optimal performance parameters analytically without the need to run the code. Further we show that YaskSite's features can be exploited by external tuning frameworks to reliably select the most efficient kernel(s) for the application at hand. To demonstrate this, we integrate YaskSite into Offsite, an offline tuner for explicit ordinary differential equation methods, and show that the generated performance predictions are reliable and accurate, leading to considerable performance gains at minimal code generation time and autotuning costs on the latest Intel Cascade Lake and AMD Rome CPUs. Christie L. Alappat, Johannes Seiferth, Georg Hager, Matthias Korch, Thomas Rauber, Gerhard Wellein |
CGO | 6 |
| 2021 | Opening the Black Box: Performance Estimation during Code Generation for GPUsabstractAutomatic code generation is frequently used to create implementations of algorithms specifically tuned to particular hardware and application parameters. The code generation process involves the selection of adequate code transformations, tuning parameters, and parallelization strategies. To cover the huge search space, code generation frameworks may apply time-intensive autotuning, exploit scenario-specific performance models, or treat performance as an intangible black box that must be described via machine learning. This paper addresses the selection problem by identifying the relevant performance-defining mechanisms through a performance model coupled with an analytic hardware metric estimator. This enables a quick exploration of large configuration spaces to identify highly efficient candidates with high accuracy. Our current approach targets memory-intensive GPGPU applications and focuses on the correct modeling of data transfer volumes to all levels of the memory hierarchy. We show how our method can be coupled to the “pystencils” stencil code generator, which is used to generate kernels for a range four 3D25pt stencil and a complex two phase fluid solver based on the Lattice Boltzmann Method. For both, it delivers a ranking that can be used to select the best performing candidate. The method is not limited to stencil kernels, but can be integrated into any code generator that can generate the required address expressions. Dominik Ernst, Georg Hager, Matthias Knorr 0002, Gerhard Wellein, Markus Holzer 0005 |
SBAC-PAD | 4 |
| 2020 | PHIST: A Pipelined, Hybrid-Parallel Iterative Solver ToolkitabstractThe increasing complexity of hardware and software environments in high-performance computing poses big challenges on the development of sustainable and hardware-efficient numerical software. This article addresses these challenges in the context of sparse solvers. Existing solutions typically target sustainability, flexibility, or performance, but rarely all of them. Our new library PHIST provides implementations of solvers for sparse linear systems and eigenvalue problems. It is a productivity platform for performance-aware developers of algorithms and application software with abstractions that do not obscure the view on hardware-software interaction. The PHIST software architecture and the PHIST development process were designed to overcome shortcomings of existing packages. An interface layer for basic sparse linear algebra functionality that can be provided by multiple backends ensures sustainability, and PHIST supports common techniques for improving scalability and performance of algorithms such as blocking and kernel fusion. We showcase these concepts using the PHIST implementation of a block Jacobi-Davidson solver for non-Hermitian and generalized eigenproblems. We study its performance on a multi-core CPU, a GPU, and a large-scale many-core system. Furthermore, we show how an existing implementation of a block Krylov-Schur method in the Trilinos package Anasazi can benefit from the performance engineering techniques used in PHIST. Jonas Thies, Melven Röhrig-Zöllner, Nigel Overmars, Achim Basermann, Dominik Ernst, Georg Hager, Gerhard Wellein |
ACM Trans. Math. Softw. | 7 |
| 2019 | Propagation and Decay of Injected One-Off Delays on Clusters: A Case StudyabstractAnalytic, first-principles performance modeling of distributed-memory applications is difficult due to a wide spectrum of random disturbances caused by the application and the system. These disturbances (commonly called “noise”) run contrary to the assumptions about regularity that one usually employs when constructing simple analytic models. Despite numerous efforts to quantify, categorize, and reduce such effects, a comprehensive quantitative understanding of their performance impact is not available, especially for long, one-off delays of execution periods that have global consequences for the parallel application. In this work, we investigate various traces collected from synthetic benchmarks that mimic real applications on simulated and real message-passing systems in order to pin-point the mechanisms behind delay propagation. We analyze the dependence of the propagation speed of “idle waves,” i.e., propagating phases of inactivity, emanating from injected delays with respect to the execution and communication properties of the application, study how such delays decay under increased noise levels, and how they interact with each other. We also show how fine-grained noise can make a system immune against the adverse effects of propagating idle waves. Our results contribute to a better understanding of the collective phenomena that manifest themselves in distributed-memory parallel applications. Ayesha Afzal, Georg Hager, Gerhard Wellein |
CLUSTER | 3 |
| 2019 | ClusterCockpit - A web application for job-specific performance monitoringabstractMonitoring is a common component of HPC system software. Up to now, monitoring focused mainly on health checking and system level performance as well as on job scheduler information and was targeted towards system administrators. Recently job-specific performance monitoring based on hardware performance counter metrics has gained attention at academic HPC computing centers. HPC is becoming a mainstream tool that is also used by non-HPC experts, and HPC centers see a demand to check for pathological jobs and jobs with large optimization potential. The possibility to measure hardware performance counter data with negligible overhead allows assessment of efficient resource utilization and detection of pathological jobs. Pathological jobs are, e.g. jobs with errors in the batch script, jobs which do not terminate, jobs with severe load imbalance, or jobs that do not use any resources. This paper introduces ClusterCockpit, a web front-end tailor-made tool for job-specific performance monitoring. While many recent job-specific performance monitoring efforts concentrate on the measurement and data collection layers, ClusterCockpit provides a modern user interface targeted towards performance analysts as well as application users. Jan Eitzinger, Thomas Gruber 0007, Ayesha Afzal, Thomas Zeiser, Gerhard Wellein |
CLUSTER | 5 |
| 2019 | Code generation for massively parallel phase-field simulationsabstractThis article describes the development of automatic program generation technology to create scalable phase-field methods for material science applications. To simulate the formation of microstructures in metal alloys, we employ an advanced, thermodynamically consistent phase-field method. A state-of-the-art large-scale implementation of this model requires extensive, time-consuming, manual code optimization to achieve unprecedented fine mesh resolution. Our new approach starts with an abstract description based on free-energy functionals which is formally transformed into a continuous PDE and discretized automatically to obtain a stencil-based time-stepping scheme. Subsequently, an automatized performance engineering process generates highly optimized, performance-portable code for CPUs and GPUs. We demonstrate the efficiency for real-world simulations on large-scale GPU-based (PizDaint) and CPU-based (SuperMUC-NG) supercomputers. Our technique simplifies program development and optimization for a wide class of models. Martin Bauer 0003, Johannes Hötzer, Dominik Ernst, Julian Hammer, Marco Seiz, Henrik Hierl, Jan Hönig, Harald Köstler, Gerhard Wellein, Britta Nestler, Ulrich Rüde |
SC | 9 |
| 2019 | CRAFT: A Library for Easier Application-Level Checkpoint/Restart and Automatic Fault ToleranceabstractIn order to efficiently use the future generations of supercomputers, fault tolerance and power consumption are two of the prime challenges anticipated by the High Performance Computing (HPC) community. Checkpoint/Restart (CR) has been and still is the most widely used technique to deal with hard failures. Application-level CR is the most effective CR technique in terms of overhead efficiency but it takes a lot of implementation effort. This work presents the implementation of our C++ based library CRAFT (Checkpoint-Restart and Automatic Fault Tolerance), which serves two purposes. First, it provides an extendable library that significantly eases the implementation of application-level checkpointing. The most basic and frequently used checkpoint data-types are already part of CRAFT and can be directly used out of the box. The library can be easily extended to add more data-types. As means of overhead reduction, the library offers a built-in asynchronous checkpointing mechanism and also supports the Scalable Checkpoint/Restart (SCR) library for node level checkpointing. Second, CRAFT provides an easier interface for User-Level Failure Mitigation (ULFM) based dynamic process recovery, which significantly reduces the complexity and effort of failure detection and communication recovery mechanism. By utilizing both functionalities together, applications can write application-level checkpoints and recover dynamically from process failures with very limited programming effort. This work presents the challenges addressed by the library, its design, and its use. The associated overheads are analyzed using benchmarks. Faisal Shahzad 0001, Jonas Thies, Moritz Kreutzer, Thomas Zeiser, Georg Hager, Gerhard Wellein |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2018 | Multicore Performance Engineering of Sparse Triangular Solves Using a Modified Roofline ModelabstractThe Roofline model is widely used to visualize the performance of executed code together with the upper performance bounds given by the memory bandwidth and the processor peak performance. The model can thus provide an insightful visualization of bottlenecks. In this paper, we try to establish realistic bandwidth ceilings for the sparse triangular solve step of PARDISO, a leading sparse direct solver package, which is also part of the Intel MKL library. The performance of the forward and backward substitution process is analyzed and benchmarked for a representative set of sparse matrices on seven modern x86-type multicore architectures and the Knights Landing manycore architecture. It is shown how to accurately measure the necessary quantities also for threaded code, and the measurement approach, its validation, as well as limitations are discussed. Our modeling approach covers the serial and parallel execution phases, allowing for in-socket performance predictions. Markus Wittmann, Georg Hager, Radim Janalík, Martin Lanser, Axel Klawonn, Oliver Rheinbach, Olaf Schenk, Gerhard Wellein |
SBAC-PAD | 8 |
| 2017 | LIKWID Monitoring Stack: A Flexible Framework Enabling Job Specific Performance monitoring for the massesabstractSystem monitoring is an established tool to measure the utilization and health of HPC systems. Usually system monitoring infrastructures make no connection to job information and do not utilize hardware performance monitoring (HPM) data. To increase the efficient use of HPC systems automatic and continuous performance monitoring of jobs is an essential component. It can help to identify pathological cases, provides instant performance feedback to the users, offers initial data to judge on the optimization potential of applications and helps to build a statistical foundation about application specific system usage. The LIKWID monitoring stack is a modular framework build on top of the LIKWID tools library. It aims on enabling job specific performance monitoring using HPM data, system metrics and applicationlevel data for small to medium sized commodity clusters. Moreover, it is designed to integrate in existing monitoring infrastructures to speed up the change from pure system monitoring to job-aware monitoring. Thomas Röhl, Jan Eitzinger, Georg Hager, Gerhard Wellein |
CLUSTER | 4 |
| 2017 | Performance analysis of the Kahan-enhanced scalar product on current multi-core and many-core processorsabstractSummary We investigate the performance characteristics of a numerically enhanced scalar product (dot) kernel loop that uses the Kahan algorithm to compensate for numerical errors, and describe efficient single instruction multiple data‐vectorized implementations on recent multi‐core and many‐core processors. Using low‐level instruction analysis and the execution‐cache‐memory performance model, we pinpoint the relevant performance bottlenecks for single‐core and thread‐parallel execution and predict performance and saturation behavior. We show that the Kahan‐enhanced scalar product comes at almost no additional cost compared with the naive (non‐Kahan) scalar product if appropriate low‐level optimizations, notably single instruction multiple data vectorization and unrolling, are applied. The execution‐cache‐memory model is extended appropriately to accommodate not only modern Intel multicore chips but also the Intel Xeon Phi ‘Knights Corner’ coprocessor and an IBM POWER8 CPU. This allows us to discuss the impact of processor features on the performance across four modern architectures that are relevant for high performance computing. Copyright © 2016 John Wiley & Sons, Ltd. Johannes Hofmann 0001, Dietmar Fey, Michael Riedmann, Jan Eitzinger, Georg Hager, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 6 |
| 2017 | Preconditioned Krylov solvers on GPUs
Hartwig Anzt, Mark Gates, Jack J. Dongarra, Moritz Kreutzer, Gerhard Wellein, Martin Koehler |
Parallel Comput. | 5 |
| 2016 | Performance and power for highly parallel systemsabstractThis special issue is the result of an open call for papers initiated after the minisymposium ‘Analysis and Modeling: Techniques and Tools’, conducted at the Society for Industrial and Applied Mathematics (SIAM) Conference on Parallel Processing in Scientific Computing in Savannah, GA, in February 2012. The minisymposium brought together tool developers, performance and power modeling experts, and application analysts to present the state of the art on performance analysis and modeling techniques. This unique combination of expertise is highly needed in a time where complex, hierarchical architectures are the standard for all highly parallel computer systems. Without a good grip on relevant performance limitations, any ptimization attempt is just a shot in the dark. Hence, it is crucial to fully understand the performance properties and bottlenecks that come about with clustered multicore/many-core, multisocket nodes. Another aspect of modern systems is the complicated interplay between power constraints and the need for compute performance, which leads to complicated trade-offs. The challenges ahead are many-fold as systems scale in size. While parallelism is increasing, memory systems, interconnection networks, storage, and uncertainties in programming models all add to the complexities. More rapid realization of energy savings will require significant increases in measurement resolution and optimization techniques. This special issue is focused on how performance and power properties of modern highly parallel systems can be analyzed using state-of-the-art modeling and analysis techniques and real-world applications and tools. G. Hager, J. Treibig, J. Habich, and G. Wellein 1 introduce simple but insightful analytic models for execution performance and energy consumption of multicore CPUs. Automatic dynamic voltage and frequency scaling is leveraged by the “Green Queue” framework presented by J. Peraza, A. Tiwari, M. Laurenzano, L. Carrington, and A. E. Snavely 2 in their article. They show that significant energy savings at low performance loss are in reach if dynamic voltage and frequency scaling is used in an application-aware manner. A. D. Breslow, L. Porter, A. Tiwari, M. Laurenzano, L. Carrington, D. M. Tullsen, and A. E. Snavely 3 investigate the potential of job striping, a technique for co-locating HPC workloads with different characteristics on the same CPU chip, and demonstrate increased throughput and energy efficiency for a mix of typical simulation codes on a production cluster. The problem of how to deal with coarse-grained power measurements is tackled in the paper by H. Servat, G. Llort, J. Giménez, and J. Labarta 4. They present a tool that can derive fine-grained power and performance data for code with quickly alternating phases. The power usage and power variability of workloads on production supercomputers at Los Alamos National Laboratory are studied by S. Pakin, C. Storlie, M. Lang, R. E. Fields, E. E. Romero, C. Idler, S. Michalak, H. Greenberg, J. Loncaric, R. Rheinheimer, G. Grider, and J. Wendelberger in their paper 5. One of their central findings is that real power dissipation under real-world workloads is significantly lower than what the power infrastructure can handle, which opens interesting possibilities for saving cost via power capping. We think that this selection of papers is unique in providing several very different views on the problem of performance and power efficiency on present-day parallel machines from the core to the computing center level. Georg Hager, Darren J. Kerbyson, Abhinav Vishnu, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Exploring performance and power properties of modern multi-core chips via simple machine modelsabstractSummary Modern multi‐core chips show complex behavior with respect to performance and power. Starting with the Intel Sandy Bridge processor, it has become possible to directly measure the power dissipation of a CPU chip and correlate this data with the performance properties of the running code. Going beyond a simple bottleneck analysis, we employ the recently published Execution‐Cache‐Memory (ECM) model to describe the single‐core and multi‐core performance of streaming kernels. The model refines the well‐known roofline model, because it can predict the scaling and the saturation behavior of bandwidth‐limited loop kernels on a multi‐core chip. The saturation point is especially relevant for considerations of energy consumption. From power dissipation measurements of benchmark programs with vastly different requirements to the hardware, we derive a simple, phenomenological power model for the Sandy Bridge processor. Together with the ECM model, we are able to explain many peculiarities in the performance and power behavior of multi‐core processors and derive guidelines for energy‐efficient execution of parallel programs. Finally, we show that the ECM and power models can be successfully used to describe the scaling and power behavior of a lattice Boltzmann flow solver code. Copyright © 2013 John Wiley & Sons, Ltd. Georg Hager, Jan Eitzinger, Johannes Habich, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Chip-level and multi-node analysis of energy-optimized lattice Boltzmann CFD simulationsabstractSummary Memory‐bound algorithms show complex performance and energy consumption behavior on multicore processors. We choose the lattice Boltzmann method on an Intel Sandy Bridge cluster as a prototype scenario to investigate if and how single‐chip performance and power characteristics can be generalized to the highly parallel case. First, we perform an analysis of a sparse‐lattice lattice Boltzmann method implementation for complex geometries. Using a single‐core performance model, we predict the intra‐chip saturation characteristics and the optimal operating point in terms of energy‐to‐solution as a function of implementation details, clock frequency, vectorization, and number of active cores per chip. We show that high single‐core performance and a correct choice of the number of active cores per chip are the essential optimizations for the lowest energy‐to‐solution at minimal performance degradation. Then we extrapolate to the Message Passing Interface (MPI)‐parallel level and quantify the energy‐saving potential of various optimizations and execution modes, where we find these guidelines to be even more important, especially when communication overhead is non‐negligible. In our setup, we could achieve energy savings of 35% in this case, compared with a naive approach. We also demonstrate that a simple non‐reflective reduction of the clock speed leaves most of the energy‐saving potential unused. Copyright © 2015 John Wiley & Sons, Ltd. Markus Wittmann, Georg Hager, Thomas Zeiser, Jan Eitzinger, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Building a Fault Tolerant Application Using the GASPI Communication LayerabstractIt is commonly agreed that highly parallel software on Exascale computers will suffer from many more runtime failures due to the decreasing trend in the mean time to failures (MTTF). Therefore, it is not surprising that a lot of research is going on in the area of fault tolerance and fault mitigation. Applications should survive a failure and/or be able to recover with minimal cost. MPI is not yet very mature in handling failures, the User-Level Failure Mitigation (ULFM) proposal being currently the most promising approach is still in its prototype phase. In our work we use GASPI, which is a relatively new communication library based on the PGAS model. It provides the missing features to allow the design of fault-tolerant applications. Instead of introducing algorithm-based fault tolerance in its true sense, we demonstrate how we can build on (existing) clever checkpointing and extend applications to allow integrate a low cost fault detection mechanism and, if necessary, recover the application on the fly. The aspects of process management, the restoration of groups and the recovery mechanism is presented in detail. We use a sparse matrix vector multiplication based application to perform the analysis of the overhead introduced by such modifications. Our fault detection mechanism causes no overhead in failure-free cases, whereas in case of failure(s), the failure detection and recovery cost is of reasonably acceptable order and shows good scalability. Faisal Shahzad 0001, Moritz Kreutzer, Thomas Zeiser, Andreas Pieper, Georg Hager, Gerhard Wellein |
CLUSTER | 7 |
| 2015 | Quantifying Performance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory ModelabstractStencil algorithms on regular lattices appear in many fields of computational science, and much effort has been put into optimized implementations. Such activities are usually not guided by performance models that provide estimates of expected speedup. Understanding the performance properties and bottlenecks by performance modeling enables a clear view on promising optimization opportunities. In this work we refine the recently developed Execution-Cache-Memory (ECM) model and use it to quantify the performance bottlenecks of stencil algorithms on a contemporary Intel processor. This includes applying the model to arrive at single-core performance and scalability predictions for typical "corner case" stencil loop kernels. Guided by the ECM model we accurately quantify the significance of "layer conditions," which are required to estimate the data traffic through the memory hierarchy, and study the impact of typical optimization approaches such as spatial blocking, strength reduction, and temporal blocking for their expected benefits. We also compare the ECM model to the widely known Roofline model. Holger Stengel, Jan Eitzinger, Georg Hager, Gerhard Wellein |
ICS | 4 |
| 2015 | Performance Engineering of the Kernel Polynomal Method on Large-Scale CPU-GPU SystemsabstractThe Kernel Polynomial Method (KPM) is a well-established scheme in quantum physics and quantum chemistry to determine the Eigen value density and spectral properties of large sparse matrices. In this work we demonstrate the high optimization potential and feasibility of peta-scale heterogeneous CPU-GPU implementations of the KPM. At the node level we show that it is possible to decouple the sparse matrix problem posed by KPM from main memory bandwidth both on CPU and GPU. To alleviate the effects of scattered data access we combine loosely coupled outer iterations with tightly coupled block sparse matrix multiple vector operations, which enables pure data streaming. All optimizations are guided by a performance analysis and modelling process that indicates how the computational bottlenecks change with each optimization step. Finally we use the optimized node-level KPM with a hybrid-parallel framework to perform large-scale heterogeneous electronic structure calculations for novel topological materials on a pet scale-class Cray XC30 system. Moritz Kreutzer, Andreas Pieper, Georg Hager, Gerhard Wellein, Andreas Alvermann, Holger Fehske |
IPDPS | 4 |
| 2012 | Asynchronous Checkpointing by Dedicated Checkpoint Threads
Faisal Shahzad 0001, Markus Wittmann, Thomas Zeiser, Gerhard Wellein |
EuroMPI | 4 |
| 2011 | A flexible Patch-based lattice Boltzmann parallelization approach for heterogeneous GPU-CPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Georg Hager, Ulrich Rüde, Gerhard Wellein |
Parallel Comput. | 6 |
| 2009 | The world's fastest CPU and SMP node: Some performance results from the NEC SX-9abstractClassic vector systems have all but vanished from recent TOP500 lists. Looking at the newly introduced NEC SX-9 series, we benchmark its memory subsystem using the low level vector triad and employ an advanced lattice Boltzmann flow solver kernel to demonstrate that classic vectors still combine excellent performance with a well-established optimization approach. Results for commodity x86-based systems are provided for reference. Thomas Zeiser, Georg Hager, Gerhard Wellein |
IPDPS | 3 |
| 2008 | Data access optimizations for highly threaded multi-core CPUs with multiple memory controllersabstractProcessor and system architectures that feature multiple memory controllers are prone to show bottlenecks and erratic performance numbers on codes with regular access patterns. Although such effects are well known in the form of cache thrashing and aliasing conflicts, they become more severe when memory access is involved. Using the new Sun UltraSPARC T2 processor as a prototypical multi-core design, we analyze performance patterns in low-level and application benchmarks and show ways to circumvent bottlenecks by careful data layout and padding. Georg Hager, Thomas Zeiser, Gerhard Wellein |
IPDPS | 3 |
| 2004 | Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing ArchitecturesabstractComputationally intensive programs with moderate communication requirements such as CFD codes suffer from the standard slow interconnects of commodity "off the shelf" (COTS) hardware. We will introduce different large-scale applications of the Lattice Boltzmann Method (LBM) in fluid dynamics, material science, and chemical engineering and present results of the parallel performance on different architectures. It will be shown that a high speed communication network in combination with an efficient CPU is mandatory in order to achieve the required performance. An estimation of the necessary CPU count to meet the performance of 1 TFlop/s will be given as well as a prediction as to which architecture is the most suitable for LBM. Finally, ratios of costs to application performance for tailored HPC systems and COTS architectures will be presented. Thomas Pohl, Frank Deserno, Nils Thürey, Ulrich Rüde, Peter Lammers, Gerhard Wellein, Thomas Zeiser |
SC | 6 |