VLDB 2026 Research / reviewers in the wild / expert
Dirk Pleiter
dblp:70/3241
· DBLP profile ↗
21ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-7296-7817ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Portable BLAS3 Micro-Kernel Generator
Stepan Nassyr, Daniel Seibel, Prateek Chawla, Jayesh Badwaik, Andreas Herten, Dirk Pleiter |
Euro-Par (1) | 6 |
| 2025 | Performance-Portable Optimization and Analysis of Multiple Right-Hand Sides in a Lattice QCD SolverabstractManaging the high computational cost of iterative solvers for sparse linear systems is a known challenge in scientific computing. Moreover, scientific applications often face memory bandwidth constraints, making it critical to optimize data locality and enhance the efficiency of data transport. We extend the lattice QCD solver DD-$\alpha$AMG to incorporate multiple right-hand sides (rhs) for both the Wilson-Dirac operator evaluation and the GMRES solver, with and without odd-even preconditioning. To optimize auto-vectorization, we introduce a flexible interface that supports various data layouts and implement a new data layout for better SIMD utilization. We evaluate our optimizations on both x86 and Arm clusters, demonstrating performance portability with similar speedups. A key contribution of this work is the performance analysis of our optimizations, which reveals the complexity introduced by architectural constraints and compiler behavior. Additionally, we explore different implementations leveraging a new matrix instruction set for Arm called SME and provide an early assessment of its potential benefits. Shiting Long, Gustavo Ramirez-Hidalgo, Stepan Nassyr, Jose Jimenez-Merchan, Andreas Frommer, Dirk Pleiter |
HiPC | 6 |
| 2025 | The European master for HPC curriculumabstractInternational audience Pascal Bouvry, Mats Brorsson, Ramon Canal, Aryan Eftekhari, Siegfried Höfinger, Didier Smets, Harald Köstler, Tomás Kozubek, Ezhilmathi Krishnasamy, Josep Llosa, Alexandra Lukas-Rother, Xavier Martorell, Dirk Pleiter, Ana Proykova, Maria-Ribera Sancho, Olaf Schenk, Cristina Silvano |
J. Parallel Distributed Comput. | 13 |
| 2024 | Exploring Processor Micro-architectures Optimised for BLAS3 Micro-kernels
Stepan Nassyr, Dirk Pleiter |
Euro-Par (2) | 2 |
| 2023 | DECICE: Device-Edge-Cloud Intelligent Collaboration FrameworkabstractDECICE is a Horizon Europe project that is developing an AI-enabled open and portable management framework for automatic and adaptive optimization and deployment of applications in computing continuum encompassing from IoT sensors on the Edge to large-scale Cloud / HPC computing infrastructures. In this paper, we describe the DECICE framework and architecture. Furthermore, we highlight use-cases for framework evaluation: intelligent traffic intersection, magnetic resonance imaging, and emergency response. Julian M. Kunkel, Christian Boehme, Jonathan Decker, Fabrizio Magugliani, Dirk Pleiter, Bastian Koller, Karthee Sivalingam, Sabri Pllana, Alexander Nikolov, Müjdat Soytürk, Christian Racca, Andrea Bartolini, Adrian Tate, Berkay Yaman |
CF | 5 |
| 2022 | Assessing the State of Autovectorization Support based on SVEabstractSo-called SIMD instructions, which trigger operations that process in each clock cycle a data tuple, have become widespread in modern processor architectures. In particular, processors for high-performance computing (HPC) systems rely on this additional level of parallelism to reach a high throughput of arithmetic operations. Leveraging these SIMD instructions can still be challenging for application software developers. This challenge has become simpler due to a compiler technique called auto-vectorization. In this paper, we explore the current state of auto-vectorization capabilities using state-of-the-art compilers using a recent extension of the Arm instruction set architecture, called SVE. We measure the performance gains on a recent processor architecture supporting SVE, namely the Fujitsu A64FX processor. Bine Brank, Dirk Pleiter |
CLUSTER | 2 |
| 2022 | Strong Scaling of OpenACC enabled Nek5000 on several GPU based HPC systemsabstractWe present new results on the strong parallel scaling for the OpenACC-accelerated implementation of the high-order spectral element fluid dynamics solver Nek5000. The test case considered consists of a direct numerical simulation of fully-developed turbulent flow in a straight pipe, at two different Reynolds numbers Reτ = 360 and Reτ = 550, based on friction velocity and pipe radius. The strong scaling is tested on several GPU-enabled HPC systems, including the Swiss Piz Daint system, TACC’s Longhorn, Jülich’s JUWELS Booster, and Berzelius in Sweden. The performance results show that speed-up between 3-5 can be achieved using the GPU accelerated version compared with the CPU version on these different systems. The run-time for 20 timesteps reduces from 43.5 to 13.2 seconds with increasing the number of GPUs from 64 to 512 for Reτ = 550 case on JUWELS Booster system. This illustrates the GPU accelerated version the potential for high throughput. At the same time, the strong scaling limit is significantly larger for GPUs, at about 2000 − 5000 elements per rank; compared to about 50 − 100 for a CPU-rank. Jonathan Vincent, Martin Karp, Adam Peplinski, Niclas Jansson, Artur Podobas, Andreas Jocksch, Fazle Hussain, Stefano Markidis, Matts Karlsson, Dirk Pleiter, Erwin Laure, Philipp Schlatter |
HPC Asia | 12 |
| 2021 | Mont-Blanc 2020: Towards Scalable and Power Efficient European HPC ProcessorsabstractThe Mont-Blanc 2020 (MB2020) project has triggered the development of the next generation industrial processor for Big Data and High Performance Computing (HPC). MB2020 is paving the way to the future low-power European processor for exascale, defining the System-on-Chip (SoC) architecture and implementing new critical building blocks to be integrated in such an SoC. In this paper, we first present an overview of the MB2020 project, then we describe our experimental infrastructure, the requirements of relevant applications, and the IP blocks developed in the project. Finally, we present our emulation-based final demonstrator and explain how it integrates within our first generation of HPC processors. Adrià Armejach, Bine Brank, Jordi Cortina, François Dolique, Timothy Hayes 0001, Nam Ho, Pierre-Axel Lagadec, Romain Lemaire, Guillem López-Paradís, Laurent Marliac, Miquel Moretó, Pedro Marcuello, Dirk Pleiter, Xubin Tan, Said Derradji |
DATE | 13 |
| 2020 | Porting Applications to Arm-based ProcessorsabstractArm-based server processors are becoming increasingly used for building massively parallel HPC systems. This triggers the need for porting HPC applications to such architectures and to collect more knowledge about performance benefits and challenges. In this contribution, we report on experiences made during our ongoing efforts to port applications to Arm-based platforms at the Jülich Supercomputing Centre (JSC). The performance of these applications is explored on different Arm-based node architectures plus an x86-based architecture for reference. Bine Brank, Stepan Nassyr, Fatemeh Pouyan, Dirk Pleiter |
CLUSTER | 4 |
| 2020 | Performance Evaluation of ParalleX Execution model on Arm-based PlatformsabstractThe HPC community shows a keen interest in creating diversity in the CPU ecosystem. The advent of Arm-based processors provides an alternative to the existing HPC ecosystem, which is primarily dominated by x86 processors. In this paper, we port an Asynchronous Many-Task runtime system based on the ParalleX model, i.e., High Performance ParalleX (HPX), and evaluate it on the Arm ecosystem with a suite of benchmarks. We wrote these benchmarks with an emphasis on vectorization and distributed scaling. We present the performance results on a variety of Arm processors and compare it with their x86 brethren from Intel. We show that the results obtained are equally good or better than their x86 brethren. Finally, we also discuss a few drawbacks of the present Arm ecosystem. Nikunj Gupta, Rohit Ashiwal, Bine Brank, Sateesh K. Peddoju, Dirk Pleiter |
CLUSTER | 5 |
| 2019 | IO Challenges for Human Brain Atlasing Using Deep Learning Methods - An In-Depth AnalysisabstractThe use of Deep Learning methods have been identified as a key opportunity for enabling processing of extreme-scale scientific datasets. Feeding data into compute nodes equipped with several high-end GPUs at sufficiently high rate is a known challenge. Facilitating processing of these datasets thus requires the ability to store petabytes of data as well as to access the data with very high bandwidth. In this work, we look at two Deep Learning use cases for cytoarchitectonic brain mapping. These applications are very challenging for the underlying IO system. We present an in depth analysis of their IO requirements and performance. Both applications are limited by the IO performance, as the training processes often have to wait several seconds for new training data. Both applications read random patches from a collection of large HDF5 datasets or TIFF files, which result in many small non-consecutive accesses to the parallel file systems. By using a chunked data format or storing temporally copies of the required patches, the IO performance can be improved significantly. These leads to a decrease of the total runtime of up to 80%. Lena Oden, Christian Schiffer, Hannah Spitzer, Timo Dickscheid, Dirk Pleiter |
PDP | 5 |
| 2019 | Performance of ODROID-MC1 for scientific flow problems
Andreas Lintermann, Dirk Pleiter, Wolfgang Schröder 0001 |
Future Gener. Comput. Syst. | 2 |
| 2019 | SAGE: Percipient Storage for Exascale Data Centric Computing
Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Stefano Markidis, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Dirk Pleiter, Shaun De Witt |
Parallel Comput. | 9 |
| 2018 | The SAGE project: a storage centric approach for exascale computing: invited paperabstractSAGE (Percipient StorAGe for Exascale Data Centric Computing) is a European Commission funded project towards the era of Exascale computing. Its goal is to design and implement a Big Data/Extreme Computing (BDEC) capable infrastructure with associated software stack. The SAGE system follows a storage centric approach as it is capable of storing and processing large data volumes at the Exascale regime. Sai Narasimhamurthy, Nikita Danilov, Sining Wu, Ganesan Umanesan, Steven W. D. Chien, Sergio Rivas-Gomez, Ivy Bo Peng, Erwin Laure, Shaun De Witt, Dirk Pleiter, Stefano Markidis |
CF | 10 |
| 2018 | SVE-Enabling Lattice QCD CodesabstractOptimization of applications for supercomputers of the highest performance class requires parallelization at multiple levels using different techniques. In this contribution we focus on parallelization of particle physics simulations through vector instructions. With the advent of the Scalable Vector Extension (SVE) ISA, future ARM-based processors are expected to provide a significant level of parallelism at this level. Nils Meyer, Peter Georg, Dirk Pleiter, Stefan Solbrig, Tilo Wettig |
CLUSTER | 3 |
| 2018 | Mainstream vs. Emerging HPC: Metrics, Trade-Offs and Lessons LearnedabstractVarious servers with different characteristics and architectures are hitting the market, and their evaluation and comparison in terms of HPC features is complex and multidimensional. In this paper, we share our experience of evaluating a diverse set of HPC systems, consisting of three mainstream and five emerging architectures. We evaluate the performance and power efficiency using prominent HPC benchmarks, High-Performance Linpack (HPL) and High Performance Conjugate Gradients (HPCG), and expand our analysis using publicly available specialized kernel benchmarks, targeting specific system components. In addition to a large body of quantitative results, we emphasize six usually overlooked aspects of the HPC platforms evaluation, and share our conclusions and lessons learned. Overall, we believe that this paper will improve the evaluation and comparison of HPC platforms, making a first step towards a more reliable and uniform methodology. Milan Radulovic, Kazi Asifuzzaman, Darko Zivanovic, Nikola Rajovic, Guillaume Colin de Verdière, Dirk Pleiter, Manolis Marazakis, Nikolaos D. Kallimanis, Paul M. Carpenter, Petar Radojkovic, Eduard Ayguadé |
SBAC-PAD | 6 |
| 2017 | Paving the Way Towards a Highly Energy-Efficient and Highly Integrated Compute Node for the Exascale Revolution: The ExaNoDe ApproachabstractPower consumption and high compute density are the key factors to be considered when building a compute node for the upcoming Exascale revolution. Current architectural design and manufacturing technologies are not able to provide the requested level of density and power efficiency to realise an operational Exascale machine. A disruptive change in the hardware design and integration process is needed in order to cope with the requirements of this forthcoming computing target. This paper presents the ExaNoDe H2020 research project aiming to design a highly energy efficient and highly integrated heterogeneous compute node targeting Exascale level computing, mixing low-power processors, heterogeneous co-processors and using advanced hardware integration technologies with the novel UNIMEM Global Address Space memory system. Alvise Rigo, Christian Pinto, Kevin Pouget, Daniel Raho, Denis Dutoit, Pierre-Yves Martinez, Chris Doran, Luca Benini, Iakovos Mavroidis, Manolis Marazakis, Valeria Bartsch, Guy Lonsdale, Antoniu Pop, John Goodacre, Annaik Colliot, Paul M. Carpenter, Petar Radojkovic, Dirk Pleiter, Dominique Drouin, Benoît Dupont de Dinechin |
DSD | 18 |
| 2016 | Addressing Materials Science Challenges Using GPU-accelerated POWER8 Nodes
Paul F. Baumeister, Marcel Bornemann, Markus Bühler, Thorsten Hater, Benjamin Krill, Dirk Pleiter, Rudolf Zeller |
Euro-Par | 6 |
| 2015 | A Performance Model for GPU-Accelerated FDTD ApplicationsabstractIn this work we develop, validate and use a performance model for a Finite-Difference Time-Domain (FDTD) application which is parallelized on multiple GPUs. FDTD is a method for simulating electrodynamic interaction and is applied in a number of research and engineering areas. In this work we focus on a particular implementation called B-CALM (Belgium-California Light Machine). We adopt a simple, semi-empirical modelling approach to design a model which we validate for different hardware architectures. Using the model allows making implementation decisions and exploring the architectural design space with the goal of optimizing HPC systems for this application. Paul F. Baumeister, Thorsten Hater, Jiri Kraus, Dirk Pleiter, Pierre Wahl |
HiPC | 4 |
| 2014 | Modeling CPU Energy Consumption of HPC Applications on the IBM POWER7abstractEnergy consumption optimization of HPC applications inherently requires measurements for reference and comparison. However, most of today's systems lack the necessary hardware support for power or energy measurements. Furthermore, in-band data availability is preferred for specific optimization techniques such as auto-tuning. For this reason, we present in-band energy consumption models for the IBM POWER7 processor based on hardware counters. We demonstrate that linear regression is a suitable means for modeling energy consumption, and we rely on already available, high-level benchmarks for training instead of self-written or hand-tuned micro-kernels. We compare modeling efforts for different instruction mixes caused by two compilers (GCC and IBM XL) as well as various multi-threading usage scenarios, and validate across our training benchmarks and two real-world applications. Results show mean errors of approximately 1% and overall max errors of 5.3% for GCC. Philipp Gschwandtner, Michael Knobloch, Bernd Mohr, Dirk Pleiter, Thomas Fahringer |
PDP | 4 |
| 2013 | GPUMAFIA: Efficient Subspace Clustering with MAFIA on GPUs
Andrew V. Adinetz, Jiri Kraus, Jan H. Meinke, Dirk Pleiter |
Euro-Par | 4 |