VLDB 2026 Research / reviewers in the wild / expert
Dirk Pflüger
dblp:96/5029
· DBLP profile ↗
24ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0002-4360-0212ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Safety and Reliability: Validation for Automated Driving Functions through Scenario-Based TestingabstractThis paper presents an infrastructure for defining scenarios and utilizing them to validate automated driving systems. It addresses various aspects of scenario-based testing, with a focus on lane keeping, lane changing, and traffic light scenarios. We define the scenarios using the OpenSCENARIO 2.0 format, as well as directly through Python scripts. These scenarios are integrated into two distinct simulators: an in-house simulator based on the Intelligent Driver Model (IDM), and the CARLA simulator. In these simulators, two agents are subjected to a range of challenging conditions, and the risk of failure is assessed. This evaluation provides insights into the agents' performance and their safety compliance, acting as a benchmark for safety assessment in the different scenarios. Naya Baslan, Alexander Kerschl, Julian Schmidt, Dirk Pflüger |
IV | 4 |
| 2024 | Training Large Language Models for System-Level Test Program Generation Targeting Non-functional PropertiesabstractSystem-Level Test (SLT) has been an integral part of integrated circuit test flows for over a decade and continues to be significant. Nevertheless, there is a lack of systematic approaches for generating test programs, specifically focusing on the non-functional aspects of the Device under Test (DUT). Currently, test engineers manually create test suites using commercially available software to simulate the end-user environment of the DUT. This process is challenging and laborious and does not assure adequate control over non-functional properties. This paper proposes to use Large Language Models (LLMs) for SLT program generation. We use a pre-trained LLM and fine-tune it to generate test programs that optimize non-functional properties of the DUT, e.g., instructions per cycle. Therefore, we use Gem5, a microarchitectural simulator, in conjunction with Reinforcement Learning-based training. Finally, we write a prompt to generate C code snippets that maximize the instructions per cycle of the given architecture. In addition, we apply hyperparameter optimization to achieve the best possible results in inference. Denis Schwachhofer, Peter Domanski, Steffen Becker 0001, Stefan Wagner 0001, Matthias Sauer 0002, Dirk Pflüger, Ilia Polian |
ETS | 6 |
| 2024 | Short Paper: Evaluation of stdpar Compilers on a Kernel Matrix Assembly and a BLAS Level 3 SYMM KernelabstractThe introduction of stdpar in C++ introduces parallel algorithms into the standard library, which allows developers to easily harness the power of parallelism in their applications, leading to potential performance improvements, scalability, increased productivity, and improved reproducibility in scientific computing tasks. At least theoretically, as its performance depends on the compiler in use. In this paper, we examine the usability of stdpar and demonstrate the significant differences between different stdpar compilers. To this end, we study a kernel matrix assembly algorithm and a BLAS level 3 SYMM function, both implemented in the Parallel Least Squares Support Vector Machine (PLSSVM) library. The same code is compiled with an NVIDIA nvc++, gcc, Intel oneAPI icpx, and AdaptiveCpp compiler. We analyzed the code on an NVIDIA A30 GPU, an AMD MI210 GPU, and a dual-socket AMD EPYC 9274F CPU machine. First, we report surprisingly large runtime differences between the different stdpar compilers: up to 11 percent on the NVIDIA A30 GPU and even up to 335% on two AMD EPYC 9274F. Second, we show how stdpar enables energy reduction by a factor of up to 20.7 with the very same implementation by utilizing different hardware platforms.The code, utility scripts, and documentation are all available on GitHub: https://github.com/SC-SGS/PLSSVM Marcel Breyer, Alexander Van Craen, Dirk Pflüger |
ISPDC | 3 |
| 2024 | Realizing Joint Extreme-Scale Simulations on Multiple Supercomputers - Two Superfacility Case StudiesabstractHigh-dimensional grid-based simulations serve as both a tool and a challenge in researching various domains. The main challenge of these approaches is the well-known curse of dimensionality, amplified by the need for fine resolutions in high-fidelity applications. The combination technique (CT) provides a straightforward way of performing such simulations while alleviating the curse of dimensionality. Recent work demonstrated the potential of the CT to join multiple systems simultaneously to perform a single high-dimensional simulation. This paper shows how to extend this to three or more systems and addresses some remaining challenges: load balancing on heterogeneous hardware; utilizing compression to maximize the communication bandwidth; efficient I/O management through hardware mapping; and improving memory utilization through algorithmic optimizations. Combining these contributions, we demonstrate the feasibility of the CT for extreme-scale Superfacility scenarios of 46 trillion DOF on two systems and 35 trillion DOF on three systems. Scenarios at these resolutions would be intractable with full-grid solvers ($\gt1,000$ nonillion DOF each). Theresa Pollinger, Alexander Van Craen, Philipp Offenhäuser, Dirk Pflüger |
SC | 4 |
| 2024 | Simulating stellar merger using HPX/Kokkos on A64FX on Supercomputer Fugaku
Patrick Diehl, Gregor Daiß, Kevin A. Huck, Dominic Marcello, Sagiv Shiber, Hartmut Kaiser, Dirk Pflüger |
J. Supercomput. | 7 |
| 2023 | Learn to Tune: Robust Performance Tuning in Post-Silicon ValidationabstractPost-silicon validation is a crucial yet challenging problem primarily due to the increasing complexity of the semi-conductor value chain. Existing techniques cannot keep up with the rapid increase in the complexity of designs. Therefore, post-silicon validation is becoming an expensive bottleneck. Robust performance tuning is relevant to compensate impacts of process variations and non-ideal design implementations. We propose a novel approach based on Deep Reinforcement Learning and Learn to Optimize. The method automatically learns flexible tuning strategies tailored to specific circuits. Additionally, it addresses high-dimensional tuning tasks, including mixed data types and dependencies, e.g., on operating conditions. In this work, we introduce Learn to Tune and demonstrate its appealing properties in post-silicon validation, e.g., lower computational cost or faster time-to-optimize, allowing a more efficient adaption of the tuning to changing tuning conditions than classical methods. Peter Domanski, Dirk Pflüger, Raphaël Latty |
ETS | 2 |
| 2023 | Leveraging the Compute Power of Two HPC Systems for Higher-Dimensional Grid-Based Simulations with the Widely-Distributed Sparse Grid Combination TechniqueabstractGrid-based simulations of hot fusion plasmas are often severely limited by computational and memory resources; the grids live in four- to six-dimensional space and thus suffer the curse of dimensionality. However, high resolutions are required to fully capture the physics of interest. The sparse grid combination technique is a multi-scale method in which many anisotropically coarse resolved grids are used to approximate a fine-scale solution---and it alleviates the curse of dimensionality. Theresa Pollinger, Alexander Van Craen, Christoph Niethammer, Marcel Breyer, Dirk Pflüger |
SC | 5 |
| 2022 | Intelligent Methods for Test and ReliabilityabstractTest methods that can keep up with the ongoing increase in complexity of semiconductor products and their underlying technologies are an essential prerequisite for maintaining quality and safety of our daily lives and for continued success of our economies and societies. There is a huge potential how test methods can benefit from recent breakthroughs in domains such as artificial intelligence, data analytics, virtual/augmented reality, and security. The Graduate School on “Intelligent Methods for Semiconductor Test and Reliability” (GS-IMTR) at the University of Stuttgart is a large-scale, radically interdisciplinary effort to address the scientific-technological challenges in this domain. It is funded by Advantest, one of the world leaders in automatic test equipment. In this paper, we describe the overall philosophy of the Graduate School and the specific scientific questions targeted by its ten projects. Hussam Amrouch, Jens Anders, Steffen Becker 0001, Maik Betka, Gerd Bleher, Peter Domanski, Nourhan Elhamawy, Thomas Ertl, Athanasios Gatzastras, Paul R. Genssler, Sebastian Hasler, Martin Heinrich, André van Hoorn, Hanieh Jafarzadeh, Ingmar Kallfass, Florian Klemme, Steffen Koch 0001, Ralf Küsters, Andrés Lalama, Raphaël Latty, Yiwen Liao, Natalia Lylina, Zahra Paria Najafi-Haghi, Dirk Pflüger, Ilia Polian, Jochen Rivoir, Matthias Sauer 0002, Denis Schwachhofer, Steffen Templin, Christian Volmer, Stefan Wagner 0001, Daniel Weiskopf, Hans-Joachim Wunderlich, Bin Yang 0009 |
DATE | 24 |
| 2022 | Machine Learning for Test, Diagnosis, Post-Silicon Validation and Yield OptimizationabstractRecent breakthroughs in machine learning (ML) technology are shifting the boundaries of what is technologically possible in several areas of Computer Science and Engineering. This paper discusses ML in the context of test-related activities, including fault diagnosis, post-silicon validation and yield optimization. ML is by now an established scientific discipline, and a large number of successful ML techniques have been developed over the years. This paper focuses on how to adapt ML approaches that were originally developed with other applications in mind to test-related problems. We consider two specific applications of learning in more depth: delay fault diagnosis in three-dimensional integrated circuits and tuning performed during post-silicon validation. Moreover, we examine the emerging concept of brain-inspired hyperdimensional computing (HDC) and its potential for addressing test and reliability questions. Finally, we show how to integrate ML into actual industrial test and yield-optimization flows. Hussam Amrouch, Krishnendu Chakrabarty, Dirk Pflüger, Ilia Polian, Matthias Sauer 0002, Matteo Sonza Reorda |
ETS | 3 |
| 2022 | PDEBench: An Extensive Benchmark for Scientific Machine LearningabstractMachine learning-based modeling of physical systems has experienced increased interest in recent years. Despite some impressive progress, there is still a lack of benchmarks for Scientific ML that are easy to use but still challenging and repre- sentative of a wide range of problems. We introduce PDEBENCH, a benchmark suite of time-dependent simulation tasks based on Partial Differential Equations (PDEs). PDEBENCH comprises both code and data to benchmark the performance of novel machine learning models against both classical numerical simulations and machine learning baselines. Our proposed set of benchmark problems con- tribute the following unique features: (1) A much wider range of PDEs compared to existing benchmarks, ranging from relatively common examples to more real- istic and difficult problems; (2) much larger ready-to-use datasets compared to prior work, comprising multiple simulation runs across a larger number of ini- tial and boundary conditions and PDE parameters; (3) more extensible source codes with user-friendly APIs for data generation and baseline results with popular machine learning models (FNO, U-Net, PINN, Gradient-Based Inverse Method). PDEBENCH allows researchers to extend the benchmark freely for their own pur- poses using a standardized API and to compare the performance of new models to existing baseline methods. We also propose new evaluation metrics with the aim to provide a more holistic understanding of learning methods in the context of Scientific ML. With those metrics we identify tasks which are challenging for recent ML methods and propose these tasks as future challenges for the community. The code is available at https://github.com/pdebench/PDEBench. Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Daniel MacKinlay, Francesco Alesiani, Dirk Pflüger, Mathias Niepert |
NeurIPS | 6 |
| 2021 | Octo-Tiger's New Hydro Module and Performance Using HPX+CUDA on ORNL's SummitabstractOcto-Tiger is a code for modeling three-dimensional self-gravitating astrophysical fluids. It was particularly designed for the study of dynamical mass transfer between interacting binary stars. Octo-Tiger is parallelized for distributed systems using the asynchronous many-task runtime system, the C++ standard library for parallelism and concurrency (HPX) and utilizes CUDA for its gravity solver. Recently, we have remodeled Octo-Tiger’s hydro solver to use a three-dimensional reconstruction scheme. In addition, we have ported the hydro solver to GPU using CUDA kernels. We present scaling results for the new hydro kernels on ORNL’s Summit machine using a Sedov-Taylor blast wave problem. We also compare Octo-Tiger’s new hydro scheme with its old hydro scheme, using a rotating star as a test problem. Patrick Diehl, Gregor Daiß, Dominic Marcello, Kevin A. Huck, Sagiv Shiber, Hartmut Kaiser, Juhan Frank, Geoffrey C. Clayton, Dirk Pflüger |
CLUSTER | 9 |
| 2021 | ORSA: Outlier Robust Stacked Aggregation for Best- and Worst-Case Approximations of Ensemble SystemsabstractIn recent years, the usage of ensemble learning in applications has grown significantly due to increasing computational power allowing the training of large ensembles in reasonable time frames. Many applications, e.g., malware detection, face recognition, or financial decision-making, use a finite set of learning algorithms and do aggregate them in a way that a better predictive performance is obtained than any other of the individual learning algorithms. In the field of Post-Silicon Validation for semiconductor devices (PSV), data sets are typically provided that consist of various devices like, e.g., chips of different manufacturing lines. In PSV, the task is to approximate the underlying function of the data with multiple learning algorithms, each trained on a device-specific subset, instead of improving the performance of arbitrary classifiers on the entire data set. Furthermore, the expectation is that an unknown number of subsets describe functions showing very different characteristics. Corresponding ensemble members, which are called outliers, can heavily influence the approximation. Our method aims to find a suitable approximation that is robust to outliers and represents the best or worst case in a way that will apply to as many types as possible. A ‘softmax’ or ‘soft-min’ function is used in place of a maximum or minimum operator. A Neural Network (NN) is trained to learn this ‘soft-function’ in a two-stage process. First, we select a subset of ensemble members that is representative of the best or worst case. Second, we combine these members and define a weighting that uses the properties of the Local Outlier Factor (LOF) to increase the influence of non-outliers and to decrease outliers. The weighting ensures robustness to outliers and makes sure that approximations are suitable for most types. Peter Domanski, Dirk Pflüger, Raphaël Latty, Jochen Rivoir |
ICMLA | 2 |
| 2021 | Learning Free-Surface Flow with Physics-Informed Neural NetworksabstractThe interface between data-driven learning methods and classical simulation poses an interesting field offering a multitude of new applications. In this work, we build on the notion of physics-informed neural networks (PINNs) and employ them in the area of shallow-water equation (SWE) models. These models play an important role in modeling and simulating free-surface flow scenarios such as in flood-wave propagation or tsunami waves. Different formulations of the PINN residual are compared to each other and multiple optimizations are being evaluated to speed up the convergence rate. We test these with different 1-D and 2-D experiments and finally demonstrate that regarding a SWE scenario with varying bathymetry, the method is able to produce competitive results in comparison to the direct numerical simulation with a total relative L2error of 8.9e−3. Raphael Leiteritz, Marcel Hurler, Dirk Pflüger |
ICMLA | 3 |
| 2019 | From piz daint to the stars: simulation of stellar mergers using high-level abstractionsabstractWe study the simulation of stellar mergers, which requires complex simulations with high computational demands. We have developed Octo-Tiger, a finite volume grid-based hydrodynamics simulation code with Adaptive Mesh Refinement which is unique in conserving both linear and angular momentum to machine precision. To face the challenge of increasingly complex, diverse, and heterogeneous HPC systems, Octo-Tiger relies on high-level programming abstractions. Gregor Daiß, Parsa Amini, John Biddiscombe, Patrick Diehl, Juhan Frank, Kevin A. Huck, Hartmut Kaiser, Dominic Marcello, David Pfander, Dirk Pflüger |
SC | 10 |
| 2017 | Visualization of fracture progression in peridynamics
Michael Bußler, Patrick Diehl, Dirk Pflüger, Steffen Frey, Filip Sadlo, Thomas Ertl, Marc Alexander Schweitzer |
Comput. Graph. | 3 |
| 2016 | Data mining on vast data sets as a cluster system benchmarkabstractSummary Comparing different (accelerated) cluster architectures by a single application is a tough piece of work because this application has to be optimized with respect to platform‐dependent features. In this work, we demonstrate such an optimization for a data mining algorithm which solves regression and classification problems on vast data sets. Our technique is based on least squares regression, and its major component is the iterative matrix‐free solution of a linear system of equations. By processing data sets ranging from several hundreds of thousands instances to multi‐million data points in strong‐scaling and weak‐scaling settings, we are able to estimate the amount of parallelism needed to unleash the performance of classic CPU‐based machines and clusters employing Intel Xeon Phi coprocessors and NVIDIA Kepler GPUs. Only in strong‐scaling experiments, GPUs and coprocessors suffer from their tremendous amount of needed parallelism and get outperformed by dual socket Intel Sandy Bridge nodes at large scale (more than 64 nodes/accelerators). However, in weak‐scaling scenarios, a speed‐up larger than 2X over an entire CPU node can be achieved by a single accelerator. Copyright © 2015 John Wiley & Sons, Ltd. Alexander Heinecke, Roman Karlstetter, Dirk Pflüger, Hans-Joachim Bungartz |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Global communication schemes for the numerical solution of high-dimensional PDEs
Philipp Hupp, Mario Heene, Riko Jacob, Dirk Pflüger |
Parallel Comput. | 4 |
| 2014 | Density Estimation with Adaptive Sparse Grids for Large Data SetsabstractMultifidelity approximation is an important technique in scientific computation and simulation. In this paper, we introduce a bandit-learning approach for leveraging data of varying fidelities to achieve precise estimates of the parameters of interest. Under a linear model assumption, we formulate a multifidelity approximation as a modified stochastic bandit and analyze the loss for a class of policies that uniformly explore each model before exploiting. Utilizing the estimated conditional mean-squared error, we propose a consistent algorithm, adaptive explore-then-commit (AETC), and establish a corresponding trajectorywise optimality result. These results are then extended to the case of vector-valued responses, where we demonstrate that the algorithm is efficient without the need to worry about estimating high-dimensional parameters. The main advantage of our approach is that we require neither hierarchical model structure nor a priori knowledge of statistical information (e.g., correlations) about or between models. Instead, the AETC algorithm requires only knowledge of which model is a trusted high-fidelity model, along with (relative) computational cost estimates of querying each model. Numerical experiments are provided at the end to support our theoretical findings. Benjamin Peherstorfer, Dirk Pflüger, Hans-Joachim Bungartz |
SDM | 2 |
| 2014 | Parallelizing a Black-Scholes solver based on finite elements and sparse gridsabstractSUMMARY We present the parallelization of a sparse grid finite element discretization of the Black–Scholes equation, which is commonly used for option pricing. Sparse grids allow to handle higher dimensional options than classical approaches on full grids and can be extended to a fully adaptive discretization method. We introduce the algorithmical structure of efficient algorithms operating on sparse grids and demonstrate how they can be used to derive an efficient parallelization with OpenMP of the Black–Scholes solver. We show results on different commodity hardware systems based on multi‐core architectures with up to 24 cores and discuss the parallel performance using Intel and Advanced Micro Devices (AMD) CPUs. Copyright © 2012 John Wiley & Sons, Ltd. Hans-Joachim Bungartz, Alexander Heinecke, Dirk Pflüger, Stefanie Schraufstetter |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Fast Insight into High-Dimensional Parametrized Simulation DataabstractNumerical simulation has become an inevitable tool in most industrial product development processes with simulations being used to understand the influence of design decisions (parameter configurations) on the structure and properties of the product. However, in order to allow the engineer to thoroughly explore the design space and fine-tune parameters, many -- usually very time-consuming -- simulation runs are necessary. Additionally, this results in a huge amount of data that cannot be analyzed in an efficient way without the support of appropriate tools. In this paper, we address the two-fold problem: First, instantly provide simulation results if the parameter configuration is changed, and, second, identify specific areas of the design space with concentrated change and thus importance. We propose the use of a hierarchical approach based on sparse grid interpolation or regression which acts as an efficient and cheap substitute for the simulation. Furthermore, we develop new visual representations based on the derivative information contained inherently in the hierarchical basis. They intuitively let a user identify interesting parameter regions even in higher-dimensional settings. This workflow is combined in an interactive visualization and exploration framework. We discuss examples from different fields of computational science and engineering and show how our sparse-grid-based techniques make parameter dependencies apparent and how they can be used to fine-tune parameter configurations. Daniel Butnaru, Benjamin Peherstorfer, Hans-Joachim Bungartz, Dirk Pflüger |
ICMLA (2) | 4 |
| 2012 | A Non-static Data Layout Enhancing Parallelism and Vectorization in Sparse Grid AlgorithmsabstractThe name sparse grids denotes a highly space-efficient, grid-based numerical technique to approximate high-dimensional functions. Although employed in a broad spectrum of applications from different fields, there have only been few tries to use it in real time visualization (e.g. [1]), due to complex data structures and long algorithm runtime. In this work we present a novel approach inspired by principles of I/0-efficient algorithms. Locally applied coefficient permutations lead to improved cache performance and facilitate the use of vector registers for our sparse grid benchmark problem hierarchization. Based on the compact data structure proposed for regular sparse grids in [2], we developed a new algorithm that outperforms existing implementations on modern multi-core systems by a factor of 37 for a grid size of 127 million points. For larger problems the speedup is even increasing, and with execution times below 1 s, sparse grids are well-suited for visualization applications. Furthermore, we point out how a broad class of sparse grid algorithms can benefit from our approach. Gerrit Buse, Dirk Pflüger, Alin Florindor Murarasu, Riko Jacob |
ISPDC | 2 |
| 2012 | A Parallel and Distributed Surrogate Model Implementation for Computational SteeringabstractUnderstanding the influence of multiple parameters in a complex simulation setting is a difficult task. In the ideal case, the scientist can freely steer such a simulation and is immediately presented with the results for a certain configuration of the input parameters. Such an exploration process is however not possible if the simulation is computationally too expensive. For these cases we present in this paper a scalable computational steering approach utilizing a fast surrogate model as substitute for the time-consuming simulation. The surrogate model we propose is based on the sparse grid technique, and we identify the main computational tasks associated with its evaluation and its extension. We further show how distributed data management combined with the specific use of accelerators allows us to approximate and deliver simulation results to a high-resolution visualization system in real-time. This significantly enhances the steering workflow and facilitates the interactive exploration of large datasets. Daniel Butnaru, Gerrit Buse, Dirk Pflüger |
ISPDC | 3 |
| 2011 | Compact data structure and scalable algorithms for the sparse grid techniqueabstractThe sparse grid discretization technique enables a compressed representation of higher-dimensional functions. In its original form, it relies heavily on recursion and complex data structures, thus being far from well-suited for GPUs. In this paper, we describe optimizations that enable us to implement compression and decompression, the crucial sparse grid algorithms for our application, on Nvidia GPUs. The main idea consists of a bijective mapping between the set of points in a multi-dimensional sparse grid and a set of consecutive natural numbers. The resulting data structure consumes a minimum amount of memory. For a 10-dimensional sparse grid with approximately 127 million points, it consumes up to 30 times less memory than trees or hash tables which are typically used. Compared to a sequential CPU implementation, the speedups achieved on GPU are up to 17 for compression and up to 70 for decompression, respectively. We show that the optimizations are also applicable to multicore CPUs. Alin Florindor Murarasu, Josef Weidendorfer, Gerrit Buse, Daniel Butnaru, Dirk Pflüger |
PPoPP | 5 |
| 2010 | Spatially adaptive sparse grids for high-dimensional data-driven problems
Dirk Pflüger, Benjamin Peherstorfer, Hans-Joachim Bungartz |
J. Complex. | 1 |