Ben van Werkhoven

dblp:49/10949 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-7508-3272ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Velvet: Parallel Divide-and-Conquer in Safe Rust
Anna Badia Liokouras, Ben van Werkhoven, Rob van Nieuwpoort
Euro-Par (1)2
2026 Kernel Float: Unlocking Mixed-Precision GPU Programming
abstract
Modern GPUs feature specialized hardware for low-precision floating-point arithmetic to accelerate compute-intensive workloads that do not require high numerical accuracy, such as those from artificial intelligence. However, despite the significant gains in computational throughput, memory bandwidth utilization, and energy efficiency, integrating low-precision formats into scientific applications remains difficult. We introduce Kernel Float , a header-only C++ library that simplifies the development of portable mixed-precision GPU kernels. Kernel Float provides a generic vector type, a unified interface for common mathematical operations, and fast approximations for low-precision transcendental functions that lack native hardware support. To demonstrate the potential of mixed-precision computing unlocked by our library, we integrated Kernel Float into nine GPU kernels from various domains. Our evaluation on Nvidia A100 and AMD MI250X GPUs shows performance improvements of up to \(12\times\) over double precision, while reducing source code length by up to 50% compared to handwritten kernels and having negligible runtime overhead. Our results further show that mixed-precision performance depends not only on choosing appropriate data types but also on tuning traditional optimization parameters (e.g., block size and vector width) and, when relevant, even domain-specific parameters.
Stijn Heldens, Ben van Werkhoven
ACM Trans. Math. Softw.2
2026 Accuracy-Aware Mixed-Precision GPU Auto-Tuning
abstract
Reduced-precision floating-point arithmetic has become increasingly important in GPU applications for AI and HPC, as it can deliver substantial speedups while reducing energy consumption and memory footprint. However, choosing the appropriate data formats brings a challenging tuning problem: precision parameters must be chosen to maximize performance while preserving numerical accuracy. At the same time, GPU kernels typically expose additional tunable optimization parameters, such as block size, tiling strategy, and vector width. The combination of these two kinds of parameters results in a complex trade-off between accuracy and performance, making manual exploration of the resulting design space time-consuming. In this work, we present anaccuracy-awareextension to the open-sourceKernel Tunerframework, enabling automatic tuning of floating-point precision parameters alongside conventional code-optimization parameters. We evaluate our accuracy-aware tuning solution on both Nvidia and AMD GPUs using a variety of kernels. Our results show speedups of up to$12{\times }$over double precision, demonstrate how Kernel Tuner's built-in search strategies are effective for accuracy-aware tuning, and show that our approach can be extended to other optimization objectives, such as memory footprint or energy efficiency. Moreover, we highlight that jointly tuning accuracy- and performance-affecting parameters outperforms isolated approaches in finding the best-performing configurations, despite significantly expanding the optimization space. This unified approach enables developers to trade accuracy for throughput systematically, enabling broader adoption of mixed-precision computing in scientific and industrial applications.
Stijn Heldens, Ben van Werkhoven
IEEE Trans. Parallel Distributed Syst.2
2025 Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
abstract
Automatic performance tuning (auto-tuning) is widely used to optimize performance-critical applications across many scientific domains by finding the best program variant among many choices. Efficient optimization algorithms are crucial for navigating the vast and complex search spaces in autotuning. As is well known in the context of machine learning and similar fields, hyperparameters critically shape optimization algorithm efficiency. Yet for auto-tuning frameworks, these hyperparameters are almost never tuned, and their potential performance impact has not been studied.We present a novel method for general hyperparameter tuning of optimization algorithms for auto-tuning, thus "tuning the tuner". In particular, we propose a robust statistical method for evaluating hyperparameter performance across search spaces, publish a FAIR data set and software for reproducibility, and present a simulation mode that replays previously recorded tuning data, lowering the costs of hyperparameter tuning by two orders of magnitude. We show that even limited hyperparameter tuning can improve auto-tuner performance by 94.8% on average, and establish that the hyperparameters themselves can be optimized efficiently with meta-strategies (with an average improvement of 204.7%), demonstrating the often overlooked hyperparameter tuning as a powerful technique for advancing auto-tuning research and practice.
Floris-Jan Willemsen, Rob van Nieuwpoort, Ben van Werkhoven
eScience3
2025 Efficient Construction of Large Search Spaces for Auto-Tuning
Floris-Jan Willemsen, Rob van Nieuwpoort, Ben van Werkhoven
ICPP3
2025 PowerSensor3: A Fast and Accurate Open Source Power Measurement Tool
abstract
Power consumption is a major concern in data centers and HPC applications, with GPUs typically accounting for more than half of system power usage. While accurate power measurement tools are crucial for optimizing the energy efficiency of (GPU) applications, both built-in power sensors as well as state-of-the-art power meters often lack the accuracy and temporal granularity needed, or are impractical to use. Released as open hardware, firmware, and software, PowerSensor3 provides a cost-effective solution for evaluating energy efficiency, enabling advancements in sustainable computing. The toolkit consists of a baseboard with a variety of sensor modules accompanied by host libraries with C++ and Python bindings. PowerSensor3 enables real-time power measurements of SoC boards and PCIe cards, including GPUs, FPGAs, NICs, SSDs, and domain-specific AI and ML accelerators. Additionally, it provides significant improvements over previous tools, such as a robust and modular design, current sensors resistant to external interference, simplified calibration, and a sampling rate up to 20 kHz, which is essential to identify GPU behavior at high temporal granularity. This work describes the toolkit design, evaluates its performance characteristics, and shows several use cases (GPUs, NVIDIA Jetson AGX Orin, and SSD), demonstrating PowerSensor3's potential to significantly enhance energy efficiency in modern computing environments.
Steven van der Vlugt, Leon C. Oostrum, Gijs Schoonderbeek, Ben van Werkhoven, Bram Veenboer, Krijn Doekemeijer, John W. Romein
ISPASS4
2024 Bringing Auto-Tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
Milo Lurati, Stijn Heldens, Alessio Sclocco, Ben van Werkhoven
Euro-Par (1)4
2024 A methodology for comparing optimization algorithms for auto-tuning
abstract
Adapting applications to optimally utilize available hardware is no mean feat: the plethora of choices for optimization techniques are infeasible to maximize manually. To this end, auto-tuning frameworks are used to automate this task, which in turn use optimization algorithms to efficiently search the vast searchspaces. However, there is a lack of comparability in studies presenting advances in auto-tuning frameworks and the optimization algorithms incorporated. As each publication varies in the way experiments are conducted, metrics used, and results reported, comparing the performance of optimization algorithms among publications is infeasible. The auto-tuning community identified this as a key challenge at the 2022 Lorentz Center workshop on auto-tuning. The examination of the current state of the practice in this paper further underlines this. We propose a community-driven methodology composed of four steps regarding experimental setup, tuning budget, dealing with stochasticity, and quantifying performance. This methodology builds upon similar methodologies in other fields while taking into account the constraints and specific characteristics of the auto-tuning field, resulting in novel techniques. The methodology is demonstrated in a simple case study that compares the performance of several optimization algorithms used to auto-tune CUDA kernels on a set of modern GPUs. We provide a software tool to make the application of the methodology easy for authors, and simplifies reproducibility of results.
Floris-Jan Willemsen, Richard Schoonhoven, Jiri Filipovic, Jacob Odgård Tørring, Rob van Nieuwpoort, Ben van Werkhoven
Future Gener. Comput. Syst.6
2023 Benchmarking Optimization Algorithms for Auto-Tuning GPU Kernels
abstract
Recent years have witnessed phenomenal growth in the application, and capabilities of graphical processing units (GPUs) due to their high parallel computation power at relatively low cost. However, writing a computationally efficient GPU program (kernel) is challenging and, generally, only certain specific kernel configurations lead to significant increases in performance. Auto-tuning is the process of automatically optimizing software for highly efficient execution on a target hardware platform. Auto-tuning is particularly useful for GPU programming, as a single kernel requires retuning after code changes, for different input data, and for different architectures. However, the discrete and nonconvex nature of the search space creates a challenging optimization problem. In this work, we investigate which algorithm produces the fastest kernels if the time-budget for the tuning task is varied. We conduct a survey by performing experiments on 26 different kernel spaces, from nine different GPUs, for 16 different evolutionary black-box optimization algorithms. We then analyze these results and introduce a novel metric based on the PageRank centrality concept as a tool for gaining insight into the difficulty of the optimization problem. We demonstrate that our metric correlates strongly with the observed tuning performance.
Richard Schoonhoven, Ben van Werkhoven, Kees Joost Batenburg
IEEE Trans. Evol. Comput.2
2022 Lightning: Scaling the GPU Programming Model Beyond a Single GPU
abstract
The GPU programming model is primarily aimed at the development of applications that run one GPU. However, this limits the scalability of GPU code to the capabilities of a single GPU in terms of compute power and memory capacity. To scale GPU applications further, a great engineering effort is typically required: work and data must be divided over multiple GPUs by hand, possibly in multiple nodes, and data must be manually spilled from GPU memory to higher-level memories. We present Lightning: a framework that follows the common GPU programming paradigm but enables scaling to large problems with ease. Lightning supports multi-GPU execution of GPU kernels, even across multiple nodes, and seamlessly spills data to higher-level memories (main memory and disk). Existing CUDA kernels can easily be adapted for use in Lightning, with data access annotations on these kernels allowing Lightning to infer their data requirements and the dependencies between subsequent kernel launches. Lightning efficiently distributes the work/data across GPUs and maximizes efficiency by overlapping scheduling, data movement, and kernel execution when possible. We present the design and implementation of Lightning, as well as experimental results on up to 32 GPUs for eight benchmarks and one real-world application. Evaluation shows excellent performance and scalability, such as a speedup of 57.2 x over the CPU using Lighting with 16 GPUs over 4 nodes and 80 GB of data, far beyond the memory capacity of one GPU.
Stijn Heldens, Pieter Hijma, Ben van Werkhoven, Jason Maassen, Rob van Nieuwpoort
IPDPS3
2021 GPU Optimizations for Atmospheric Chemical Kinetics
abstract
We present a series of optimizations to alleviate stack memory overflow issues and improve overall performance of GPU computational kernels in atmospheric chemical kinetics model simulations. We use heap memory in numerical solvers for stiff ODEs, move chemical reaction constants and tracer concentration arrays from stack to global memory, use direct pointer indexing for array memory access, and use CUDA streams to overlap computation with memory transfer to the device. Overall, an order of magnitude reduction in GPU memory requirements is achieved, allowing for simultaneous offloading from multiple MPI processes per node and/or increasing the chemical mechanism complexity.
Theodoros Christoudias, Timo Kirfel, Astrid Kerkweg, Domenico Taraborrelli, Georges-Emmanuel Moulard, Erwan Raffin, Victor Azizi, Gijs van den Oord, Ben van Werkhoven
HPC Asia9
2020 Rocket: efficient and scalable all-pairs computations on heterogeneous platforms
abstract
All-pairs compute problems apply a user-defined function to each combination of two items of a given data set. Although these problems present an abundance of parallelism, data reuse must be exploited to achieve good performance. Several researchers considered this problem, either resorting to partial replication with static work distribution or dynamic scheduling with full replication. In contrast, we present a solution that relies on hierarchical multi-level software-based caches to maximize data reuse at each level in the distributed memory hierarchy, combined with a divide-and-conquer approach to exploit data locality, hierarchical work-stealing to dynamically balance the workload, and asynchronous processing to maximize resource utilization. We evaluate our solution using three real-world applications, from digital forensics, localization microscopy, and bioinformatics, on different platforms, from desktop machine to a supercomputer. Results shows excellent efficiency and scalability when scaling to 96 GPUs, even obtaining super-linear speedups due to a distributed cache.
Stijn Heldens, Pieter Hijma, Ben van Werkhoven, Jason Maassen, Henri E. Bal, Rob van Nieuwpoort
SC3
2019 Kernel Tuner: A search-optimizing GPU code auto-tuner
abstract
A very common problem in GPU programming is that some combination of thread block dimensions and other code optimization parameters, like tiling or unrolling factors, results in dramatically better performance than other kernel configurations. To obtain highly-efficient kernels it is often required to search vast and discontinuous search spaces that consist of all possible combinations of values for all tunable parameters. This paper presents Kernel Tuner, an easy-to-use tool for testing and auto-tuning OpenCL, CUDA, and C kernels with support for many search optimization algorithms that accelerate the tuning process. This paper introduces the application of many new solvers and global optimization algorithms for auto-tuning GPU applications. We demonstrate that Kernel Tuner can be used in a wide range of application scenarios and drastically decreases the time spent tuning, e.g. tuning a GEMM kernel on AMD Vega Frontier Edition 71.2x faster than brute force search. • Introduces and evaluates optimization algorithms for auto-tuning GPU applications. • Presents Kernel Tuner, an easy-to-use tool for testing and auto-tuning GPU kernels. • Demonstrates effectiveness of Basin Hopping to speedup the auto-tuning of GPU kernels.
Ben van Werkhoven
Future Gener. Comput. Syst.1
2018 Development of the OMUSE/AMUSE Modeling System
abstract
The Oceanographic Multipurpose Software Environment (OMUSE, [1]) is an open source framework developed for oceanographic and other earth system modelling applications. OMUSE provides a homogeneous environment to interface with numerical simulation codes. It was developed at the IMAU (Utrecht) using coupling technology developed for astrophysical applications in the AMUSE project at Leiden Observatory[2,3]. OMUSE simplifies the use and deployment of numerical simulations codes. Furthermore, the design of the OMUSE interfaces (figure 1) allow codes that represent different physics or span different ranges of physical scales to be easily combined in novel numerical experiments. The use cases for OMUSE range from running simple numerical experiments with single codes and the addition of data analysis tools in model runs, to setting up fairly complicated and strongly coupled solvers for problems that are intrinsically multi-scale and/or require different physics. Here, we will present the design of OMUSE as well as give examples of the types of the couplings that can be implemented using OMUSE. The example provided by AMUSE and OMUSE suggests that application of the same interfacing philosophy to a more extensive set of disciplines is possible. In order to facilitate this a better separation of the core framework and domain specific code is necessary. We will present ongoing work to support meteorological and hydrological applications and the use of the framework as the computational core in the eWatercycle project [4]. For this, adaptations are made to improve the interoperability with existing interface efforts (such as the BMI) and we discuss developments regarding the encapsulation of OMUSE/AMUSE and its component models in containers. This will facilitate the installation for first time users, removing a barrier in this respect. In addition to this we anticipate this to also offer more flexible deployment options for the framework.
F. Inti Pelupessy, Ben van Werkhoven, Gijs van den Oord, Simon Portegies Zwart, Arjen van Elteren, Henk A. Dijkstra
eScience2
2018 Painting the Picture of Software Impact with the Research Software Directory
abstract
In this lightning talk we will describe the Research Software Directory; a content management system that is tailored to research software with the goal of enabling a qualitative assessment of software impact and improving software findability.
Jurriaan H. Spaaks, Tom Klaver, Stefan Verhoeven, Jason Maassen, Tom Bakker, Atze van der Ploeg, Ben van Werkhoven, Willem Robert van Hage, Rob van Nieuwpoort
eScience7
2018 Survey on Research Software Engineering in the Netherlands
abstract
This paper presents a brief overview of the Research Software Engineering landscape in the Netherlands and includes a summary of the results from a survey held in December 2017 in the Netherlands and several other countries. The results show that best practices are widely adopted. Research software is produced by small teams or individuals, is often used for scientific publications, and is frequently acknowledged in publications.
Ben van Werkhoven, Tom Bakker, Olivier Philippe, Simon Hettrick
eScience1
2018 Poster Abstracts eScience 2018 Conference
abstract
The eScience conference aims to bring together leading international researchers and research software engineers from all disciplines to present and discuss how digital technology impacts scientific research. There were many poster abstracts submitted to the conference this year, and we have also invited several authors of full papers submitted to the conference to submit their work for a poster presentation. We are very happy with the selection of posters that have been confirmed for poster presentation at the conference. The diverse set of topics covered by the posters reflects the broad impact of eScience in various domains, as well as the high quality technical work that is performed by the various teams of researchers. Posters are a great way of presenting work at a conference that really encourage discussions and interactions among the participants. We look forward to the poster sessions at this year’s eScience conference.
Ben van Werkhoven, Adriënne Mendrik, Rob van Nieuwpoort
eScience1
2017 On the complexities of utilizing large-scale lightpath-connected distributed cyberinfrastructure
abstract
Summary In Autumn 2013, we—an international team of climate scientists, computer scientists, eScience researchers, and e‐Infrastructure specialists—participated in the enlighten your research global competition, organized to showcase advanced lightpath technologies in support of state‐of‐the‐art research questions. As one of the winning entries, our enlighten your research global team embarked on a very ambitious project to run an extremely high resolution climate model on a collection of supercomputers distributed over two continents and connected using an advanced 10 G lightpath networking infrastructure. Although good progress was made, we were not able to perform all desired experiments due to a varying combination of technical problems, configuration issues, policy limitations and lack of (budget for) human resources to solve these issues. In this paper, we describe our goals, the technical and non‐technical barriers, we encountered and provide recommendations on how these barriers can be removed so future project of this kind may succeed. Copyright © 2016 John Wiley & Sons, Ltd.
Jason Maassen, Ben van Werkhoven, Maarten A. J. van Meersbergen, Henri E. Bal, Michael Kliphuis, Sandra E. Brunnabend, Henk A. Dijkstra, Gerben van Malenstein, Migiel de Vos, Sylvia Kuijpers, Sander Boele, Jules Wolfrat, Nick Hill, David Wallom, Christian Grimm, Dieter Kranzlmüller, Dinesh Ganpathi, Shantenu Jha, Yaakoub El Khamra, Frank O. Bryan, Benjamin Kirtman, Frank J. Seinstra
Concurr. Comput. Pract. Exp.2
2016 OMUSE: Oceanographic multipurpose software environment
abstract
This talk will give a brief introduction to OMUSE, the Oceanographic Multipurpose Software Environment, which is currently being developed. OMUSE is a Python framework that provides high-level object-oriented interfaces to existing or newly developed numerical ocean simulation codes, simplifying their use and development In this way, OMUSE facilitates the efficient design of numerical experiments that combine ocean models representing different physics or spanning different ranges of physical scales, for example coupling a global open ocean simulation with a regional coastal ocean model. OMUSE enables its users to write high-level Python scripts that describe simulations. The functionality provided by OMUSE takes care of the low-level integration with the code and deploying simulations on high-performance computing resources, allowing its users to focus on the physics of the simulation. We give an overview of the design of OMUSE and the modules and model components currently included. In particular, we will discuss the process of creating a new OMUSE interface to an existing code, and explain how OMUSE keeps track of the internal state of a running simulation. In addition, we will discuss the grid data types and grid remapping functionality that OMUSE provides. We also give an example of performing online data analysis on a running simulation, which is becoming increasingly important as models simulate a broader range of scales, generating large datasets that cannot be fully stored for offline analysis.
F. Inti Pelupessy, Ben van Werkhoven, Arjen van Elteren, Jan Viebahn, Adam Candy, Simon Portegies Zwart, Henk A. Dijkstra
eScience2
2016 A spatial column-store to triangulate the Netherlands on the fly
abstract
3D digital city models, important for urban planning, are currently constructed from massive point clouds obtained through airborne LiDAR (Light Detection and Ranging). They are semantically enriched with information obtained from auxiliary GIS data like Cadastral data which contains information about the boundaries of properties, road networks, rivers, lakes etc.
Romulo Goncalves, Tom van Tilburg, Kostis Kyzirakos, Foteini Alvanaki, Panagiotis Koutsourakis, Ben van Werkhoven, Willem Robert van Hage
SIGSPATIAL/GIS6
2015 An Integrated Approach to Porting Large Scientific Applications to GPUs
abstract
There are many large scientific applications that have been actively developed for several decades. However, in this time the hardware has evolved considerably. It is taking large scientific applications a very long time to get adjusted to the new computing infrastructure. This is because porting these applications to new hardware, such as Graphics Processing Units (GPUs), currently requires a huge amount of manual labor, even though the computations are very well suited for GPUs. In this paper we propose an integrated approach to semi-automatically port large long-lived scientific codes to GPUs. We propose a method that considerably reduces the effort required by experienced GPU programmers to port these applications. This approach is supported by a tool that is able to analyze, transform, and translate source code into different programming languages. We evaluate our approach by applying it to the Parallel Ocean Program, a representative, very large, and widely-used scientific application.
Ben van Werkhoven, Pieter Hijma
e-Science1
2014 Performance Models for CPU-GPU Data Transfers
abstract
Many GPU applications perform data transfers to and from GPU memory at regular intervals. For example because the data does not fit into GPU memory or because of internode communication at the end of each time step. Overlapping GPU computation with CPU-GPU communication can reduce the costs of moving data. Several different techniques exist for transferring data to and from GPU memory and for overlapping those transfers with GPU computation. It is currently not known when to apply which method. Implementing and benchmarking each method is often a large programming effort and not feasible. To solve these issues and to provide insight in the performance of GPU applications, we propose an analytical performance model that includes PCIe transfers and overlapping computation and communication. Our evaluation shows that the performance models are capable of correctly classifying the relative performance of the different implementations.
Ben van Werkhoven, Jason Maassen, Frank J. Seinstra, Henri E. Bal
CCGRID1
2014 Optimizing convolution operations on GPUs using adaptive tiling
Ben van Werkhoven, Jason Maassen, Henri E. Bal, Frank J. Seinstra
Future Gener. Comput. Syst.1
2013 User transparent data and task parallel multimedia computing with Pyxis-DT
Timo van Kessel, Ben van Werkhoven, Niels Drost, Jason Maassen, Henri E. Bal, Frank J. Seinstra
Future Gener. Comput. Syst.2