VLDB 2026 Research / reviewers in the wild / expert
Peter Thoman
dblp:44/7010
· DBLP profile ↗
27ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0002-4028-7451ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 8 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Parallelization Capabilities of State-of-the-Art Agentic Large Language Models
Peter Thoman, Philipp Gschwandtner |
Euro-Par (1) | 1 |
| 2026 | Bridging usability and performance: High-level abstractions for advanced accelerator cluster programming
Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Thomas Fahringer |
Future Gener. Comput. Syst. | 3 |
| 2026 | A Portable Compiler-Runtime Approach for Scalability PredictionabstractHighly scalable parallel applications can efficiently solve expensive computational problems when run on a large number of compute nodes. However, selecting the optimal number of nodes for a compute job of a given size is non-trivial, and allocating too few or too many nodes may not yield the expected performance. Knowing the scaling behavior of an application in advance enables us, for example, to make optimal use of the available hardware resources. We introduce a novel, portable approach to predict the scalability of parallel applications written in modern high-level programming models. We propose a predictive compiler-runtime framework based on Celerity, a task-based distributed runtime system that enables executing SYCL codes on clusters. The framework targets a broad range of computing systems, from CPU to GPU clusters, and proposes a model that combines machine learning, communication modeling and DAG heuristics. Experimental results on two large-scale clusters, JUWELS and Marconi-100, show accurate scalability prediction of unseen single and multi-task applications. Nicolai Stawinoga, Sohan Lal, Biagio Cosenza, Philip Salzmann, Peter Thoman, Thomas Fahringer |
Future Gener. Comput. Syst. | 5 |
| 2023 | An Asynchronous Dataflow-Driven Execution Model For Distributed Accelerator ComputingabstractWhile domain-specific HPC software packages continue to thrive and are vital to many scientific communities, a general purpose high-productivity GPU cluster programming model that facilitates experimentation for non-experts remains elusive. We demonstrate how Celerity, a high-level C++ programming model for distributed accelerator computing based on the open SYCL standard, allows for the quick development of - and experimentation with - distributed applications. To achieve scalability on large machines, we replace Celerity's existing master/worker scheduling model with a fully distributed scheme that reduces the worst-case scheduling complexity from quadratic to linear while maintaining the existing programming interface. We then show how this declarative, data-flow based API paired with a point-to-point communication model with eager data pushing can effectively expose and leverage opportunities for latency hiding and computation/communication overlapping with minimal or no manual guidance. We demonstrate how Celerity exhibits very good scalability on multiple benchmarks from several scientific domains and up to 128 GPUs. Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Biagio Cosenza, Thomas Fahringer |
CCGrid | 3 |
| 2023 | Tunable and Portable Extreme-Scale Drug Discovery Platform at Exascale: the LIGATE ApproachabstractToday digital revolution is having a dramatic impact on the pharmaceutical industry and the entire healthcare system. The implementation of machine learning, extreme-scale computer simulations, and big data analytics in the drug design and development process offers an excellent opportunity to lower the risk of investment and reduce the time to the patient. Gianluca Palermo, Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Cristina Silvano, Bruno Guindani, Danilo Ardagna, Andrea Beccari, Domenico Bonanni, Carmine Talarico, Filippo Lunghini, Jan Martinovic, Paulo Silva 0002, Ada Böhm, Jakub Beránek, Jan Krenek, Branislav Jansik, Biagio Cosenza, Luigi Crisci, Peter Thoman, Philip Salzmann, Thomas Fahringer, Leila Tamara Alexander, Gerardo Tauriello, Torsten Schwede, Janani Durairaj, Andrew Emerson, Federico Ficarelli, Sebastian Wingbermühle, Erik Lindahl, Daniele Gregori, Emanuele Sana, Silvano Coletti, Philipp Gschwandtner |
CF | 20 |
| 2022 | Multi-GPU room response simulation with hardware raytracingabstractSummary Time‐of‐flight camera systems are an essential component in 3D scene analysis and reconstruction for many modern computer vision applications. The development and validation of such systems require testing in a large variety of scenes and situations. Accurate room impulse response simulation greatly speeds up development and validation, as well as reducing its cost, but large computational overhead has so far limited its applicability. While the overall algorithmic requirements of this simulation differ significantly from 3D rendering, the recently introduced hardware raytracing support in GPUs nonetheless provides an interesting new implementation option. In this article, we present a new room response simulation method, implemented in a vendor‐independent fashion with Vulkan compute shaders and leveraging NVIDIA VKRay hardware raytracing. We also extend this method to multi‐GPU computation with asynchronous streaming and introduce a domain‐specific high‐performance compression scheme in order to overcome the limitations of on‐board GPU memory and PCIe bandwidth when simulating very large scenes. Our implementation is, to the best of our knowledge, the first ever combined application of Vulkan hardware raytracing and multi‐GPU compute in a non‐rendering simulation setting. Compared to a state‐of‐the‐art multicore CPU implementation running on 12 CPU cores, we achieve an overall speedup factor of up to 20 on a single consumer GPU, and 71 on four GPUs. Peter Thoman, Markus Wippler, Robert Hranitzky, Philipp Gschwandtner, Thomas Fahringer |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | ndzip: A High-Throughput Parallel Lossless Compressor for Scientific DataabstractPublikationen von Forschenden. Knorr, Fabian; Thoman, Peter; Fahringer, Thomas: ndzip: a high-throughput parallel lossless compressor for scientific data. In: Proceedings 2021 Data Compression Conference (DCC) / Bilgin, Ali; Marcellin, Michael W.; Serra-Sagrista, Joan; Storer, James A. IEEE, 2021 Fabian Knorr, Peter Thoman, Thomas Fahringer |
DCC | 2 |
| 2021 | Porting Real-World Applications to GPU Clusters: A Celerity and Cronos Case StudyabstractAccelerator clusters are an ongoing trend in high performance computing, continuously gaining traction and forming a ubiquitous hardware resource for domain scientists to run large-scale simulations on. However, there is often a gap between new hardware technologies and adoption by legacy code bases. Porting real-world applications to new programming models is a difficult undertaking, aggravated by the need for support for both distributed-memory and accelerator parallelism. In this work, we present a case study of porting Cronos, a real-world code from the field of magnetohydrodynamics, to Celerity, a high-level programming model for distributed-memory accelerator clusters. We discuss the numerical, algorithmic and implementation properties of the application and motivate our decisions for adapting them where necessary. Preliminary results show a parallel efficiency of up to 87% for 16 GPUs. Philipp Gschwandtner, Ralf Kissmann, David Huber 0003, Philip Salzmann, Fabian Knorr, Peter Thoman, Thomas Fahringer |
e-Science | 6 |
| 2021 | Optimizing Embedded Industrial Safety Systems Based on Time-of-flight Depth ImagingabstractTime-of-flight camera systems provide rapid, low-latency depth imaging. In the context of industrial safety, they can be employed to replace costly and inflexible physical safety installations with virtual safety volumes, monitored by time-of-flight cameras which turn off dangerous systems when unknown objects enter specified bounding volumes.In this context, the 3D point-shape collision detection algorithm has to be highly optimized for very specific goals. These include high real-time (worst case) performance, low power consumption and acceptable thermal characteristics, ideally with low-cost and well-established embedded hardware.We present our work in the INPACT project, which explores various algorithmic and parallelization-based methods of improving the performance, responsiveness and reliability of such systems. Our experiments are are based on a setup allowing long-term measurement of multiple thermal sampling points as well as fine-grained power consumption information and of course performance and throughput metrics. Peter Thoman, Alexander Hirsch, Markus Wippler, Robert Hranitzky |
e-Science | 1 |
| 2021 | ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUsabstractLossless data compression is a promising software approach for reducing the bandwidth requirements of scientific applications on accelerator clusters without introducing approximation errors. Suitable compressors must be able to effectively compact floating-point data while saturating the system interconnect to avoid introducing unnecessary latencies. Fabian Knorr, Peter Thoman, Thomas Fahringer |
SC | 2 |
| 2021 | The cluster coffer: Teaching HPC on the roadabstractTeaching parallel programming and HPC is a difficult task. There is a large number of sophisticated hardware and software components, each complex on their own and often showing non-intuitive interaction when used in combination. We consider education in HPC among the more difficult topics in computer science due to the fact that larger distributed memory systems are ubiquitous yet inaccessible and intangible to students. In this work, we present the Cluster Coffer, a miniature cluster computer based on 16 ARM compute boards that we believe is suitable for reducing the entry barrier to HPC in teaching and public outreach. We discuss our design goals for providing a portable, inexpensive system that is easy to maintain and repair. We outline the implementation path we took in terms of hardware and software, in order to provide others with the information required to reproduce and extend our work. Finally, we present two use cases for which the Cluster Coffer has been used multiple times, and will continue to be used in the upcoming years. Philipp Gschwandtner, Alexander Hirsch, Peter Thoman, Peter Zangerl, Herbert Jordan, Thomas Fahringer |
J. Parallel Distributed Comput. | 3 |
| 2020 | SYCL-Bench: A Versatile Cross-Platform Benchmark Suite for Heterogeneous Computing
Sohan Lal, Aksel Alpay, Philip Salzmann, Biagio Cosenza, Alexander Hirsch, Nicolai Stawinoga, Peter Thoman, Thomas Fahringer, Vincent Heuveline |
Euro-Par | 7 |
| 2020 | The allscale framework architecture
Herbert Jordan, Philipp Gschwandtner, Peter Thoman, Peter Zangerl, Alexander Hirsch, Thomas Fahringer, Thomas Heller, Dietmar Fey |
Parallel Comput. | 3 |
| 2019 | The AllScale APIabstractEffectively implementing scientific algorithms in distributed memory parallel applications is a difficult task for domain scientists, as evident by the large number of domain-specific languages and libraries available today attempting to facilitate the process. However, they usually provide a closed set of parallel patterns and are not open for extension without vast modifications to the underlying system. In this work, we present the AllScale API, a programming interface for developing distributed memory parallel applications with the ease of shared memory programming models. The AllScale API is closed for modification but open for extension, allowing new, user-defined parallel patterns and data structures to be implemented based on existing core primitives and therefore fully supported in the AllScale framework. Focusing on high-level functionality directly offered to application developers, we present the design advantages of such an API design, detail some of its specifications and evaluate it using three real-world use cases. Our results show that AllScale decreases the complexity of implementing scientific applications for distributed memory while attaining comparable or higher performance compared to MPI reference implementations. Philipp Gschwandtner, Herbert Jordan, Peter Thoman, Thomas Fahringer |
eScience | 3 |
| 2019 | Celerity: High-Level C++ for Accelerator Clusters
Peter Thoman, Philip Salzmann, Biagio Cosenza, Thomas Fahringer |
Euro-Par | 1 |
| 2018 | The AllScale Runtime Application ModelabstractContemporary state-of-the-art runtime systems underlying widely utilized general purpose parallel programming languages and libraries like OpenMP, MPI, or OpenCL provide the foundation for accessing the parallel capabilities of modern computing architectures. In the tradition of their respective imperative host languages those runtime systems' main focus is on providing means for the distribution and synchronization of operations - while the organization and management of manipulated data is left to application developers. Consequently, the distribution of data remains inaccessible to those runtime systems. However, many desirable system-level features depend on a runtime system's ability to exercise control on the distribution of data. Thus, program models underlying traditional systems lack the potential for the support of those features. In this paper, we present a novel application model granting parallel runtime systems system-wide control over the distribution of user-defined shared data structures. Our model utilizes the high-level nature of parallel programming languages, in particular, the usage of well-typed data structures and the associated hiding of implementation details from the application developers. By being based on a generalization of such data structures and extending the resulting abstraction with features facilitating the automated management of the distribution of those, our model enables runtime systems to dynamically influence the placement and replication of shared data. This paper covers a rigorous formal description of our application model, as well as details on our prototype implementation and experimental results demonstrating its ability to efficiently and scalably manage various data structures in real-world environments. Herbert Jordan, Thomas Heller, Philipp Gschwandtner, Peter Zangerl, Peter Thoman, Dietmar Fey, Thomas Fahringer |
CLUSTER | 5 |
| 2018 | A taxonomy of task-based parallel programming technologies for high-performance computingabstractTask-based programming models for shared memory—such as Cilk Plus and OpenMP 3—are well established and documented. However, with the increase in parallel, many-core, and heterogeneous systems, a number of research-driven projects have developed more diversified task-based support, employing various programming and runtime features. Unfortunately, despite the fact that dozens of different task-based systems exist today and are actively used for parallel and high-performance computing (HPC), no comprehensive overview or classification of task-based technologies for HPC exists. In this paper, we provide an initial task-focused taxonomy for HPC technologies, which covers both programming interfaces and runtime mechanisms. We demonstrate the usefulness of our taxonomy by classifying state-of-the-art task-based environments in use today. Peter Thoman, Kiril Dichev, Thomas Heller, Roman Iakymchuk, Xavier Aguilar, Khalid Hasanov, Philipp Gschwandtner, Pierre Lemarinier, Stefano Markidis, Herbert Jordan, Thomas Fahringer, Kostas Katrinis, Erwin Laure, Dimitrios S. Nikolopoulos |
J. Supercomput. | 1 |
| 2017 | Characterizing Performance and Cache Impacts of Code Multi-versioning on Multicore ArchitecturesabstractCode multi-versioning is an increasingly widely adopted tool for implementing optimizations which respond to unknown or dynamically changing runtime conditions, without the performance overhead of just-in-time compilation. A common concern in its use is instruction cache performance, due to larger binary sizes increasing cache pressure on the one hand and more unpredictable branching on the other. Despite this ongoing interest, there has been no comprehensive study of the impact of multi-versioning so far - particularly in a multi-threaded setting. In this paper, we present a categorization of the parameter space potentially affecting multi-versioned performance, a toolset for exploring this space, and an in-depth characterization of three hardware platforms using this toolset. Peter Zangerl, Peter Thoman, Thomas Fahringer |
PDP | 2 |
| 2017 | SCALO: Scalability-Aware Parallelism Orchestration for Multi-Threaded WorkloadsabstractShared memory machines continue to increase in scale by adding more parallelism through additional cores and complex memory hierarchies. Often, executing multiple applications concurrently, dividing among them hardware threads, provides greater efficiency rather than executing a single application with large thread counts. However, contention for shared resources can limit the improvement of concurrent application execution: orchestrating the number of threads used by each application and is essential. In this article, we contribute SCALO, a solution to orchestrate concurrent application execution to increase throughput. SCALO monitors co-executing applications at runtime to evaluate their scalability. Its optimizing thread allocator analyzes these scalability estimates to adapt the parallelism of each program. Unlike previous approaches, SCALO differs by including dynamic contention effects on scalability and by controlling the parallelism during the execution of parallel regions. Thus, it improves throughput when other state-of-the-art approaches fail and outperforms them by up to 40% when they succeed. Giorgis Georgakoudis, Hans Vandierendonck, Peter Thoman, Bronis R. de Supinski, Thomas Fahringer, Dimitrios S. Nikolopoulos |
ACM Trans. Archit. Code Optim. | 3 |
| 2015 | Optimizing Task Parallelism with Library-Semantics-Aware Compilation
Peter Thoman, Stefan Moosbrugger, Thomas Fahringer |
Euro-Par | 1 |
| 2015 | On the Quality of Implementation of the C++11 Thread Support LibraryabstractProviding standardized building blocks for task-parallel programs within a language and its standard library has several advantages over other solutions. Close integration with compilers and runtime systems allows for potentially higher performance and portability facilitates wide-spread use. In the recently ratified C++11 standard, language constructs have been added along with a memory model to provide the developer with such building blocks. They allow accessing task parallelism and synchronization in a flexible and standardized way, potentially removing the need for third-party solutions. Nevertheless, since parallelization aims at high performance, an examination of the quality of implementation of these standardized means is necessary to determine their suitability for replacing established solutions. To that end, we present INNCABS, a new cross-platform cross-library benchmark suite consisting of 14 benchmarks with varying task granularities and synchronization requirements. Based on these benchmarks, we demonstrate that the performance of C++11 parallelism constructs in the three most commonly employed C++ runtime libraries prevents their use as a full replacement for third-party solutions due to simplistic parallelism implementations and high synchronization overheads. Peter Thoman, Philipp Gschwandtner, Thomas Fahringer |
PDP | 1 |
| 2014 | Compiler multiversioning for automatic task granularity controlabstractSUMMARY Task parallelism is a programming technique that has been shown to be applicable in a wide variety of problem domains. A central parameter that needs to be controlled to ensure efficient execution of task parallel programs is the granularity of tasks. When they are too coarse grained, scalability and load balance suffer, while very fine‐grained tasks introduce execution overheads. We present a combined compiler and runtime approach that enables automatic granularity control. Starting from recursive, task parallel programs, our compiler generates multiple versions of each task, increasing granularity by task unrolling. Subsequently, we apply a parallelism‐aware optimizing transformation to remove superfluous task synchronization primitives in all generated versions. A runtime system then selects among these task versions of varying granularity by locally tracking task demand. Benchmarking on a set of task parallel programs using a work‐stealing scheduler demonstrates that our approach is generally effective. For fine‐grained tasks, we can achieve reductions in execution time exceeding a factor of 6, compared with state‐of‐the‐art implementations. Additionally, we evaluate the impact of two crucial algorithmic parameters, the number of generated code versions and the task queue length, on the performance of our method. Copyright © 2014 John Wiley & Sons, Ltd. Peter Thoman, Herbert Jordan, Thomas Fahringer |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | INSPIRE: The insieme parallel intermediate representationabstractProgramming standards like OpenMP, OpenCL and MPI are frequently considered programming languages for developing parallel applications for their respective kind of architecture. Nevertheless, compilers treat them like ordinary APIs utilized by an otherwise sequential host language. Their parallel control flow remains hidden within opaque runtime library calls which are embedded within a sequential intermediate representation lacking the concepts of parallelism. Consequently, the tuning and coordination of parallelism is clearly beyond the scope of conventional optimizing compilers and hence left to the programmer or the runtime system. The main objective of the Insieme compiler is to overcome this limitation by utilizing INSPIRE, a unified, parallel, highlevel intermediate representation. Instead of mapping parallel constructs and APIs to external routines, their behavior is modeled explicitly using a unified and fixed set of parallel language constructs. Making the parallel control flow accessible to the compiler lays the foundation for the development of reusable, static and dynamic analyses and transformations bridging the gap between a variety of parallel paradigms. Within this paper we describe the structure of INSPIRE and elaborate the considerations which influenced its design. Furthermore, we demonstrate its expressiveness by illustrating the encoding of a variety of parallel language constructs and we evaluate its ability to preserve performance relevant aspects of input codes. Herbert Jordan, Simone Pellegrini, Peter Thoman, Klaus Kofler, Thomas Fahringer |
PACT | 3 |
| 2013 | Adaptive Granularity Control in Task Parallel Programs Using Multiversioning
Peter Thoman, Herbert Jordan, Thomas Fahringer |
Euro-Par | 1 |
| 2012 | A multi-objective auto-tuning framework for parallel codesabstractIn this paper we introduce a multi-objective autotuning framework comprising compiler and runtime components. Focusing on individual code regions, our compiler uses a novel search technique to compute a set of optimal solutions, which are encoded into a multi-versioned executable. This enables the runtime system to choose specifically tuned code versions when dynamically adjusting to changing circumstances. We demonstrate our method by tuning loop tiling in cache-sensitive parallel programs, optimizing for both runtime and efficiency. Our static optimizer finds solutions matching or surpassing those determined by exhaustively sampling the search space on a regular grid, while using less than 4% of the computational effort on average. Additionally, we show that parallelism-aware multi-versioning approaches like our own gain a performance improvement of up to 70% over solutions tuned for only one specific number of threads. Herbert Jordan, Peter Thoman, Juan José Durillo, Simone Pellegrini, Philipp Gschwandtner, Thomas Fahringer, Hans Moritsch |
SC | 2 |
| 2011 | Automatic OpenCL Device Characterization: Guiding Optimized Kernel Design
Peter Thoman, Klaus Kofler, Heiko Studt, John Thomson, Thomas Fahringer |
Euro-Par (2) | 1 |
| 2008 | GPU-Based Multigrid: Real-Time Performance in High Resolution Nonlinear Image Processing
Harald Grossauer, Peter Thoman |
ICVS | 2 |