VLDB 2026 Research / reviewers in the wild / expert
Matthias Korch
dblp:k/MatthiasKorch
· DBLP profile ↗
33ranked-venue papers
21as first author
6since 2021 · last 2026
0000-0001-7267-0378ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 17 first-author · 3 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Efficient are the Efficient Cores? - An Experimental Evaluation of Energy Efficiency in Asymmetric Multicore Processors
Hana Shatri Ahmeti, Matthias Korch, Tim Werner 0001, Thomas Rauber |
Euro-Par (1) | 2 |
| 2026 | Can Microbenchmark-Derived Insights Guide Energy-Efficient Execution of Real Applications on Asymmetric Multicore Processors?
Hana Shatri Ahmeti, Matthias Korch, Tim Werner 0001, Thomas Rauber |
ISPDC | 2 |
| 2026 | ONACADI: An Online Autotuner for Compiled Applications Based on Debugger Interfaces
Fabian Mikula, Matthias Korch |
ISPDC | 2 |
| 2023 | Generation of logic designs for efficiently solving ordinary differential equations on field programmable gate arraysabstractAbstract Ordinary differential equations can be used to describe simulation models. As such, solving these equations is an important task in the high performance computing (HPC) domain. Field programmable gate arrays (FPGAs) are a promising platform, expected to be usable as efficient accelerators for such computations. While the use of hardware description languages (HDLs) can produce very efficient logic designs, their unique concept is hard to adopt for scientists and software engineers. High‐level synthesis (HLS) tools promise faster development, but bear the risk of lower performance and increased resource consumption of the final design. But even when using HLS tools the user still requires specialized knowledge about FPGAs and circuit design. In order to reach a wide adoption of FPGAs in HPC applications, a need for simple to use tools which enable performant designs was identified. This article proposes a framework that is able to automatically generate specific and optimized solver logic from easy to handle configuration files. No manual development, nor special FPGA or programming knowledge is required. To measure the capability of the proposed tool, the performance was evaluated for different solver methods and compared with an alternative hand optimized HLS implementation. The logic generated by this improved approach is up to 43 times faster than its hand optimized HLS counterpart, depending on the solution method. Silas Bartel, Matthias Korch |
Softw. Pract. Exp. | 2 |
| 2021 | YaskSite: Stencil Optimization Techniques Applied to Explicit ODE Methods on Modern ArchitecturesabstractThe landscape of multi-core architectures is growing more complex and diverse. Optimal application performance tuning parameters can vary widely across CPUs, and finding them in a possibly multidimensional parameter search space can be time consuming, expensive and potentially infeasible. In this work, we introduce YaskSite, a tool capable of tackling these challenges for stencil computations. YaskSite is built upon Intel's YASK framework. It combines YASK's flexibility to deal with different target architectures with the Execution-Cache-Memory performance model, which enables identifying optimal performance parameters analytically without the need to run the code. Further we show that YaskSite's features can be exploited by external tuning frameworks to reliably select the most efficient kernel(s) for the application at hand. To demonstrate this, we integrate YaskSite into Offsite, an offline tuner for explicit ordinary differential equation methods, and show that the generated performance predictions are reliable and accurate, leading to considerable performance gains at minimal code generation time and autotuning costs on the latest Intel Cascade Lake and AMD Rome CPUs. Christie L. Alappat, Johannes Seiferth, Georg Hager, Matthias Korch, Thomas Rauber, Gerhard Wellein |
CGO | 4 |
| 2021 | An in-depth introduction of multi-workgroup tiling for improving the locality of explicit one-step methods for ODE systems with limited access distance on GPUsabstractSummary This article considers a locality optimization technique for the parallel solution of a special class of large systems of ordinary differential equations (ODEs) by explicit one‐step methods on GPUs. This technique is based on tiling across the stages of the one‐step method and is enabled by the special structure of the class of ODE systems considered, that is, the limited access distance. The focus of this article is on increasing the range of access distances for which the tiling technique can provide a speedup by joining the memory resources and the computational power of multiple workgroups for the computation of one tile (multi‐workgroup tiling). In particular, this article provides an extended in‐depth introduction and discussion of the multi‐workgroup tiling technique and its theoretical and technical foundations together with a new tuning option (mapping stride) and new experiments. The experiments performed show speedups of the multi‐workgroup tiling technique compared with traditional single‐workgroup tiling for two different Runge–Kutta methods on NVIDIAs Kepler and Volta architectures. Matthias Korch, Tim Werner 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Implementation and Optimization of a 1D2V PIC Method for Nonlinear Kinetic Models on GPUsabstractThis paper considers the parallel numerical simulation of the time evolution of galaxies and globular clusters on GPUs. The model used is the Einstein-Vlasov system, which is designed, in particular, to study the formation of black holes and spacetime singularities in a general relativistic framework.First, a reference implementation is derived using NVIDIA CUDA as programming model, which is then optimized in several steps. Bottlenecks are identified by profiling, and different approaches, namely particle sort, improved treatment of atomic operations, and kernel fusion are investigated to overcome these bottlenecks. Each optimized variant is evaluated in relation to the other variants using detailed runtime experiments and profiling results. Using in the order of 107to 108particles, speedups between 1.84 and 2.38 w.r.t. the reference implementation have been observed. Matthias Korch, Philipp Raithel, Tim Werner 0001 |
PDP | 1 |
| 2020 | Improving locality of explicit one-step methods on GPUs by tiling across stages and time steps
Matthias Korch, Tim Werner 0001 |
Future Gener. Comput. Syst. | 1 |
| 2019 | Performance Prediction of Explicit ODE Methods on Multi-Core Cluster SystemsabstractWhen migrating a scientific application to a new HPC system, the program code usually has to be re-tuned to achieve the best possible performance. Auto-tuning techniques are a promising approach to support the portability of performance. Often, a large pool of possible implementation variants exists from which the most efficient variant needs to be selected. Ideally, auto-tuning approaches should be capable of undertaking this task in an efficient manner for a new HPC system and new characteristics of the input data by applying suitable analytic models and program transformations. Markus Scherg, Johannes Seiferth, Matthias Korch, Thomas Rauber |
ICPE | 3 |
| 2018 | Exploiting Limited Access Distance for Kernel Fusion Across the Stages of Explicit One-Step Methods on GPUsabstractThe performance of explicit parallel methods solving large systems of ordinary differential equations (ODEs) on GPUs is often memory bound. Therefore, locality optimizations, such as kernel fusion, are desirable. This paper exploits a special property of a large class of right-hand-side (RHS) functions to enable the fusion of computations of blocks of components across multiple stages of the method. This leads to a tiling of the stages within one time step. Our approach is based on a representation of the ODE method by a data flow graph and allows efficient GPU code with fused kernels to be generated automatically for user-defined tilings. In particular, we investigate two generalized tiling strategies, trapezoidal and hexagonal tiling, which are evaluated experimentally for several different high-order Runge-Kutta (RK) methods. Matthias Korch, Tim Werner 0001 |
SBAC-PAD | 1 |
| 2018 | Accelerating explicit ODE methods on GPUs by kernel fusionabstractSummary Graphics processing units (GPUs) have a promising architecture for implementing highly parallel solution methods for systems of ordinary differential equations (ODEs). However, their high performance comes at the price of caveats such as small caches or wide SIMD. For ODE methods, optimizing the memory access pattern is often crucial. In this article, instead of considering only one specific method, we generalize the description of explicit ODE methods by using data flow graphs consisting of basic operations that are suitable to cover the types of computations occurring in all common explicit methods. After showing that the straightforward approach for processing the data flow graph by calling one kernel per basic operation is memory bound, we explain how the number of memory accesses can be reduced by the kernel fusion technique, which fuses several basic operations into one kernel. Moreover, we will present enabling transformations that allow additional fusions and thus can reduce the number of memory accesses even further. We apply these optimizations to three different classes of explicit ODE methods: embedded Runge–Kutta (RK) methods, parallel iterated RK (PIRK) methods, and peer methods. A detailed experimental evaluation on three modern GPUs showed speedups between 1.86 and 3.51 compared to unfused implementations. Matthias Korch, Tim Werner 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Influence of Locality on the Scalability of Method-and System-Parallel Explicit Peer MethodsabstractBecause the numerical solution of initial value problems (IVPs) of systems of ordinary differential equations (ODEs) can be computationally intensive, several parallel methods have been proposed in the past.One class of modern parallel IVP methods are the peer methods proposed by Schmitt and Weiner, some of which are publicly available in the software package EPPEER released in 2012.Since they possess eight independent stages, these methods offer natural parallelism across the method suitable for the typical numbers of CPU cores in modern multicore workstations.EPPEER is written in FORTRAN95 and uses OpenMP as parallel programming model.In this paper, we investigate the influence of the locality of memory references on the scalability of method-and systemparallel explicit peer methods.In particular, we investigate the interplay between the linear combination of the stages and the function evaluations by applying different program transformations to the loop structure and by evaluating their performance in detailed runtime experiments.These experiments point out that loop tiling is required to improve cache utilization while still allowing the compiler to vectorize along the system dimension.To show that for certain classes of right-hand-side functions a stage-parallel execution is not optimal, and to enhance the scalability of the peer methods to core numbers larger than the number of stages of a method, system-parallel implementations have been derived.Runtime experiments show that there are IVPs for which these new implementations outperform stage-parallel implementations on numbers of cores less than or equal to the number of stages.Moreover, by exploiting the ability to utilize higher core numbers, higher speedups than the number of stages have been reached. Matthias Korch, Thomas Rauber, Matthias Stachowski, Tim Werner 0001 |
FedCSIS | 1 |
| 2014 | Online auto-tuning for the time-step-based parallel solution of ODEs on shared-memory systems
Natalia Kalinnik, Matthias Korch, Thomas Rauber |
J. Parallel Distributed Comput. | 2 |
| 2013 | MAP: Mobile Assistance Platform with a VM Type Selection AbilityabstractThe usage of remote compute and storage resources is becoming popular to assist mobile devices. However, existing frameworks that support mobile client/server applications do not consider the provision of resources in a large scale. Although a cloud-based infrastructure may provide a large number of resources on demand, most cloud offerings do not provide the right level of abstraction to assist mobile devices. Infrastructure-as-a-Service (IaaS) offerings are too complex to set up on demand and Platform-as-a-Service (PaaS) offerings are not flexible enough regarding Quality-of-Service (QoS) adjustment. To overcome these issues, we propose a middleware named Mobile Assistance Platform (MAP), which serves as an extended PaaS layer. It provides an on-demand execution platform for program code, but with the additional ability to select the underlying VM quality for execution. Furthermore, each mobile request is assigned to a single VM instance for processing. MAP provides configurable compute resources for mobile users in a large scale. For this demo, we have set up MAP on the Amazon EC2 infrastructure and we present two mobile applications that can benefit from the on-demand resource quality selection. Marvin Ferber, Natalia Kalinnik, Matthias Korch, Andreas Prell, Thomas Rauber, Matthias Witzgall |
ICPADS | 3 |
| 2013 | Parallelization of Particle-in-Cell Codes for Nonlinear Kinetic Models from Mathematical PhysicsabstractThis paper considers the parallelization of two Particle-in-Cell (PIC) codes which simulate the time evolution of galaxies and globular clusters in the Newtonian or the general relativistic framework. The corresponding models are known as the Vlasov-Poisson or the Einstein-Vlasov system, and the latter is designed in particular to study the formation of black holes and space time singularities. We start with a step-by-step shared-memory parallelization of the Vlasov-Poisson code using POSIX Threads and finally develop message passing codes using MPI. The parallel codes have been investigated on three modern supercomputer systems using up to 4096 cores, and speedups above 1300 have been reached. The speedup obtained through parallelization has already helped in finding new numerical results, such as oscillating solutions of the Vlasov-Poisson system. Matthias Korch, Tobias Ramming, Gerhard Rein |
ICPP | 1 |
| 2012 | Locality Improvement of Data-Parallel Adams-Bashforth Methods through Block-Based Pipelining of Time Steps
Matthias Korch |
Euro-Par | 1 |
| 2012 | Diamond-Like Tiling Schemes for Efficient Explicit Euler on GPUsabstractGPU computing offers a high potential of raw processing power at comparatively low costs. This paper investigates optimization techniques for solving initial value problems (IVPs) of ordinary differential equations (ODEs) on GPUs. Different techniques, especially for exploiting the GPU memory hierarchy, are discussed, and corresponding OpenCL implementations of the explicit Euler method are compared using runtime experiments. The results show considerable performance improvements in many situations. Due to the basic character of the explicit Euler method, the results of this investigation can guide the optimization of more complex ODE methods with higher order and better stability on GPUs. Matthias Korch, Julien Kulbe, Carsten Scholtes |
ISPDC | 1 |
| 2011 | Memory-Intensive Applications on a Many-Core ProcessorabstractFuture micro-processors are expected to contain an increasing number of cores. Different models exist for efficiently organizing the cores of the resulting many-core processors. The Single-Chip Cloud Computer (SCC) is an experimental processor created by Intel Labs. It is optimized for providing to each core a programming model similar to that of the nodes of a message-passing distributed system. We have examined the performance of a memory-intensive application on the SCC. The application solves Initial Value Problems (IVPs) of Ordinary Differential Equations (ODEs). Experiments with different configurations and optimizations of this application have been performed. The evaluation of these experiments reveals bottlenecks and provides hints for optimizing applications for similar many-core architectures. Matthias Korch, Thomas Rauber, Carsten Scholtes |
HPCC | 1 |
| 2011 | Dynamic selection of implementation variants of sequential iterated runge-kutta methods with tile size samplingabstractThis paper describes an efficient self-adaptive procedure for iterated Runge-Kutta (IRK) methods, a class of solution methods for initial value problems (IVPs) of ordinary differential equations (ODEs). IRK methods execute a potentially large number of discrete time steps to compute the solution of the IVP. The performance of an IRK solver may strongly depend on the specific characteristics of the given IVP and the hardware architecture on which the solver is executed. To address this problem, this paper applies dynamic auto-tuning to the sequential execution of IRK methods. Auto-tuning is a promising technique to avoid time consuming and extensive manual tuning. Our self-adaptive IRK solver utilizes the time-stepping nature of the IRK method. It selects the fastest implementation variant for the given IVP on the target architecture from a candidate pool during the first time steps. Then, the fastest implementation variant is used to compute all remaining time steps. The different implementation variants in the candidate pool have been developed by modifications of the loop structure of the basic algorithm. For those implementation variants that use loop tiling, we consider different tile sizes during the auto-tuning phase to further improve the performance of the self-adaptive IRK solver. Runtime experiments demonstrate the efficiency of the self-adaptive IRK solver for different IVPs on different hardware architectures. Natalia Kalinnik, Matthias Korch, Thomas Rauber |
ICPE | 2 |
| 2011 | Scalability and locality of extrapolation methods on large parallel systemsabstractAbstract Time‐dependent processes can often be modeled by systems of ordinary differential equations (ODEs). Solving such a system for a detailed model can be highly computationally intensive. We investigate explicit extrapolation methods for solving such systems efficiently on current highly parallel supercomputer systems with shared‐or distributed‐memory architecture. We analyze and compare the scalability of several parallelization variants, some of them using multiple levels of parallelization. For a large class of ODE systems, data access costs are reduced considerably by exploiting the special structure of the ODE system. Furthermore, by employing a pipeline‐like loop structure, the locality of memory references is increased for such systems resulting in a better utilization of the cache hierarchy. Runtime experiments show that the optimized implementations can deliver a high scalability. Copyright © 2011 John Wiley & Sons, Ltd. Matthias Korch, Thomas Rauber, Carsten Scholtes |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Scalability and Locality of Extrapolation Methods for Distributed-Memory Architectures
Matthias Korch, Thomas Rauber, Carsten Scholtes |
Euro-Par (2) | 1 |
| 2010 | Mixed-Parallel Implementations of Extrapolation Methods with Reduced Synchronization Overhead for Large Shared-Memory ComputersabstractExtrapolation methods belong to the class of one-step methods for the solution of systems of ordinary differential equations (ODEs). In this paper, we present parallel implementation variants of extrapolation methods for large shared-memory computer systems which exploit pure data parallelism or mixed task and data parallelism and make use of different load balancing strategies and different loop structures. In addition to general implementation variants suitable for ODE systems with arbitrary access structure, we devise specialized implementation variants which exploit the specific access structure of a large class of ODE systems to reduce synchronization costs and to improve the locality of memory references. We analyze and compare the scalability and the locality behavior of the implementation variants on an SGI Altix 4700 using up to 500 threads. Matthias Korch, Thomas Rauber, Carsten Scholtes |
ICPADS | 1 |
| 2009 | Parallel Implementation of Runge-Kutta Integrators with Low Storage Requirements
Matthias Korch, Thomas Rauber |
Euro-Par | 1 |
| 2009 | Scalability of Time- and Space-Efficient Embedded Runge-Kutta Solvers for Distributed Address SpaceabstractEmbedded Runge-Kutta methods are well-known and efficient solution methods for initial value problems of ordinary differential equations (ODEs). In this paper, we discuss the parallel implementation of embedded Runge-Kutta methods for distributed address space. Our focus lies on the exploitation of a special structure of commonly appearing ODE systems to improve scalability and memory usage. We show how the memory space of a pipeline-like computation scheme can be reduced to less than three storage registers by an overlapping of vectors without compromising the choice of method coefficients or the potential for efficient stepsize control. We analyze and compare the scalability of different implementation strategies in detailed runtime experiments on different modern parallel architectures. These experiments show that our approach leads to a good scalability behavior even on large numbers of processors. Matthias Korch, Thomas Rauber |
ICPP | 1 |
| 2008 | Transformation of Legacy Software into Client/Server Applications through Pattern-Based RearchitecturingabstractIn this article, we address the problem of modularizing legacy applications with monolithic structure, primarily focusing on business software written in an object-oriented programming language. We introduce theTransFormr toolkit that guides the developer through the entire incremental transformation process. It is the goal of the transformation to separate the original software into several independent replaceable components to support the migration of legacy code to new hardware or to integrate legacy components into modern enterprise applications. We show the effectiveness of our approach by demonstrating a pattern-based transformation of classes in a case study. Sascha Hunold, Matthias Korch, Björn Krellner, Thomas Rauber, Thomas Reichel, Gudula Rünger |
COMPSAC | 2 |
| 2007 | Locality Optimized Shared-Memory Implementations of Iterated Runge-Kutta Methods
Matthias Korch, Thomas Rauber |
Euro-Par | 1 |
| 2006 | Applicability of Load Balancing Strategies to Data-Parallel Embedded Runge-Kutta Integrators
Matthias Korch, Thomas Rauber |
Euro-Par | 1 |
| 2006 | Optimizing locality and scalability of embedded Runge-Kutta solvers using block-based pipelining
Matthias Korch, Thomas Rauber |
J. Parallel Distributed Comput. | 1 |
| 2004 | Using Hardware Operations to Reduce the Synchronization Overhead of Task PoolsabstractWe consider the task-based execution of parallel irregular applications, which are characterized by an unpredictable computational structure induced by the input data. The dynamic load balancing required to execute such applications efficiently can be provided by task pools. Thus, the performance of a task-based irregular application is tightly coupled to the scalability and the overhead of the task pool used to execute it. In order to reduce this overhead this article considers the use of the hardware-specific synchronization operations compare & swap and load & reserve/store conditional. We present several different realizations of task pools using these operations. Runtime experiments on two shared-memory machines, a SunFire 6800 and an IBM p690, show that the new implementations obtain a significantly higher performance than implementations relying on the POSIX thread library for synchronization. Ralf Hoffmann, Matthias Korch, Thomas Rauber |
ICPP | 2 |
| 2004 | Performance Evaluation of Task Pools Based on Hardware SynchronizationabstractA task-based execution provides a universal approach to dynamic load balancing for irregular applications. Tasks are arbitrary units of work that are created dynamically at run-time and that are stored in a parallel data structure, the task pool, until they are scheduled onto a processor for execution. In this paper, we evaluate the performance of different task pool implementations for shared-memory computer systems using several realistic applications. We consider task pools with different data structures, different load balancing strategies and a specialized memory management. In particular, we use synchronization operations based on hardware support that is available on many modern micro-processors. We show that the resulting task pool implementations lead to a much better performance than implementations using Pthreads library calls for synchronization. The applications considered are parallel quicksort, volume rendering, ray tracing, and hierarchical radiosity. The target machines are an IBM p690 server and a SunFire 6800. Ralf Hoffmann, Matthias Korch, Thomas Rauber |
SC | 2 |
| 2004 | A comparison of task pools for dynamic load balancing of irregular algorithmsabstractAbstract Since a static work distribution does not allow for satisfactory speed‐ups of parallel irregular algorithms, there is a need for a dynamic distribution of work and data that can be adapted to the runtime behavior of the algorithm. Task pools are data structures which can distribute tasks dynamically to different processors where each task specifies computations to be performed and provides the data for these computations. This paper discusses the characteristics of task‐based algorithms and describes the implementation of selected types of task pools for shared‐memory multiprocessors. Several task pools have been implemented in C with POSIX threads and in Java. The task pools differ in the data structures to store the tasks, the mechanism to achieve load balance, and the memory manager used to store the tasks. Runtime experiments have been performed on three different shared‐memory systems using a synthetic algorithm, the hierarchical radiosity method, and a volume rendering algorithm. Copyright © 2004 John Wiley & Sons, Ltd. Matthias Korch, Thomas Rauber |
Concurr. Comput. Pract. Exp. | 1 |
| 2003 | Scalable Parallel RK Solvers for ODEs Derived by the Method of Lines
Matthias Korch, Thomas Rauber |
Euro-Par | 1 |
| 2002 | Pipelining for Locality Improvement in RK Methods
Matthias Korch, Thomas Rauber, Gudula Rünger |
Euro-Par | 1 |