VLDB 2026 Research / reviewers in the wild / expert
Ari Rasch
dblp:204/7105
· DBLP profile ↗
10ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-0286-0755ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | pyATF: Constraint-Based Auto-Tuning in PythonabstractWe introduce pyATF -- a new, language-independent, open-source auto-tuning tool that fully automatically determines optimized values of performance-critical program parameters. A major feature of pyATF is its support for constrained parameters, e.g., the value of one parameter has to divide the value of another parameter. A further major feature of pyATF is its user interface which is designed with a particular focus on expressivity and usability for real-world demands, and which is offered in the increasingly popular Python programming language. We experimentally confirm the practicality of pyATF using real-world studies from the areas of quantum chemistry, image processing, data mining, and deep learning: we show that pyATF auto-tunes the complex parallel implementations of our studies to higher performance than achieved by state-of-practice approaches, including hand-optimized vendor libraries. Richard Schulze, Sergei Gorlatch, Ari Rasch |
CC | 3 |
| 2025 | Scheduling Language Chronology: Past, Present, and FutureabstractScheduling languages express to a compiler—or equivalently, a code generator—a sequence of optimizations to apply. Performance tools that support a scheduling language interface allow exploration of optimizations, i.e., exploratory compilers . While scheduling languages have become a common feature of tools for experts, the proliferation of these languages without unifying common features may be confusing to users. Moreover, we recognize a need to organize the compiler developer community around common exploratory compiler infrastructure, and future advances to address, for example, data layout and data movement. To support a broader set of users may require raising the level of abstraction. This article provides a chronology of scheduling languages, discussing their origins in iterative compilation and autotuning, noting the common features that are used in existing frameworks, and calling for changes to increase their utility and portability. Mary W. Hall, Cosmin E. Oancea, Anne C. Elster, Ari Rasch, Sameeran Joshi, Amir Mohammad Tavakkoli, Richard Schulze |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | (De/Re)-Composition of Data-Parallel Computations via Multi-Dimensional HomomorphismsabstractData-parallel computations, such as linear algebra routines and stencil computations, constitute one of the most relevant classes in parallel computing, e.g., due to their importance for deep learning. Efficiently de-composing such computations for the memory and core hierarchies of modern architectures and re-composing the computed intermediate results back to the final result—we say (de/re)-composition for short—is key to achieve high performance for these computations on, e.g., GPU and CPU. Current high-level approaches to generating data-parallel code are often restricted to a particular subclass of data-parallel computations and architectures (e.g., only linear algebra routines on only GPU or only stencil computations), and/or the approaches rely on a user-guided optimization process for a well-performing (de/re)-composition of computations, which is complex and error prone for the user. We formally introduce a systematic (de/re)-composition approach, based on the algebraic formalism of Multi-Dimensional Homomorphisms (MDHs). Our approach is designed as general enough to be applicable to a wide range of data-parallel computations and for various kinds of target parallel architectures. To efficiently target the deep and complex memory and core hierarchies of contemporary architectures, we exploit our introduced (de/re)-composition approach for a correct-by-construction, parametrized cache blocking, and parallelization strategy. We show that our approach is powerful enough to express, in the same formalism, the (de/re)-composition strategies of different classes of state-of-the-art approaches (scheduling-based, polyhedral, etc.), and we demonstrate that the parameters of our strategies enable systematically generating code that can be fully automatically optimized (auto-tuned) for the particular target architecture and characteristics of the input and output data (e.g., their sizes and memory layouts). Particularly, our experiments confirm that via auto-tuning, we achieve higher performance than state-of-the-art approaches, including hand-optimized solutions provided by vendors (such as NVIDIA cuBLAS/cuDNN and Intel oneMKL/oneDNN), on real-world datasets and for a variety of data-parallel computations, including linear algebra routines, stencil and quantum chemistry computations, data mining algorithms, and computations that recently gained high attention due to their relevance for deep learning. Ari Rasch |
ACM Trans. Program. Lang. Syst. | 1 |
| 2023 | (De/Re)-Compositions Expressed Systematically via MDH-Based SchedulesabstractWe introduce a new scheduling language, based on the formalism of Multi-Dimensional Homomorphisms (MDH). In contrast to existing scheduling languages, our MDH-based language is designed to systematically "de-compose" computations for the memory and core hierarchies of architectures, and "re-compose" the computed intermediate results back to the final result -- we say "(de/re)-composition" for short. We argue that our scheduling langauge is easy to use and yet expressive enough to express well-performing (de/re)-compositions of popular related approaches, e.g., the TVM compiler, for MDH-supported computations (such as linear algebra routines and stencil computations). Moreover, our language is designed as auto-tunable, i.e., any optimization decision can optionally be left to the auto-tuning engine of our system, and our system can automatically recommend schedules for the user, based on its auto-tuning capabilities. Also, by relying on the MDH approach, we can formally guarantee the correctness of optimizations expressed in our language, thereby further enhancing user experience. Our experiments on GPU and CPU confirm that we can express optimizations that cannot be expressed straightforwardly (or at all) in TVM's scheduling language, thereby achieving higher performance than TVM, and also vendor libraries provided by NVIDIA and Intel, for time-intensive computations used in real-world deep learning neural networks. Ari Rasch, Richard Schulze, Denys Shabalin, Anne C. Elster, Sergei Gorlatch, Mary W. Hall |
CC | 1 |
| 2021 | Efficient Auto-Tuning of Parallel Programs with Interdependent Tuning Parameters via Auto-Tuning Framework (ATF)abstractAuto-tuning is a popular approach to program optimization: it automatically finds good configurations of a program’s so-called tuning parameters whose values are crucial for achieving high performance for a particular parallel architecture and characteristics of input/output data. We present three new contributions of the Auto-Tuning Framework (ATF), which enable a key advantage in general-purpose auto-tuning : efficiently optimizing programs whose tuning parameters have interdependencies among them. We make the following contributions to the three main phases of general-purpose auto-tuning: (1) ATF generates the search space of interdependent tuning parameters with high performance by efficiently exploiting parameter constraints; (2) ATF stores such search spaces efficiently in memory, based on a novel chain-of-trees search space structure; (3) ATF explores these search spaces faster, by employing a multi-dimensional search strategy on its chain-of-trees search space representation. Our experiments demonstrate that, compared to the state-of-the-art, general-purpose auto-tuning frameworks, ATF substantially improves generating, storing, and exploring the search space of interdependent tuning parameters, thereby enabling an efficient overall auto-tuning process for important applications from popular domains, including stencil computations, linear algebra routines, quantum chemistry computations, and data mining algorithms. Ari Rasch, Richard Schulze, Michel Steuwer, Sergei Gorlatch |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | dOCAL: high-level distributed programming with OpenCL and CUDA
Ari Rasch, Julian Bigge, Martin Wrodarczyk, Richard Schulze, Sergei Gorlatch |
J. Supercomput. | 1 |
| 2019 | Generating Portable High-Performance Code via Multi-Dimensional HomomorphismsabstractWe address a key challenge in programming high-performance applications - achieving portable performance, i.e., the same source code achieves a consistent, high level of performance over the variety of modern parallel processors, including multi-core CPU and many-core Graphics Processing Unit (GPU), and over the variety of input sizes. Our approach relies on the algebraic formalism of Multi-Dimensional Homomorphisms (MDH), which enables expressing data-parallel computations uniformly via a higher-order function (a.k.a. parallel pattern). For MDHs, we develop a novel code generation approach based on a generic OpenCL implementation. Our implementation efficiently exploits the OpenCL's abstract platform and memory model, generically for arbitrary MDH functions, by incorporating a parameterized parallelization and tiling strategy - on both layers of the OpenCL's two models and in all dimensions of the multi-dimensional input. We achieve performance portability for MDHs by auto-tuning the parameters of our two strategies, thereby enabling fully automatically optimizing our code for any combination of an MDH function, target device, and input size. We demonstrate for computations from four popular domains - dense linear algebra (BLAS), stencil computations, data mining, and tensor contractions - how we express them in the MDH formalism, and we experimentally show that our automatically generated and auto-tuned code for them achieves competitive and often significantly better performance than several state-of-practice approaches on both Intel multi-core CPU and NVIDIA many-core GPU - speedups of up to 5x over the state-of-the-art performance-portable approaches, and competitive or even better performance as compared to hand-optimized approaches such as Intel MKL and NVIDIA cuBLAS on real-world input data as used in deep learning. Ari Rasch, Richard Schulze, Sergei Gorlatch |
PACT | 1 |
| 2019 | ATF: A generic directive-based auto-tuning frameworkabstractSummary We describe the Auto‐Tuning Framework (ATF) — a simple‐to‐use, generic approach and its implementation, as a framework for automatic program optimization by choosing the most suitable values of program parameters such as the number of parallel threads, tile sizes, etc. ATF combines four major advantages over the state‐of‐the‐art auto‐tuning: i) it is generic regarding the programming language, application domain, tuning objective (eg, high performance and/or low energy consumption), and search technique; ii) it can auto‐tune a broader class of applications by allowing tuning parameters to be interdependent, eg, when one parameter is divisible by another parameter; iii) it allows tuning parameters to have substantially larger ranges by implementing an optimized search space generation process; and iv) it is arguably simpler to use, eg, the ATF user prepares an application for auto‐tuning by annotating its source code with simple tuning directives. We demonstrate ATF's efficacy by comparing it to the state‐of‐the‐art auto‐tuning approaches, OpenTuner and CLTune; ATF shows better tuning results with less programmer's effort. Ari Rasch, Sergei Gorlatch |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | OCAL: An Abstraction for Host-Code Programming with OpenCL and CUDAabstractThe state-of-the-art parallel programming approaches OpenCL and CUDA require so-called host code for pro-gram's execution. Implementing host code is often a cumbersome task, especially when executing OpenCL and CUDA programs on systems with multiple devices, e.g., multi-core CPU and Graphics Processing Units (GPUs): the programmer is responsible for explicitly managing system's main memory and devices' memories, synchronizing computations with data transfers between main and/or devices' memories, and optimizing data transfers, e.g., by using pinned main memory for accelerating data transfers and overlapping the transfers with comnutations. In this paper, we present OCAL (DpenCL/CUDA Abstraction Layer) - a high-level approach to simplify the development of host code. OCAL combines five major advantages over the state-of-the-art high-level approaches: 1) it simplifies implementing both OpenCL and CUDA host code by providing a simple-to-use, uniform high-level host code abstraction API; 2) it supports executing arbitrary OpenCL and CUDA programs; 3) it simplifies implementing data-transfer optimizations by providing specially-optimized memory buffers, e.g., for conveniently using pinned main memory; 4) it optimizes memory management by automatically avoiding unnecessary data transfers; 5) it enables interoperability between OpenCL and CUDA host code by automatically managing the communication between OpenCL and CUDA data structures and by automatically translating between the OpenCL. and CUDA programming constructs. Our experiments demonstrate that OCAL significantly simplifies implementing host code with a low runtime overhead for abstraction. Ari Rasch, Martin Wrodarczyk, Richard Schulze, Sergei Gorlatch |
ICPADS | 1 |
| 2017 | eccCL: parallelized GPU implementation of Ensemble Classifier ChainsabstractBACKGROUND: Multi-label classification has recently gained great attention in diverse fields of research, e.g., in biomedical application such as protein function prediction or drug resistance testing in HIV. In this context, the concept of Classifier Chains has been shown to improve prediction accuracy, especially when applied as Ensemble Classifier Chains. However, these techniques lack computational efficiency when applied on large amounts of data, e.g., derived from next-generation sequencing experiments. By adapting algorithms for the use of graphics processing units, computational efficiency can be greatly improved due to parallelization of computations. RESULTS: Here, we provide a parallelized and optimized graphics processing unit implementation (eccCL) of Classifier Chains and Ensemble Classifier Chains. Additionally to the OpenCL implementation, we provide an R-Package with an easy to use R-interface for parallelized graphics processing unit usage. CONCLUSION: eccCL is a handy implementation of Classifier Chains on GPUs, which is able to process up to over 25,000 instances per second, and thus can be used efficiently in high-throughput experiments. The software is available at http://www.heiderlab.de . Mona Riemenschneider, Alexander Herbst, Ari Rasch, Sergei Gorlatch, Dominik Heider |
BMC Bioinform. | 3 |