EDBT 2026 Demo / reviewers in the wild / expert
Tobias Gysi
dblp:157/2868
· DBLP profile ↗
8ranked-venue papers
6as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 45% GPUs and heterogeneous computing · 15% Performance modeling and evaluation · 10% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 85% Program synthesis and code generation · 15% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.7 | 2 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
Compilers and program optimization
compiler infrastructure |
0.5 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
Compilers and program optimization › compiler infrastructure
MLIR |
0.5 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
High-performance computing › scientific computing systems
climate modeling |
0.5 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
Memory systems › cache
cache performance |
0.4 | 1 | 2019 | A fast analytical model of fully associative caches · PLDI 2019 |
Performance modeling and evaluation
cache performance modeling |
0.4 | 1 | 2019 | A fast analytical model of fully associative caches · PLDI 2019 |
Distributed systems › communication optimization
communication-computation overlap |
0.2 | 1 | 2016 | dCUDA: hardware supported overlap of computation and communication · SC 2016 |
Processor architecture and microarchitecture
latency hiding |
0.2 | 1 | 2016 | dCUDA: hardware supported overlap of computation and communication · SC 2016 |
Program synthesis and code generation
domain-specific code generation |
0.2 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
Compilers and program optimization › loop optimization
stencil computation optimization |
0.2 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
High-performance computing › large-scale simulation
climate and weather simulation |
0.2 | 1 | 2014 | Application Centric Energy-Efficiency Study of Distributed Multi-Core and Hybrid CPU-GPU Systems · SC 2014 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
High-performance computing
stencil computation |
0.1 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster |
0.1 | 1 | 2016 | dCUDA: hardware supported overlap of computation and communication · SC 2016 |
Environmental and earth informatics › atmospheric modeling
weather and climate modeling |
0.1 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
High-performance computing
many-core acceleration |
0.1 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
Methods — techniques the papers use, named apart from their topics
MLIR · 1.0LLVM · 1.0domain-specific language · 0.7architecture-dependent code generation · 0.7symbolic counting · 0.4target notification · 0.2over-decomposition · 0.2CUDA · 0.2strong scaling · 0.2energy measurement · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate SimulationabstractMost compilers have a single core intermediate representation (IR) (e.g., LLVM) sometimes complemented with vaguely defined IR-like data structures. This IR is commonly low-level and close to machine instructions. As a result, optimizations relying on domain-specific information are either not possible or require complex analysis to recover the missing information. In contrast, multi-level rewriting instantiates a hierarchy of dialects (IRs), lowers programs level-by-level, and performs code transformations at the most suitable level. We demonstrate the effectiveness of this approach for the weather and climate domain. In particular, we develop a prototype compiler and design stencil- and GPU-specific dialects based on a set of newly introduced design principles. We find that two domain-specific optimizations (500 lines of code) realized on top of LLVM’s extensible MLIR compiler infrastructure suffice to outperform state-of-the-art solutions. In essence, multi-level rewriting promises to herald the age of specialized compilers composed from domain- and target-specific dialects implemented on top of a shared infrastructure. Tobias Gysi, Oleksandr Zinenko, Stephan Herhut, Eddie Davis, Tobias Wicky, Oliver Fuhrer, Torsten Hoefler, Tobias Grosser |
ACM Trans. Archit. Code Optim. | 1 |
| 2020 | Automatic Generation of Multi-Objective Polyhedral Compiler TransformationsabstractTo this day, polyhedral optimizing compilers use either extremely rigid (but accurate) cost models, one-size-fits-all general-purpose heuristics, or auto-tuning strategies to traverse and evaluate large optimization spaces. In this paper, we introduce an adaptive and automatic scheduler that permits to generate novel loop transformation sequences (or recipes) capable of delivering strong performance for a variety of different architectures without relying on auto-tuning, nor on pre-determined transformation strategies. We evaluate our approach using the Polybench/C benchmark suite against two modern state-of-the-art optimizers on three different architectures: An AMD ThreadRipper, an Intel Xeon Phi, and an Intel Xeon Platinum. Our results provide evidence that a set of high-level objectives backed up by an automatic adaptive scheduler (i.e., not hard-wired) is capable of achieving competitive performance, while only resorting to evaluating a handful of tuned variants. Lorenzo Chelini, Tobias Gysi, Tobias Grosser, Martin Kong, Henk Corporaal |
PACT | 2 |
| 2019 | Absinthe: Learning an Analytical Performance Model to Fuse and Tile Stencil Codes in One ShotabstractExpensive data movement makes the optimal target-specific selection of data-locality transformations essential. Loop fusion and tiling are the most important data-locality transformations. Their optimal selection is hard since good tile size choices inherently depend on the fusion choices and vice versa. Existing approaches avoid this difficulty by optimizing independent analytical models or by reverting to heuristics. Absinthe formulates the first unified linear optimization problem to derive single shot fusion and tile size decisions for stencil codes. At the core of our optimization problem, we place a learned analytic performance model that captures the characteristics of the target system. The tuned application kernels demonstrate excellent performance within 10% of exhaustively auto-tuned versions and up to 74% faster than the results of independent optimization with max fusion heuristic and Absinthe tile size selection. While the full search space is non-linear, bounding it to relevant solutions enables the efficient exploration of the exponential search space using linear solvers. As a result, the tuning of our application kernels takes less than one minute. Our approach thus establishes the foundations for next-generation compilers, which exploit empirical information to guide target-specific code transformations. Tobias Gysi, Tobias Grosser, Torsten Hoefler |
PACT | 1 |
| 2019 | A fast analytical model of fully associative cachesabstractWhile the cost of computation is an easy to understand local property, the cost of data movement on cached architectures depends on global state, does not compose, and is hard to predict. As a result, programmers often fail to consider the cost of data movement. Existing cache models and simulators provide the missing information but are computationally expensive. We present a lightweight cache model for fully associative caches with least recently used (LRU) replacement policy that gives fast and accurate results. We count the cache misses without explicit enumeration of all memory accesses by using symbolic counting techniques twice: 1) to derive the stack distance for each memory access and 2) to count the memory accesses with stack distance larger than the cache size. While this technique seems infeasible in theory, due to non-linearities after the first round of counting, we show that the counting problems are sufficiently linear in practice. Our cache model often computes the results within seconds and contrary to simulation the execution time is mostly problem size independent. Our evaluation measures modeling errors below 0.6% on real hardware. By providing accurate data placement information we enable memory hierarchy aware software development. Tobias Gysi, Tobias Grosser, Laurin Brandner, Torsten Hoefler |
PLDI | 1 |
| 2016 | dCUDA: hardware supported overlap of computation and communicationabstractOver the last decade, CUDA and the underlying GPU hardware architecture have continuously gained popularity in various high-performance computing application domains such as climate modeling, computational chemistry, or machine learning. Despite this popularity, we lack a single coherent programming model for GPU clusters. We therefore introduce the dCUDA programming model, which implements device-side remote memory access with target notification. To hide instruction pipeline latencies, CUDA programs over-decompose the problem and over-subscribe the device by running many more threads than there are hardware execution units. Whenever a thread stalls, the hardware scheduler immediately proceeds with the execution of another thread ready for execution. This latency hiding technique is key to make best use of the available hardware resources. With dCUDA, we apply latency hiding at cluster scale to automatically overlap computation and communication. Our benchmarks demonstrate perfect overlap for memory bandwidth-bound tasks and good overlap for compute-bound tasks. Tobias Gysi, Jeremia Bär, Torsten Hoefler |
SC | 1 |
| 2015 | MODESTO: Data-centric Analytic Optimization of Complex Stencil Programs on Heterogeneous ArchitecturesabstractCode transformations, such as loop tiling and loop fusion, are of key importance for the efficient implementation of stencil computations. However, their direct application to a large code base is costly and severely impacts program maintainability. While recently introduced domain-specific languages facilitate the application of such transformations, they typically still require manual tuning or auto-tuning techniques to select the transformations that yield optimal performance. In this paper, we introduce MODESTO, a model-driven stencil optimization framework, that for a stencil program suggests program transformations optimized for a given target architecture. Initially, we review and categorize data locality transformations for stencil programs and introduce a stencil algebra that allows the expression and enumeration of different stencil program implementation variants. Combining this algebra with a compile-time performance model, we show how to automatically tune stencil programs. We use our framework to model the STELLA library and optimize kernels used by the COSMO atmospheric model on multi-core and hybrid CPU-GPU architectures. Compared to naive and expert-tuned variants, the automatically tuned kernels attain a 2.0-3.1x and a 1.0-1.8x speedup respectively. Tobias Gysi, Tobias Grosser, Torsten Hoefler |
ICS | 1 |
| 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate modelsabstractMany high-performance computing applications solving partial differential equations (PDEs) can be attributed to the class of kernels using stencils on structured grids. Due to the disparity between floating point operation throughput and main memory bandwidth these codes typically achieve only a low fraction of peak performance. Unfortunately, stencil computation optimization techniques are often hardware dependent and lead to a significant increase in code complexity. We present a domain-specific tool, STELLA, which eases the burden of the application developer by separating the architecture dependent implementation strategy from the user-code and is targeted at multi- and manycore processors. On the example of a numerical weather prediction and regional climate model (COSMO) we demonstrate the usefulness of STELLA for a real-world production code. The dynamical core based on STELLA achieves a speedup factor of 1.8x (CPU) and 5.8x (GPU) with respect to the legacy code while reducing the complexity of the user code. Tobias Gysi, Carlos Osuna, Oliver Fuhrer, Mauro Bianco, Thomas C. Schulthess |
SC | 1 |
| 2014 | Application Centric Energy-Efficiency Study of Distributed Multi-Core and Hybrid CPU-GPU SystemsabstractWe study the energy used by a production-level regional climate and weather simulation code on a distributed memory system with hybrid CPU-GPU nodes. The code is optimised for both processor architectures, for which we investigate both time and energy to solution. Operational constraints for time to solution can be met with both processor types, although on different numbers of nodes. Energy to solution is a factor 3 lower with GPUs, but strong scaling can be pushed to larger node counts with CPUs to minimize time to solution. Our data shows that an affine relationship exists between energy and node hours consumed by the simulation. We use this property to devise a simple and practical methodology for optimising for energy efficiency that can be applied to other applications, which we demonstrate with the HPCG benchmark. We conclude with a discussion about the relationship to the commonly-used GF/Watt metric. Ben Cumming, Gilles Fourestey, Oliver Fuhrer, Tobias Gysi, Massimiliano Fatica, Thomas C. Schulthess |
SC | 4 |