Tobias Gysi

dblp:157/2868 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
High-performance computing · 45% GPUs and heterogeneous computing · 15% Performance modeling and evaluation · 10%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 85% Program synthesis and code generation · 15%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.722021
Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
Compilers and program optimization
compiler infrastructure
0.512021
Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021
Compilers and program optimization › compiler infrastructure
MLIR
0.512021
Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021
High-performance computing › scientific computing systems
climate modeling
0.512021
Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021
Memory systems › cache
cache performance
0.412019
A fast analytical model of fully associative caches · PLDI 2019
Performance modeling and evaluation
cache performance modeling
0.412019
A fast analytical model of fully associative caches · PLDI 2019
Distributed systems › communication optimization
communication-computation overlap
0.212016
dCUDA: hardware supported overlap of computation and communication · SC 2016
Processor architecture and microarchitecture
latency hiding
0.212016
dCUDA: hardware supported overlap of computation and communication · SC 2016
Program synthesis and code generation
domain-specific code generation
0.212015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
Compilers and program optimization › loop optimization
stencil computation optimization
0.212015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
High-performance computing › large-scale simulation
climate and weather simulation
0.212014
Application Centric Energy-Efficiency Study of Distributed Multi-Core and Hybrid CPU-GPU Systems · SC 2014
GPUs and heterogeneous computing
GPU computing
0.112021
Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021
High-performance computing
stencil computation
0.112021
Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster
0.112016
dCUDA: hardware supported overlap of computation and communication · SC 2016
Environmental and earth informatics › atmospheric modeling
weather and climate modeling
0.112015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015
High-performance computing
many-core acceleration
0.112015
STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015

Methods — techniques the papers use, named apart from their topics

MLIR · 1.0LLVM · 1.0domain-specific language · 0.7architecture-dependent code generation · 0.7symbolic counting · 0.4target notification · 0.2over-decomposition · 0.2CUDA · 0.2strong scaling · 0.2energy measurement · 0.2
YearPublicationVenuePosition
2021 Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation
abstract
Most compilers have a single core intermediate representation (IR) (e.g., LLVM) sometimes complemented with vaguely defined IR-like data structures. This IR is commonly low-level and close to machine instructions. As a result, optimizations relying on domain-specific information are either not possible or require complex analysis to recover the missing information. In contrast, multi-level rewriting instantiates a hierarchy of dialects (IRs), lowers programs level-by-level, and performs code transformations at the most suitable level. We demonstrate the effectiveness of this approach for the weather and climate domain. In particular, we develop a prototype compiler and design stencil- and GPU-specific dialects based on a set of newly introduced design principles. We find that two domain-specific optimizations (500 lines of code) realized on top of LLVM’s extensible MLIR compiler infrastructure suffice to outperform state-of-the-art solutions. In essence, multi-level rewriting promises to herald the age of specialized compilers composed from domain- and target-specific dialects implemented on top of a shared infrastructure.
Tobias Gysi, Oleksandr Zinenko, Stephan Herhut, Eddie Davis, Tobias Wicky, Oliver Fuhrer, Torsten Hoefler, Tobias Grosser
ACM Trans. Archit. Code Optim.1
2020 Automatic Generation of Multi-Objective Polyhedral Compiler Transformations
abstract
To this day, polyhedral optimizing compilers use either extremely rigid (but accurate) cost models, one-size-fits-all general-purpose heuristics, or auto-tuning strategies to traverse and evaluate large optimization spaces. In this paper, we introduce an adaptive and automatic scheduler that permits to generate novel loop transformation sequences (or recipes) capable of delivering strong performance for a variety of different architectures without relying on auto-tuning, nor on pre-determined transformation strategies. We evaluate our approach using the Polybench/C benchmark suite against two modern state-of-the-art optimizers on three different architectures: An AMD ThreadRipper, an Intel Xeon Phi, and an Intel Xeon Platinum. Our results provide evidence that a set of high-level objectives backed up by an automatic adaptive scheduler (i.e., not hard-wired) is capable of achieving competitive performance, while only resorting to evaluating a handful of tuned variants.
Lorenzo Chelini, Tobias Gysi, Tobias Grosser, Martin Kong, Henk Corporaal
PACT2
2019 Absinthe: Learning an Analytical Performance Model to Fuse and Tile Stencil Codes in One Shot
abstract
Expensive data movement makes the optimal target-specific selection of data-locality transformations essential. Loop fusion and tiling are the most important data-locality transformations. Their optimal selection is hard since good tile size choices inherently depend on the fusion choices and vice versa. Existing approaches avoid this difficulty by optimizing independent analytical models or by reverting to heuristics. Absinthe formulates the first unified linear optimization problem to derive single shot fusion and tile size decisions for stencil codes. At the core of our optimization problem, we place a learned analytic performance model that captures the characteristics of the target system. The tuned application kernels demonstrate excellent performance within 10% of exhaustively auto-tuned versions and up to 74% faster than the results of independent optimization with max fusion heuristic and Absinthe tile size selection. While the full search space is non-linear, bounding it to relevant solutions enables the efficient exploration of the exponential search space using linear solvers. As a result, the tuning of our application kernels takes less than one minute. Our approach thus establishes the foundations for next-generation compilers, which exploit empirical information to guide target-specific code transformations.
Tobias Gysi, Tobias Grosser, Torsten Hoefler
PACT1
2019 A fast analytical model of fully associative caches
abstract
While the cost of computation is an easy to understand local property, the cost of data movement on cached architectures depends on global state, does not compose, and is hard to predict. As a result, programmers often fail to consider the cost of data movement. Existing cache models and simulators provide the missing information but are computationally expensive. We present a lightweight cache model for fully associative caches with least recently used (LRU) replacement policy that gives fast and accurate results. We count the cache misses without explicit enumeration of all memory accesses by using symbolic counting techniques twice: 1) to derive the stack distance for each memory access and 2) to count the memory accesses with stack distance larger than the cache size. While this technique seems infeasible in theory, due to non-linearities after the first round of counting, we show that the counting problems are sufficiently linear in practice. Our cache model often computes the results within seconds and contrary to simulation the execution time is mostly problem size independent. Our evaluation measures modeling errors below 0.6% on real hardware. By providing accurate data placement information we enable memory hierarchy aware software development.
Tobias Gysi, Tobias Grosser, Laurin Brandner, Torsten Hoefler
PLDI1
2016 dCUDA: hardware supported overlap of computation and communication
abstract
Over the last decade, CUDA and the underlying GPU hardware architecture have continuously gained popularity in various high-performance computing application domains such as climate modeling, computational chemistry, or machine learning. Despite this popularity, we lack a single coherent programming model for GPU clusters. We therefore introduce the dCUDA programming model, which implements device-side remote memory access with target notification. To hide instruction pipeline latencies, CUDA programs over-decompose the problem and over-subscribe the device by running many more threads than there are hardware execution units. Whenever a thread stalls, the hardware scheduler immediately proceeds with the execution of another thread ready for execution. This latency hiding technique is key to make best use of the available hardware resources. With dCUDA, we apply latency hiding at cluster scale to automatically overlap computation and communication. Our benchmarks demonstrate perfect overlap for memory bandwidth-bound tasks and good overlap for compute-bound tasks.
Tobias Gysi, Jeremia Bär, Torsten Hoefler
SC1
2015 MODESTO: Data-centric Analytic Optimization of Complex Stencil Programs on Heterogeneous Architectures
abstract
Code transformations, such as loop tiling and loop fusion, are of key importance for the efficient implementation of stencil computations. However, their direct application to a large code base is costly and severely impacts program maintainability. While recently introduced domain-specific languages facilitate the application of such transformations, they typically still require manual tuning or auto-tuning techniques to select the transformations that yield optimal performance. In this paper, we introduce MODESTO, a model-driven stencil optimization framework, that for a stencil program suggests program transformations optimized for a given target architecture. Initially, we review and categorize data locality transformations for stencil programs and introduce a stencil algebra that allows the expression and enumeration of different stencil program implementation variants. Combining this algebra with a compile-time performance model, we show how to automatically tune stencil programs. We use our framework to model the STELLA library and optimize kernels used by the COSMO atmospheric model on multi-core and hybrid CPU-GPU architectures. Compared to naive and expert-tuned variants, the automatically tuned kernels attain a 2.0-3.1x and a 1.0-1.8x speedup respectively.
Tobias Gysi, Tobias Grosser, Torsten Hoefler
ICS1
2015 STELLA: a domain-specific tool for structured grid methods in weather and climate models
abstract
Many high-performance computing applications solving partial differential equations (PDEs) can be attributed to the class of kernels using stencils on structured grids. Due to the disparity between floating point operation throughput and main memory bandwidth these codes typically achieve only a low fraction of peak performance. Unfortunately, stencil computation optimization techniques are often hardware dependent and lead to a significant increase in code complexity. We present a domain-specific tool, STELLA, which eases the burden of the application developer by separating the architecture dependent implementation strategy from the user-code and is targeted at multi- and manycore processors. On the example of a numerical weather prediction and regional climate model (COSMO) we demonstrate the usefulness of STELLA for a real-world production code. The dynamical core based on STELLA achieves a speedup factor of 1.8x (CPU) and 5.8x (GPU) with respect to the legacy code while reducing the complexity of the user code.
Tobias Gysi, Carlos Osuna, Oliver Fuhrer, Mauro Bianco, Thomas C. Schulthess
SC1
2014 Application Centric Energy-Efficiency Study of Distributed Multi-Core and Hybrid CPU-GPU Systems
abstract
We study the energy used by a production-level regional climate and weather simulation code on a distributed memory system with hybrid CPU-GPU nodes. The code is optimised for both processor architectures, for which we investigate both time and energy to solution. Operational constraints for time to solution can be met with both processor types, although on different numbers of nodes. Energy to solution is a factor 3 lower with GPUs, but strong scaling can be pushed to larger node counts with CPUs to minimize time to solution. Our data shows that an affine relationship exists between energy and node hours consumed by the simulation. We use this property to devise a simple and practical methodology for optimising for energy efficiency that can be applied to other applications, which we demonstrate with the HPCG benchmark. We conclude with a discussion about the relationship to the commonly-used GF/Watt metric.
Ben Cumming, Gilles Fourestey, Oliver Fuhrer, Tobias Gysi, Massimiliano Fatica, Thomas C. Schulthess
SC4