VLDB 2026 Research / reviewers in the wild / expert
Oliver Fuhrer
dblp:144/2155
· DBLP profile ↗
5ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0002-0682-1374ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 88% GPUs and heterogeneous computing · 6% Energy-efficient computing · 6% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 76% Program synthesis and code generation · 14% Programming languages and type systems · 11% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › large-scale simulation
climate and weather simulation |
0.8 | 2 | 2022 | Productive Performance Engineering for Weather and Climate Modeling with Python · SC 2022 Application Centric Energy-Efficiency Study of Distributed Multi-Core and Hybrid CPU-GPU Systems · SC 2014 |
High-performance computing
scientific computing systems |
0.7 | 2 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
High-performance computing
scientific computing |
0.6 | 1 | 2022 | Productive Performance Engineering for Weather and Climate Modeling with Python · SC 2022 |
Compilers and program optimization
compiler infrastructure |
0.5 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
Compilers and program optimization › compiler infrastructure
MLIR |
0.5 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
High-performance computing › scientific computing systems
climate modeling |
0.5 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
Program synthesis and code generation
domain-specific code generation |
0.2 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
Compilers and program optimization › loop optimization
stencil computation optimization |
0.2 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
Programming languages and type systems
domain-specific languages |
0.2 | 1 | 2022 | Productive Performance Engineering for Weather and Climate Modeling with Python · SC 2022 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
High-performance computing
stencil computation |
0.1 | 1 | 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate Simulation · ACM Trans. Archit. Code Optim. 2021 |
Environmental and earth informatics › atmospheric modeling
weather and climate modeling |
0.1 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
High-performance computing
many-core acceleration |
0.1 | 1 | 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate models · SC 2015 |
Methods — techniques the papers use, named apart from their topics
MLIR · 1.0LLVM · 1.0domain-specific language · 0.7architecture-dependent code generation · 0.7strong scaling · 0.2energy measurement · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Productive Performance Engineering for Weather and Climate Modeling with PythonabstractEarth system models are developed with a tight coupling to target hardware, often containing specialized code predicated on processor characteristics. This coupling stems from using imperative languages that hard-code computation schedules and layout. We present a detailed account of optimizing the Finite Volume Cubed-Sphere Dynamical Core (FV3), improving productivity and performance. By using a declarative Python-embedded stencil domain-specific language and data-centric optimization, we abstract hardware-specific details and define a semi-automated workflow for analyzing and optimizing weather and climate applications. The workflow utilizes both local and full-program optimization, as well as user-guided fine-tuning. To prune the infeasible global optimization space, we automatically utilize repeating code motifs via a novel transfer tuning approach. On the Piz Daint supercomputer, we scale to 2,400 GPUs, achieving speedups of up to 3.92× over the tuned production implementation at a fraction of the original code. Tal Ben-Nun, Linus Groner, Florian Deconinck, Tobias Wicky, Eddie Davis, Johann Dahm, Oliver Elbert, Rhea George, Jeremy McGibbon, Lukas Trümper, Elynn Wu, Oliver Fuhrer, Thomas C. Schulthess, Torsten Hoefler |
SC | 12 |
| 2021 | Domain-Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU-accelerated Climate SimulationabstractMost compilers have a single core intermediate representation (IR) (e.g., LLVM) sometimes complemented with vaguely defined IR-like data structures. This IR is commonly low-level and close to machine instructions. As a result, optimizations relying on domain-specific information are either not possible or require complex analysis to recover the missing information. In contrast, multi-level rewriting instantiates a hierarchy of dialects (IRs), lowers programs level-by-level, and performs code transformations at the most suitable level. We demonstrate the effectiveness of this approach for the weather and climate domain. In particular, we develop a prototype compiler and design stencil- and GPU-specific dialects based on a set of newly introduced design principles. We find that two domain-specific optimizations (500 lines of code) realized on top of LLVM’s extensible MLIR compiler infrastructure suffice to outperform state-of-the-art solutions. In essence, multi-level rewriting promises to herald the age of specialized compilers composed from domain- and target-specific dialects implemented on top of a shared infrastructure. Tobias Gysi, Oleksandr Zinenko, Stephan Herhut, Eddie Davis, Tobias Wicky, Oliver Fuhrer, Torsten Hoefler, Tobias Grosser |
ACM Trans. Archit. Code Optim. | 7 |
| 2015 | STELLA: a domain-specific tool for structured grid methods in weather and climate modelsabstractMany high-performance computing applications solving partial differential equations (PDEs) can be attributed to the class of kernels using stencils on structured grids. Due to the disparity between floating point operation throughput and main memory bandwidth these codes typically achieve only a low fraction of peak performance. Unfortunately, stencil computation optimization techniques are often hardware dependent and lead to a significant increase in code complexity. We present a domain-specific tool, STELLA, which eases the burden of the application developer by separating the architecture dependent implementation strategy from the user-code and is targeted at multi- and manycore processors. On the example of a numerical weather prediction and regional climate model (COSMO) we demonstrate the usefulness of STELLA for a real-world production code. The dynamical core based on STELLA achieves a speedup factor of 1.8x (CPU) and 5.8x (GPU) with respect to the legacy code while reducing the complexity of the user code. Tobias Gysi, Carlos Osuna, Oliver Fuhrer, Mauro Bianco, Thomas C. Schulthess |
SC | 3 |
| 2014 | Designing Bit-Reproducible Portable High-Performance ApplicationsabstractBit-reproducibility has many advantages in the context of high-performance computing. Besides simplifying and making more accurate the process of debugging and testing the code, it can allow the deployment of applications on heterogeneous systems, maintaining the consistency of the computations. In this work we analyze the basic operations performed by scientific applications and identify the possible sources of non-reproducibility. In particular, we consider the tasks of evaluating transcendental functions and performing reductions using non-associative operators. We present a set of techniques to achieve reproducibility and we propose improvements over existing algorithms to perform reproducible computations in a portable way, at the same time obtaining good performance and accuracy. By applying these techniques to more complex tasks we show that bit-reproducibility can be achieved on a broad range of scientific applications. Andrea Arteaga, Oliver Fuhrer, Torsten Hoefler |
IPDPS | 2 |
| 2014 | Application Centric Energy-Efficiency Study of Distributed Multi-Core and Hybrid CPU-GPU SystemsabstractWe study the energy used by a production-level regional climate and weather simulation code on a distributed memory system with hybrid CPU-GPU nodes. The code is optimised for both processor architectures, for which we investigate both time and energy to solution. Operational constraints for time to solution can be met with both processor types, although on different numbers of nodes. Energy to solution is a factor 3 lower with GPUs, but strong scaling can be pushed to larger node counts with CPUs to minimize time to solution. Our data shows that an affine relationship exists between energy and node hours consumed by the simulation. We use this property to devise a simple and practical methodology for optimising for energy efficiency that can be applied to other applications, which we demonstrate with the HPCG benchmark. We conclude with a discussion about the relationship to the commonly-used GF/Watt metric. Ben Cumming, Gilles Fourestey, Oliver Fuhrer, Tobias Gysi, Massimiliano Fatica, Thomas C. Schulthess |
SC | 3 |