Ludger Paehler

dblp:239/5188 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0002-7200-7637ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 61% Parallel and multicore computing · 21% High-performance computing · 18%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
automatic differentiation
1.122022
Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler Augmentation · SC 2022
Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme · SC 2021
Compilers and program optimization › automatic differentiation
reverse-mode automatic differentiation
0.512021
Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme · SC 2021
GPUs and heterogeneous computing
GPU kernel optimization
0.512021
Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme · SC 2021
Parallel and multicore computing › parallel programming models and runtimes
parallel programming frameworks
0.212022
Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler Augmentation · SC 2022
High-performance computing
scientific computing systems
0.112021
Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme · SC 2021

Methods — techniques the papers use, named apart from their topics

enzyme · 1.1LLVM · 1.1DAG-based parallelism · 1.1automatic differentiation · 1.0LLVM compiler plugin · 1.0
YearPublicationVenuePosition
2022 Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler Augmentation
abstract
Derivatives are key to numerous science, engineering, and machine learning applications. While existing tools generate derivatives of programs in a single language, modern parallel applications combine a set of frameworks and languages to leverage available performance and function in an evolving hardware landscape. We propose a scheme for differentiating arbitrary DAG-based parallelism that preserves scalability and efficiency, implemented into the LLVM-based Enzyme automatic differentiation framework. By integrating with a full-fledged compiler backend, Enzyme can differentiate numerous parallel frameworks and directly control code generation. Combined with its ability to differentiate any LLVM-based language, this flexibility permits Enzyme to leverage the compiler tool chain for parallel and differentiation-specitic optimizations. We differentiate nine distinct versions of the LULESH and miniBUDE applications, written in different programming languages (C++, Julia) and parallel frameworks (OpenMP, MPI, RAJA, Julia tasks, MPI.jl), demonstrating similar scalability to the original program. On benchmarks with 64 threads or nodes, we find a differentiation overhead of 3.4–6.8× on C++ and 5.4–12.5× on Julia.
William S. Moses, Sri Hari Krishna Narayanan, Ludger Paehler, Valentin Churavy, Michel Schanen, Jan Hückelheim, Johannes Doerfert, Paul D. Hovland
SC3
2021 Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme
abstract
Computing derivatives is key to many algorithms in scientific computing and machine learning such as optimization, uncertainty quantification, and stability analysis. Enzyme is a LLVM compiler plugin that performs reverse-mode automatic differentiation (AD) and thus generates high performance gradients of programs in languages including C/C++, Fortran, Julia, and Rust. Prior to this work, Enzyme and other AD tools were not capable of generating gradients of GPU kernels. Our paper presents a combination of novel techniques that make Enzyme the first fully automatic reversemode AD tool to generate gradients of GPU kernels. Since unlike other tools Enzyme performs automatic differentiation within a general-purpose compiler, we are able to introduce several novel GPU and AD-specific optimizations. To show the generality and efficiency of our approach, we compute gradients of five GPU-based HPC applications, executed on NVIDIA and AMD GPUs. All benchmarks run within an order of magnitude of the original program's execution time. Without GPU and AD-specific optimizations, gradients of GPU kernels either fail to run from a lack of resources or have infeasible overhead. Finally, we demonstrate that increasing the problem size by either increasing the number of threads or increasing the work per thread, does not substantially impact the overhead from differentiation.
William S. Moses, Valentin Churavy, Ludger Paehler, Jan Hückelheim, Sri Hari Krishna Narayanan, Michel Schanen, Johannes Doerfert
SC3