VLDB 2026 Research / reviewers in the wild / expert
Nicolas R. Gauger
dblp:126/0474
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2027
0000-0002-5863-7384ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Performance benchmarking of Tensor Trains for quantum-inspired homogenization on TPU, GPU, and CPU architecturesabstractRecent advances in high-resolution CT-imaging technology are creating a new class of ultra-high resolved microstructural datasets that challenge the limits of traditional homogenization approaches. While state-of-the-art FFT-based homogenization techniques remain effective for moderate datasets, their memory footprint and computational cost grow rapidly with increasing resolution, making them progressively inefficient for industrial-scale problems. To address these challenges, the recently developed Superfast-Fourier Transform (SFFT)-based homogenization algorithm leverages the memory-efficient low-rank representations of Tensor Trains (TTs), which reduce the storage and computational requirements of large-scale homogenization problems. Developed for CPU usage, SFFT-based Homogenization efficiently handles high-resolution datasets, assuming the underlying data is well-behaved.In this work, we investigate the performance of fundamental TT operations on modern hardware accelerators using the JAX framework. A benchmarking study across CPUs, GPUs, and TPUs evaluates execution times and computational efficiency, highlighting the strengths and limitations of TT operations on different architectures and motivating future hybrid approaches. Building on these insights, we adapt the SFFT-based homogenization algorithm for accelerator execution, enabling homogenization at high resolutions ranging from 300 million to 70 billion grid points, which are infeasible for the best available GPU-based FFT reference implementation. While the observed scaling behavior is geometry-dependent, the results demonstrate the potential of accelerator-based quantum-inspired homogenization for high-performance multiscale simulations. Sascha Hauck, Matthias Kabel, Nicolas R. Gauger |
Future Gener. Comput. Syst. | 3 |
| 2025 | Forward-Mode Automatic Differentiation of Compiled ProgramsabstractAlgorithmic differentiation (AD) is a set of techniques that provide partial derivatives of computer-implemented functions. Such functions can be supplied to state-of-the-art AD tools via their source code , or via intermediate representations produced while compiling their source code. We present the novel AD tool Derivgrind, which augments the machine code of compiled programs with forward-mode AD logic. Derivgrind leverages the Valgrind instrumentation framework for structured access to the machine code, and a shadow memory tool to store dot values. Access to the source code is required at most for the files in which input and output variables are defined. Derivgrind’s versatility mainly comes at the price of reduced run-time performance. According to our extensive regression test suite, Derivgrind produces correct results on GCC- and Clang-compiled programs, including a Python interpreter, with a small number of exceptions. We provide a list of “bit-tricks” that Derivgrind does not handle correctly, some of which actually appear in highly optimized math libraries. As long as differentiating those is avoided, Derivgrind enables black-box forward-mode AD for an unprecedentedly wide range of cross-language software with little integration efforts. Max Aehle, Johannes Blühdorn, Max Sagebaum, Nicolas R. Gauger |
ACM Trans. Math. Softw. | 4 |
| 2023 | Towards Neural Charged Particle Tracking in Digital Tracking Calorimeters With Reinforcement LearningabstractWe propose a novel technique for reconstructing charged particles in digital tracking calorimeters using reinforcement learning aiming to benefit from the rapid progress and success of neural network architectures without the dependency on simulated or manually-labeled data. Here we optimize by trial-and-error a behavior policy acting as an approximation to the full combinatorial optimization problem, maximizing the physical plausibility of sampled trajectories. In modern processing pipelines used in high energy physics and related applications, tracking plays an essential role allowing to identify and follow charged particle trajectories traversing particle detectors. Due to the high multiplicity of charged particles and their physical interactions, randomly deflecting the particles, the reconstruction is a challenging undertaking, requiring fast, accurate and robust algorithms. Our approach works on graph-structured data, capturing track hypotheses through edge connections between particles in the detector layers. We demonstrate in a comprehensive study on simulated data for a particle detector used for proton computed tomography, the high potential as well as the competitiveness of our approach compared to a heuristic search algorithm and a model trained on ground truth. Finally, we point out limitations of our approach, guiding towards a robust foundation for further development of reinforcement learning based tracking. Tobias Kortus, Ralf Keidel, Nicolas R. Gauger |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Event-Based Automatic Differentiation of OpenMP with OpDiLibabstractWe present the new software OpDiLib, a universal add-on for classical operator overloading AD tools that enables the automatic differentiation (AD) of OpenMP parallelized code. With it, we establish support for OpenMP features in a reverse mode operator overloading AD tool to an extent that was previously only reported on in source transformation tools. We achieve this with an event-based implementation ansatz that is unprecedented in AD. Combined with modern OpenMP features around OMPT, we demonstrate how it can be used to achieve differentiation without any additional modifications of the source code; neither do we impose a priori restrictions on the data access patterns, which makes OpDiLib highly applicable. For further performance optimizations, restrictions like atomic updates on adjoint variables can be lifted in a fine-grained manner. OpDiLib can also be applied in a semi-automatic fashion via a macro interface, which supports compilers that do not implement OMPT. We demonstrate the applicability of OpDiLib for a pure operator overloading approach in a hybrid parallel environment. We quantify the cost of atomic updates on adjoint variables and showcase the speedup and scaling that can be achieved with the different configurations of OpDiLib in both the forward and the reverse pass. Johannes Blühdorn, Max Sagebaum, Nicolas R. Gauger |
ACM Trans. Math. Softw. | 3 |
| 2019 | High-Performance Derivative Computations using CoDiPackabstractThere are several AD tools available that all implement different strategies for the reverse mode of AD. The most common strategies are primal value taping (implemented e.g. by ADOL-C) and Jacobian taping (implemented e.g. by Adept and dco/c++). Particulary for Jacobian taping, recent advances using expression templates make it very attractive for large scale software. However, the current implementations are either closed source or miss essential features and flexibility. Therefore, we present the new AD tool CoDiPack (Code Differentiation Package) in this paper. It is specifically designed for minimal memory consumption and optimal runtime, such that it can be used for the differentiation of large scale software. An essential part of the design of CoDiPack is the modular layout and the recursive data structures which not only allow the efficient implementation of the Jacobian taping approach but will also enable other approaches like the primal value taping or new research ideas. We will finally present the performance values of CoDiPack on a generic PDE example and on the SU2 code. Max Sagebaum, Tim Albring, Nicolas R. Gauger |
ACM Trans. Math. Softw. | 3 |