Christian von Elm

dblp:290/4001 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-4741-7201ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 From Microbenchmarks to LLM Inference an End-To-End Analysis on the Energy Efficiency of the Grace Hopper Superchip
Markus Velten, Christian von Elm, Lena Jurkschat, Gülçin Gedik, Sebastian Döbel, Daniel Hackenberg
ISPDC2
2026 Improving Energy Efficiency and Performance of Weather and Climate Simulations by Leveraging the Heterogeneity of Modern Systems
abstract
The increasing need for higher resolution and greater accuracy in weather forecasts and climate simulations continues to drive software and hardware developments in high-performance computing (HPC) systems. To achieve increasingly faster simulations, the adoption of hardware accelerators has proven highly effective in recent years. However, these HPC codes exhibit, in parts, divergent memory access and computation patterns. Hence, their performance benefits strongly depend on how well the software's computational characteristics align with the underlying hardware architecture. While an accelerator may be well-suited for some code regions, other architectures may achieve better performance and energy efficiency elsewhere. We therefore propose a highly heterogeneous setup for a performant and energy-efficient execution of complex HPC codes such as those used in the weather and climate domain, incorporating a variety of processor and accelerator architectures.
Julius Plehn, Christian von Elm, Pay Gießelmann, Carsten Clauss, Hendryk Bockelmann, Robert Schöne, Jan Frederik Engels
ICPE2
2023 RTASS: a RunTime Adaptable and Scalable System for Network-on-Chip-Based Architectures
abstract
In an ever-evolving digital world with complex algorithms like machine learning, we need new strategies for more flexibility to cope with the ever-changing environment. For this, runtime scalability and runtime adaptability for low-power and highly efficient hardware is a promising solution. By combining the runtime reconfiguration of FPGAs with the efficient communication of Networks-on-Chip (NoC), we are able to implement a highly scalable, high-performance, and energy-efficient computing architecture that fixed-function units and specialized static accelerators lack. In this work, we introduce a RunTime Adaptable and Scalable System for NoC-based architectures called RTASS. The hardware architecture includes a master subsystem, a network adapter, and an NoC subsystem with parametrizable routers and several various routing algorithms. Furthermore, RTASS provides a software architecture that includes advanced drivers for runtime management. The key benefit of RTASS is the ability to dynamically adjust the number of routers within the NoC at runtime based on the current application's requirements. That allows the system to support both homogeneous and inhomogeneous types of processing elements as well as regular and irregular shapes. The development of this runtime scalable and flexible architecture will establish the foundation for future highly adaptable applications such as machine learning and computer vision in the embedded computing field. We implemented and evaluated the proposed work with the Xilinx Zynq-7000 FPGA, with the possibility of porting it to other FPGAs that support runtime reconfiguration.
Najdet Charaf, Julian Haase, Adrian Kulisch, Christian von Elm, Diana Göhringer
DSD4
2023 Software-Defined CPU Modes
abstract
Our CPUs contain a compute instruction set, which regular applications use. But they also feature an intricate underworld of different CPU modes, combined with trap and exception handling to transition between these modes. These mechanisms are manifold and complex, yet the layering and functionality offered by the CPU modes is fixed. We have to take what CPU vendors provide, including potential security problems from unneeded modes. This paper explores the question, whether CPU modes could instead be defined entirely by software. We show how such a design would function and explore the advantages it enables. We believe that pushing all existing modes under a common design umbrella would enforce a cleaner structure and more control over exposed functionality. At the same time, the flexibility of software-defined modes enables interesting new use cases.
Michael Roitzsch, Till Miemietz, Christian von Elm, Nils Asmussen
HotOS3
2022 Bridging the Gap between Application Performance Analysis and System Monitoring
abstract
Performance analysis has a long history in the high-performance computing community. On the one hand, the traditional application analysis focuses on scalable yet detailed instrumentation of parallel execution. On the other hand, per-node or cluster-wide monitoring solutions are used in data center operation. However, performance anomalies resulting from the interaction between applications, background processes and the operating system, are difficult to analyze with tools that reveal only part of the issue. In this paper, we present a novel approach that covers all aspects of individual nodes. We extend a monitoring tool to combine call stack sampling, process monitoring, and syscall recording into a symbiotic view of the application execution, background activity, and the operating system.
Thomas Ilsche, Mario Bielert, Christian von Elm
CLUSTER3