Christoph Gerum

dblp:161/0144 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1715-567XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Constrained NAS via Symbolic Expressions in Declarative Hierarchical Search Spaces
abstract
Neural Architecture Search (NAS) automates the design of deep neural networks (DNNs), but the design of the search space remains crucial: manually designed spaces require significant engineering effort, while overly flexible designs often lead to invalid or inefficient architectures. This paper introduces a novel NAS search space design that centers on symbolic constraint modeling, enabling fine-grained parametrization while ensuring architecture validity and resource efficiency. By representing parameter dependencies as symbolic expressions, our method supports automatic resolution of interdependent attributes and the specification of hard constraints, such as limits on parameter count or MAC operations, directly within the search space. This mechanism allows efficient exploration of valid architectures under strict deployment budgets. The search space itself is constructed declaratively using hierarchical, composable topology patterns, drawing from common DNN motifs and enabling intuitive and scalable definition. We demonstrate the effectiveness of our approach through evolutionary NAS under multiple resource constraints, showing that symbolic constraint enforcement improves search efficiency and robustness without sacrificing accuracy.
Moritz Reiber, Christoph Gerum, Oliver Bringmann 0001
ASP-DAC2
2024 Energy-Efficient Seizure Detection Suitable for Low-Power Applications
abstract
Epilepsy is the most common, chronic, neurological disease worldwide and is typically accompanied by reoccurring seizures. Neuro implants can be used for effective treatment by suppressing an upcoming seizure upon detection. Due to the restricted size and limited battery lifetime of those medical devices, the employed approach also needs to be limited in size and have low energy requirements. We present an energy-efficient seizure detection approach involving a TC-ResNet and time-series analysis which is suitable for low-power edge devices. The presented approach allows for accurate seizure detection without preceding feature extraction while considering the stringent hardware requirements of neural implants. The approach is validated using the CHB-MIT Scalp EEG Database with a 32-bit floating point model and a hardware suitable 4-bit fixed point model. The presented method achieves an accuracy of 95.28%, a sensitivity of 92.34% and an AUC score of 0.9384 on this dataset with 4-bit fixed point representation. Furthermore, the power consumption of the model is measured with the low-power AI accelerator UltraTrail, which only requires 495nW on average. Due to this low-power consumption this classification approach is suitable for real-time seizure detection on low-power wearable devices such as neural implants.
Julia Werner, Bhavya Kohli, Paul Palomero Bernardo, Christoph Gerum, Oliver Bringmann 0001
IJCNN4
2024 GOURD: Tensorizing Streaming Applications to Generate Multi-Instance Compute Platforms
abstract
In this article, we rethink the dataflow processing paradigm to a higher level of abstraction to automate the generation of multi-instance compute and memory platforms with interfaces to I/O devices (sensors and actuators). Since the different compute instances (NPUs, CPUs, DSPs, etc.) and I/O devices do not necessarily have compatible interfaces on a dataflow level, an automated translation is required. However, in multidimensional dataflow scenarios, it becomes inherently difficult to reason about buffer sizes and iteration order without knowing the shape of the data access pattern (DAP) that the dataflow follows. To capture this shape and the platform composition, we define a domain-specific representation (DSR) and devise a toolchain to generate a synthesizable platform, including appropriate streaming buffers for platform-specific tensorization of the data between incompatible interfaces. This allows platforms, such as sensor edge AI devices, to be easily specified by simply focusing on the shape of the data provided by the sensors and transmitted among compute units, giving the ability to evaluate and generate different dataflow design alternatives with significantly reduced design time.
Patrick Schmid, Paul Palomero Bernardo, Christoph Gerum, Oliver Bringmann 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Enhancing Robustness of LiDAR-Based Perception in Adverse Weather using Point Cloud Augmentations
abstract
LiDAR-based perception systems have become widely adopted in autonomous vehicles. However, their performance can be severely degraded in adverse weather conditions, such as rain, snow or fog. To address this challenge, we propose a method for improving the robustness of LiDAR-based perception in adverse weather, using data augmentation techniques on point clouds. We use novel as well as established data augmentation techniques, such as realistic weather simulations, to provide a wide variety of training data for LiDAR-based object detectors. The performance of the state-of-the-art detector Voxel R-CNN using the proposed augmentation techniques is evaluated on a data set of real-world point clouds collected in adverse weather conditions. The achieved improvements in average precision (AP) are 4.00 p.p. in fog, 3.35 p.p. in snow, and 4.87 p.p. in rain at moderate difficulty. Our results suggest that data augmentations on point clouds are an effective way to improve the robustness of LiDAR-based object detection in adverse weather.
Sven Teufel, Jörg Gamerdinger, Georg Volk, Christoph Gerum, Oliver Bringmann 0001
IV4
2022 Hardware Accelerator and Neural Network Co-Optimization for Ultra-Low-Power Audio Processing Devices
abstract
The increasing spread of artificial neural networks does not stop at ultralow-power edge devices. However, these very often have high computational demand and require specialized hardware accelerators to ensure the design meets power and performance constraints. The manual optimization of neural networks along with the corresponding hardware accelerators can be very challenging. This paper presents HANNAH (Hardware Accelerator and Neural Network seArcH), a framework for automated and combined hardware/software co-design of deep neural networks and hardware accelerators for resource and power-constrained edge devices. The optimization approach uses an evolution-based search algorithm, a neural network template technique and analytical KPI models for the configurable UltraTrail hardware accelerator template in order to find an optimized neural network and accelerator configuration. We demonstrate that HANNAH can find suitable neural networks with minimized power consumption and high accuracy for different audio classification tasks such as single-class wake word detection, multi-class keyword detection and voice activity detection, which are superior to the related work.
Christoph Gerum, Adrian Frischknecht, Tobias Hald, Paul Palomero Bernardo, Konstantin Lübeck, Oliver Bringmann 0001
DSD1
2021 Behavior of Keyword Spotting Networks Under Noisy Conditions
Anwesh Mohanty, Adrian Frischknecht, Christoph Gerum, Oliver Bringmann 0001
ICANN (1)3
2021 Dynamic Range and Complexity Optimization of Mixed-Signal Machine Learning Systems
abstract
Audio processing had been in demand throughout the electronic era. Recent advances in neural networks increased the demand on audio processing for speech recognition applications. In this work, a rigorous study on the dynamic range and system complexity optimization is presented for a mixed-signal keyword spotting system. The proposed system consists of an analog feature extractor and a neural network based keyword classifier. The results showed that with the proposed method, more than an order of magnitude power saving can be achieved in the analog feature extraction compared to the digital state-of-the-art counterpart.
Naci Pekcokguler, Dominique Morche, Adrian Frischknecht, Christoph Gerum, Andreas Peter Burg, Catherine Dehollain
ISCAS4
2020 UltraTrail: A Configurable Ultralow-Power TC-ResNet AI Accelerator for Efficient Keyword Spotting
abstract
Recent advances in machine learning show the superior behavior of temporal convolutional networks (TCNs) and especially their combination with residual networks (TC-ResNet) for intelligent sensor signal processing in comparison to classical CNNs and LSTMs. In this article, we propose UltraTrail, a configurable, ultralow-power TC-ResNet AI accelerator for sensor signal processing and its application to efficient keyword spotting (KWS). Following a strict hardware/model co-design approach, we have derived an optimized low-power hardware architecture for generalized TC-ResNet topologies consisting of a configurable array of processing elements and a distributed memory with dynamic content reallocation. We additionally extend the network with conditional computing to reduce the number of operations during inference and to provide the possibility for power-gating. The final accelerator implementation in Globalfoundries' 22FDX technology achieves a power consumption of 8.2 μW for the task of always-on KWS meeting the real-time requirement of 100 ms per inference with an accuracy of 93% on the Google Speech Command Dataset.
Paul Palomero Bernardo, Christoph Gerum, Adrian Frischknecht, Konstantin Lübeck, Oliver Bringmann 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Systematic RISC-V based Firmware Design⋆
abstract
Small embedded devices are highly specialized plat forms that integrate several peripherals alongside the CPU core. Embedded devices extensively rely on Firmware (FW) to control and access the peripherals as well as other important functionality. This poses challenges to FW development since the FW must be adapted to each specific device configuration. Besides ensuring functional correctness to avoid errors and security vulnerabilities, an important design factor today is the control and adaptivity of a system with respect to non-functional properties, like for example application-specific timing budgets. Furthermore, optimizations of the FW and HW/SW interface play a very important role due to the tight resource constraints of small embedded devices. To satisfy these requirements new FW design methods are needed targeting FW generation, FW verification and FW optimization.This paper presents such new methods to enable an early, efficient and systematic FW design taking the underlying HW architecture into account. We use the RISC-V Instruction Set Architecture (ISA) as a case study to demonstrate our methods.
Vladimir Herdt, Daniel Große, Rolf Drechsler, Christoph Gerum, Alexander Louis-Ferdinand Jung, Joscha Benz, Oliver Bringmann 0001, Michael Schwarz 0010, Dominik Stoffel, Wolfgang Kunz
FDL4
2018 Advancing source-level timing simulation using loop acceleration
abstract
Source-level timing simulation (STLS) is an important technique for early examination of timing behavior, as it is very fast and accurate. A factor occasionally more important than precision is simulation speed, especially in design space exploration or very early phases of development. Additionally, practices like rapid prototyping also benefit from high-performance timing simulation. Therefore, we propose to further reduce simulation run-time by utilizing a method called loop acceleration. Accelerating a loop in the context of SLTS means deriving the timing of a loop prior to simulation to increase simulation speed of that loop. We integrated this technique in our SLTS framework and conducted an comprehensive evaluation using the Malardalen benchmark suite. We were able to reduce simulation time by up to 43% of the original time, while the introduced accuracy loss did not exceed 8 percentage points.
Joscha Benz, Christoph Gerum, Oliver Bringmann 0001
DATE2
2017 Context-sensitive timing automata for fast source level simulation
abstract
We present a novel technique for efficient source level timing simulation of embedded software execution on a target platform. In contrast to existing approaches, the proposed technique can accurately approximate time without requiring a dynamic cache model. Thereby the dramatic reduction in simulation performance inherent to dynamic cache modeling is avoided. Consequently, our approach enables an exploitation of the performance potential of source level simulation for complex microarchitectures that include caches. Our approach is based on recent advances in context-sensitive binary level timing simulation. However, a direct application of the binary level approach to source level simulation reduces simulation performance similarly to dynamic cache modeling. To overcome this performance limitation, we contribute a novel pushdown automaton based simulation technique. The proposed context-sensitive timing automata enable an efficient evaluation of complex simulation logic with little overhead. Experimental results show that the proposed technique provides a speed up of an order of magnitude compared to existing context selection techniques and simple source level cache models. Simulation performance is similar to a state of the art accelerated cache simulation. The accelerated simulation is only applicable in specific circumstances, whereas the proposed approach does not suffer this limitation.
Sebastian Ottlik, Christoph Gerum, Alexander Viehl, Wolfgang Rosenstiel, Oliver Bringmann 0001
DATE2
2015 Source level performance simulation of GPU cores
Christoph Gerum, Oliver Bringmann 0001, Wolfgang Rosenstiel
DATE1