VLDB 2026 Research / reviewers in the wild / expert
Fabian Kreß
dblp:296/0006
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-1700-5778ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DSEParted: Co-Optimization of Embedded NPU Architectures and Neural Network PartitioningabstractConvolutional Neural Networks (CNNs) have become an essential tool in the domain of vision processing. However, dedicated accelerators are needed for energy-efficient execution of these networks, especially for embedded devices with tight energy constraints. Integrating multiple of these accelerators via chiplets promises a way to scale up the performance of these emerging systems by partitioning a neural network across multiple accelerators. This approach enables the execution of different layers on an accelerator with the best-suited dataflow. However, partitioning a neural network is a non-trivial task, especially when different accelerator architectures must be considered. In this paper, we propose our framework DSEParted, which automates the co-design of network partitioning and hardware architecture optimization. It employs a hierarchical optimization approach to gradually reduce the number of design candidates until an optimal system configuration is found for a partitioned computation of a neural network. We demonstrate that our framework can design a system that, for GoogLeNet, reduces latency by 22.5% when optimizing for latency. In addition, when optimizing for energy, it reduces the system area by 7.9%, with no impact on latency or energy compared to a baseline system. Further, we show that our partitioning-aware pruning strategy can reduce the EDP of the system by up to 49.7% in the case of ResNeXt-50, compared to a strategy that only optimizes the accelerators individually. Through the provided information, designers receive feedback on the efficiency of the full system at an early development stage. Our work is available open source1.1https://github.com/itiv-kit/cnn-parted Patrick Schmidt 0003, Fabian Kreß, Alexey Serdyuk, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001 |
DSD | 2 |
| 2025 | General Compilation and Mixed-Precision Partitioning: A Combined Approach for Adaptive On-Device Learning
Iuliia Topko, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | KIHT: Kaligo-Based Intelligent Handwriting TeacherabstractKaligo-based Intelligent Handwriting Teacher (KIHT) is a bi-nationally funded research project. The aim of this joint project is to develop an intelligent learning device for automated handwriting, composed of existing components, which can be made available to as many students as possible. With KIHT, we specifically address the challenging task of using inertial sensors to retrace the trajectory of a pen without relying on external reference systems. The nearly unlimited freedom to let the pen glide over the paper has not yet provided a satisfactory solution to this challenge in the state-of-the-art methods, even with sophisticated algorithms and AI approaches. The final phase of the project is now being launched and together with partners from industry and academia, we are taking a holistic approach by considering the entire chain of components, from the pen to the embedded processing system, the algorithms and the app. Tanja Harbaum, Alexey Serdyuk, Fabian Kreß, Tim Hamann, Jens Barth, Peter Kämpf, Florent Imbert, Yann Soullard, Romain Tavenard, Éric Anquetil, Jessica Delahaie |
DATE | 3 |
| 2024 | A Challenge-Based Blended Learning Approach for an Introductory Digital Circuits and Systems CourseabstractIn the early stages of university education, frontal teaching within expansive lecture halls and paper-based assignments predominate. Students often encounter theoretical concepts whose practical relevance only emerges later, if at all. This can lead to reduced student motivation, an increased risk of academic disengagement, and a tendency toward superficial learning.Our newly developed first-semester course on digital circuits and systems employs an innovative approach that combines blended learning and challenge-based learning to address these issues effectively. Throughout the semester, we introduce four challenges, seamlessly integrated with the course lectures, designed to enhance students’ comprehension of the discussed topics. Each challenge presents a concise, well-defined task, tackled by small teams using tools such as circuit simulators, and our automated toolchain allows students to witness their circuit designs in action on FPGAs later.Through this challenge-based methodology, we aim to foster individual problem-solving skills and practical expertise, which we consider to be essential assets for students during their university education and future careers. Julian Höfer, Michael Gauß, Manuela Adams, Fabian Kreß, Fabian Kempf, Christian Maximilian Karle, Tanja Harbaum, Andreas Barth 0001, Jürgen Becker 0001 |
ISCAS | 4 |
| 2023 | ATLAS: An Approximate Time-Series LSTM Accelerator for Low-Power IoT ApplicationsabstractEnabling the use of Deep Neural Networks (DNNs) for time-series-based applications on low-power devices such as wearables opens up a wide range of new features and services. However, inference requires an enormous amount of operations to be performed by the computing platform. In addition, Long Short-Term Memory (LSTM)-based networks require memory to store the internal cell state for future calculations. In this paper, we therefore propose a hardware/software co-design based low-power LSTM hardware accelerator architecture for Internet of Things (IoT) applications called ATLAS. The design is based on approximate computing techniques to reduce the power consumption and inference latency by achieving high accuracy. Exemplary, we investigate the impact of applying our proposed architecture to a DNN for handwriting recognition. Thereby, we can show that the accuracy decreases only slightly when the inference is executed on ATLAS. The low power consumption is achieved by a minimal design requiring 173 LUTs, 67 FFs, one DSP, and one BRAM on a Xilinx FPGA. As a result, ATLAS enables the efficient use of LSTM-based DNNs in IoT devices. Fabian Kreß, Alexey Serdyuk, Micha Hiegle, Disnebio Waldmann, Tim Hotfilter, Julian Höfer, Tim Hamann, Jens Barth, Peter Kämpf, Tanja Harbaum, Jürgen Becker 0001 |
DSD | 1 |
| 2023 | SiFI-AI: A Fast and Flexible RTL Fault Simulation Framework Tailored for AI Models and AcceleratorsabstractFor AI-based systems in safety-critical domains, it is inevitable to understand the impact of random hardware faults affecting the target hardware accelerators. The high degree of data reuse makes Deep Neural Network (DNN) accelerators susceptible to significant fault propagation and hence hazardous predictions. Therefore, we present SiFI-AI, a simulation framework for fault injection in DNN accelerators. SiFI-AI proposes a hybrid simulation approach combining fast AI inference with cycle-accurate RTL simulation. Time-expensive RTL simulation is only used to accurately target registers in the hardware through condition-based fault injection. This enables to reveal vulnerable DNN layers and the related fault origin. In a resilience study with 1.5~M fault injection experiments, we analyze representative DNNs and a state-of-the-art DNN accelerator to identify vulnerable layers. The study only takes 1.15 days which is 7x faster than state-of-the-art. Our experiments show the high impact of control register faults and that narrow and deep layers are 10x more resilient compared to the wide and shallow layers of a DNN. Julian Höfer, Fabian Kempf, Tim Hotfilter, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | A Hardware-Aware Sampling Parameter Search for Efficient Probabilistic Object Detection
Julian Höfer, Tim Hotfilter, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001 |
ICVS | 3 |
| 2023 | CNNParted: An open source framework for efficient Convolutional Neural Network inference partitioning in embedded systems
Fabian Kreß, Vladimir Sidorenko, Patrick Schmidt 0003, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Tanja Harbaum, Jürgen Becker 0001 |
Comput. Networks | 1 |
| 2022 | Towards Reconfigurable Accelerators in HPC: Designing a Multipurpose eFPGA Tile for Heterogeneous SoCsabstractThe goal of modern high performance computing platforms is to combine low power consumption and high throughput. Within the European Processor Initiative (EPI), such an SoC platform to meet the novel exascale requirements is built and investigated. As part of this project, we introduce an embedded Field Programmable Gate Array (eFPGA), adding flexibility to accelerate various workloads. In this article, we show our approach to design the eFPGA tile that supports the EPI SoC. While eFPGAs are inherently reconfigurable, their initial design has to be determined for tape-out. The design space of the eFPGA is explored and evaluated with different configurations of two HPC workloads, covering control and dataflow heavy applications. As a result, we present a well-balanced eFPGA design that can host several use cases and potential future ones by allocating 1% of the total EPI SoC area. Finally, our simulation results of the architectures on the eFPGA show great performance improvements over their software counterparts. Tim Hotfilter, Fabian Kreß, Fabian Kempf, Jürgen Becker 0001, Juan Miguel De Haro Ruiz, Daniel Jiménez-González, Miquel Moretó, Carlos Álvarez 0001, Jesús Labarta, Imen Baili |
DATE | 2 |
| 2022 | Hardware-aware Partitioning of Convolutional Neural Network Inference for Embedded AI ApplicationsabstractEmbedded image processing applications like multicamera-based object detection or semantic segmentation are often based on Convolutional Neural Networks (CNNs) to provide precise and reliable results. The deployment of CNNs in embedded systems, however, imposes additional constraints such as latency restrictions and limited energy consumption in the sensor platform. These requirements have to be considered during hardware/software co-design of embedded Artifical Intelligence (AI) applications. In addition, the transmission of uncompressed image data from the sensor to a central edge node requires large bandwidth on the link, which must also be taken into account during the design phase.Therefore, we present a simulation toolchain for fast evaluation of hardware-aware CNN partitioning for embedded AI applications. This approach explores an efficient workload distribution between sensor nodes and a central edge node. Neither processing all layers close to the sensor nor transmitting all uncompressed raw data to the edge node is an optimal solution for each use case. Hence, our proposed simulation toolchain evaluates power and performance metrics for each reasonable partitioning point in a CNN. In contrast to the state of the art, our approach does not only consider the neural network architecture. In the evaluation, our simulation toolchain additionally takes into account hardware components such as special accelerators and memories that are implemented in the sensor node.Exemplary, we show the simulation results for three commonly used CNNs in embedded systems. Thereby, we identify advantageous partitioning points regarding inference latency and energy consumption. With the support of the toolchain, we are able to identify three beneficial partitioning points for FCN ResNet-50 and two for GoogLeNet as well as for SqueezeNet V1.1. Fabian Kreß, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Vladimir Sidorenko, Tanja Harbaum, Jürgen Becker 0001 |
DCOSS | 1 |