VLDB 2026 Research / reviewers in the wild / expert
Julian Höfer
dblp:304/5117
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-4904-0495ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: CeCaS Accelerator Design for Efficient Supercomputing in Automotive SystemsabstractModern vehicles integrate an increasing amount of computational functionality, driven by the growing complexity of in-vehicle applications. At the same time, automotive system architectures are becoming more centralized, requiring powerful HPC platforms at the core. These platforms must deliver the performance needed for ADAS, AI, and autonomous driving, while also meeting stringent energy efficiency and safety requirements.The CeCaS project addresses these challenges across a wide range of topics and domains of expertise, including processor design in advanced FinFET technology, the transformation of the E/E architecture, and advanced packaging for automotive supercomputing platforms. Within CeCaS, our work focuses on application-specific accelerator design to enable efficient processing of compute-intensive workloads. In this paper, we present our contributions in this area, including the design of hardware accelerators for both conventional and neuromorphic AI workloads, the development and evaluation of representative AI benchmarks, and the use of virtual platforms for early design-space exploration and hardware/software co-design. Annina Gutermann, Alexey Serdyuk, Fabian Lesniak, Julian Höfer, Hella Toto-Kiesa, Tanja Harbaum, Jürgen Becker 0001, Brian Pachideh, Sven Nitzsche, Moritz Neher, Carmen Weigelt, Jann Krausse, Victor Pazmino Betancourt, Klaus Knobloch, Lukas Groth, Andrija Neskovic, Saleh Mulhem, Mladen Berekovic |
DATE | 4 |
| 2026 | Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAIabstractArtificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps. Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl |
DATE | 3 |
| 2025 | Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-SourceabstractChip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators. Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali |
CODES+ISSS | 8 |
| 2025 | A Pixel Histogram-Based Safety Mechanism and Fault Detection Methodology for a Robust Image Signal Processor
Julian Höfer, Patrick Schmidt 0003, Hella Toto-Kiesa, Sebastian Höfer, Gregor Schewior, Dietmar Engelke, Karl-Heinz Eickel, Darius Grantz, Tanja Harbaum, Jürgen Becker 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | ZuSE-KI-Mobil: AI Chip Design Platform for Automotive and Industrial Applications
Shaown Mojumder, Simon Friedrich, Emil Matús, Matthias Lüders, Martin Friedrich, Oliver Renke, Holger Blume, Markus Kock, Gregor Schewior, Darius Grantz, Jens Benndorf, Julian Höfer, Patrick Schmidt 0003, Jürgen Becker 0001, Nael Fasfous, Pierpaolo Morì, Hans-Jörg Vögel, Samira Ahmadifarsani, Leonidas Kontopoulos, Ulf Schlichtmann, Yun-Jin Li, Gerhard P. Fettweis |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2024 | A Challenge-Based Blended Learning Approach for an Introductory Digital Circuits and Systems CourseabstractIn the early stages of university education, frontal teaching within expansive lecture halls and paper-based assignments predominate. Students often encounter theoretical concepts whose practical relevance only emerges later, if at all. This can lead to reduced student motivation, an increased risk of academic disengagement, and a tendency toward superficial learning.Our newly developed first-semester course on digital circuits and systems employs an innovative approach that combines blended learning and challenge-based learning to address these issues effectively. Throughout the semester, we introduce four challenges, seamlessly integrated with the course lectures, designed to enhance students’ comprehension of the discussed topics. Each challenge presents a concise, well-defined task, tackled by small teams using tools such as circuit simulators, and our automated toolchain allows students to witness their circuit designs in action on FPGAs later.Through this challenge-based methodology, we aim to foster individual problem-solving skills and practical expertise, which we consider to be essential assets for students during their university education and future careers. Julian Höfer, Michael Gauß, Manuela Adams, Fabian Kreß, Fabian Kempf, Christian Maximilian Karle, Tanja Harbaum, Andreas Barth 0001, Jürgen Becker 0001 |
ISCAS | 1 |
| 2023 | The ZuSE-KI-Mobil AI Accelerator SoC: Overview and a Functional Safety PerspectiveabstractZuSE-KI-Mobil (ZuKIMo) is a nationally funded research project, currently in its intermediate stage. The goal of the ZuKIMo project is to develop a new System-on-Chip (SoC) platform and corresponding ecosystem to enable efficient Artificial Intelligence (AI) applications with specific requirements. With ZuKIMo, we specifically target applications from the mobility domain, i.e. autonomous vehicles and drones. The initial ecosystem is built by a consortium consisting of seven partners from German academia and industry. We develop the SoC platform and its ecosystem around a novel AI accelerator design. The customizable accelerator is conceived from scratch to fulfill the functional and non-functional requirements derived from the ambitious use cases. A tape-out in 22 nm FDX-technology is planned in 2023. Apart from the System-on-Chip hardware design itself, the ZuKIMo ecosystem has the objective of providing software tooling for easy deployment of new use cases and hardware-CNN co-design. Furthermore, AI accelerators in safety-critical applications like our mobility use cases, necessitate the fulfillment of safety requirements. Therefore, we investigate new design methodologies for fault analysis of Deep Neural Networks (DNNs) and introduce our new redundancy mechanism for AI accelerators. Fabian Kempf, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Nael Fasfous, Alexander Frickenstein, Hans-Jörg Vögel, Simon Friedrich, Robert Wittig, Emil Matús, Gerhard P. Fettweis, Matthias Lüders, Holger Blume, Jens Benndorf, Darius Grantz, Martin Zeller, Dietmar Engelke, Karl-Heinz Eickel |
DATE | 2 |
| 2023 | ATLAS: An Approximate Time-Series LSTM Accelerator for Low-Power IoT ApplicationsabstractEnabling the use of Deep Neural Networks (DNNs) for time-series-based applications on low-power devices such as wearables opens up a wide range of new features and services. However, inference requires an enormous amount of operations to be performed by the computing platform. In addition, Long Short-Term Memory (LSTM)-based networks require memory to store the internal cell state for future calculations. In this paper, we therefore propose a hardware/software co-design based low-power LSTM hardware accelerator architecture for Internet of Things (IoT) applications called ATLAS. The design is based on approximate computing techniques to reduce the power consumption and inference latency by achieving high accuracy. Exemplary, we investigate the impact of applying our proposed architecture to a DNN for handwriting recognition. Thereby, we can show that the accuracy decreases only slightly when the inference is executed on ATLAS. The low power consumption is achieved by a minimal design requiring 173 LUTs, 67 FFs, one DSP, and one BRAM on a Xilinx FPGA. As a result, ATLAS enables the efficient use of LSTM-based DNNs in IoT devices. Fabian Kreß, Alexey Serdyuk, Micha Hiegle, Disnebio Waldmann, Tim Hotfilter, Julian Höfer, Tim Hamann, Jens Barth, Peter Kämpf, Tanja Harbaum, Jürgen Becker 0001 |
DSD | 6 |
| 2023 | SiFI-AI: A Fast and Flexible RTL Fault Simulation Framework Tailored for AI Models and AcceleratorsabstractFor AI-based systems in safety-critical domains, it is inevitable to understand the impact of random hardware faults affecting the target hardware accelerators. The high degree of data reuse makes Deep Neural Network (DNN) accelerators susceptible to significant fault propagation and hence hazardous predictions. Therefore, we present SiFI-AI, a simulation framework for fault injection in DNN accelerators. SiFI-AI proposes a hybrid simulation approach combining fast AI inference with cycle-accurate RTL simulation. Time-expensive RTL simulation is only used to accurately target registers in the hardware through condition-based fault injection. This enables to reveal vulnerable DNN layers and the related fault origin. In a resilience study with 1.5~M fault injection experiments, we analyze representative DNNs and a state-of-the-art DNN accelerator to identify vulnerable layers. The study only takes 1.15 days which is 7x faster than state-of-the-art. Our experiments show the high impact of control register faults and that narrow and deep layers are 10x more resilient compared to the wide and shallow layers of a DNN. Julian Höfer, Fabian Kempf, Tim Hotfilter, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | A Hardware-Aware Sampling Parameter Search for Efficient Probabilistic Object Detection
Julian Höfer, Tim Hotfilter, Fabian Kreß, Tanja Harbaum, Jürgen Becker 0001 |
ICVS | 1 |
| 2023 | CNNParted: An open source framework for efficient Convolutional Neural Network inference partitioning in embedded systems
Fabian Kreß, Vladimir Sidorenko, Patrick Schmidt 0003, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Tanja Harbaum, Jürgen Becker 0001 |
Comput. Networks | 4 |
| 2022 | AnaCoNGA: Analytical HW-CNN Co-Design Using Nested Genetic AlgorithmsabstractWe present AnaCoNGA, an analytical co-design methodology, which enables two genetic algorithms to evaluate the fitness of design decisions on layer-wise quantization of a neural network and hardware (HW) resource allocation. We embed a hardware architecture search (HAS) algorithm into a quantization strategy search (QSS) algorithm to evaluate the hardware design Pareto-front of each considered quantization strategy. We harness the speed and flexibility of analytical HW-modeling to enable parallel HW-CNN co-design. With this approach, the QSS is focused on seeking high-accuracy quantization strategies which are guaranteed to have efficient hardware designs at the end of the search. Through AnaCoNGA, we improve the accuracy by 2.88 p.p. with respect to a uniform 2-bit ResNet20 on CIFAR-10, and achieve a 35% and 37% improvement in latency and DRAM accesses, while reducing LUT and BRAM resources by 9% and 59% respectively, when compared to a standard edge variant of the accelerator. The nested genetic algorithm formulation also reduces the search time by 51% compared to an equivalent, sequential co-design formulation. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Julian Höfer, Anmol Singh, Naveen Shankar Nagaraja, Hans-Jörg Vögel, Nguyen Anh Vu Doan, Maurizio Martina, Jürgen Becker 0001, Walter Stechele |
DATE | 6 |
| 2022 | Hardware-aware Partitioning of Convolutional Neural Network Inference for Embedded AI ApplicationsabstractEmbedded image processing applications like multicamera-based object detection or semantic segmentation are often based on Convolutional Neural Networks (CNNs) to provide precise and reliable results. The deployment of CNNs in embedded systems, however, imposes additional constraints such as latency restrictions and limited energy consumption in the sensor platform. These requirements have to be considered during hardware/software co-design of embedded Artifical Intelligence (AI) applications. In addition, the transmission of uncompressed image data from the sensor to a central edge node requires large bandwidth on the link, which must also be taken into account during the design phase.Therefore, we present a simulation toolchain for fast evaluation of hardware-aware CNN partitioning for embedded AI applications. This approach explores an efficient workload distribution between sensor nodes and a central edge node. Neither processing all layers close to the sensor nor transmitting all uncompressed raw data to the edge node is an optimal solution for each use case. Hence, our proposed simulation toolchain evaluates power and performance metrics for each reasonable partitioning point in a CNN. In contrast to the state of the art, our approach does not only consider the neural network architecture. In the evaluation, our simulation toolchain additionally takes into account hardware components such as special accelerators and memories that are implemented in the sensor node.Exemplary, we show the simulation results for three commonly used CNNs in embedded systems. Thereby, we identify advantageous partitioning points regarding inference latency and energy consumption. With the support of the toolchain, we are able to identify three beneficial partitioning points for FCN ResNet-50 and two for GoogLeNet as well as for SqueezeNet V1.1. Fabian Kreß, Julian Höfer, Tim Hotfilter, Iris Fürst-Walter, Vladimir Sidorenko, Tanja Harbaum, Jürgen Becker 0001 |
DCOSS | 2 |
| 2021 | Binary-LoRAX: Low-Latency Runtime Adaptable XNOR Classifier for Semi-Autonomous Grasping with Prosthetic HandsabstractIntelligent, semi-autonomous prostheses take ad-vantage of combining autonomous functions and traditional myoelectric control. With the help of visual and environment sensors, intelligent prostheses achieve a level of autonomy which relieves the user from generating elaborate electromyographic (EMG) signals for grasp type and trajectory. To achieve the desired functionality, the semi-autonomous prosthesis must efficiently process the incoming environmental data at a high rate, with low power and high accuracy. In this paper, we propose Binary-LoRAX, a low-latency runtime adaptable classifier for the semi-autonomous grasping task of prosthetic hands. We offload the classification task to an efficient binary neural network accelerator which performs high-throughput XNOR operations on digital signal processing (DSP) blocks. To tailor the classifier’s performance to the current application scenario, we propose a frequency scaling approach which dynamically switches between two modes of operation, high-performance and power-saving. At high-performance, classifications are performed with a low latency of 0.45ms, high-throughput of 4999 FPS and power consumption of ∼ 2.15 W. This enables functions such as object localization and batch classification. Switching to power-saving mode, a latency of 80 ms is maintained, with up to 19% improved classifier battery-life. Our prototypes achieve a high accuracy of up to 99.82% on a 25 class problem from the YCB graspable object dataset. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Mohamed Badawy, Felix Hundhausen, Julian Höfer, Naveen Shankar Nagaraja, Christian Unger, Hans-Jörg Vögel, Jürgen Becker 0001, Tamim Asfour, Walter Stechele |
ICRA | 6 |