EDBT 2026 Demo / reviewers in the wild / expert
Christian Heidorn
dblp:205/3213
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0002-7557-0350ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entropy Sampling-Based Neural Architecture Search for Resource-Constrained Microcontroller TargetsabstractNeural architecture search (NAS) is a popular approach for the exploration of neural network (NN) architectures. Recently proposed hardware-aware NAS techniques even take resource constraints, such as FLOP count and number of weights, into account. Still, in typical NAS search spaces, a significant portion of candidate NNs may be infeasible when it comes to satisfying tight memory (i.e., RAM and ROM) and timing constraints, particularly in the case of microcontroller targets. As evaluating each design point can be quite time-intensive, we first show how to pre-process a given design space to a reduced set of only feasible (resource constraint fulfilling) solutions, and then efficiently sampling from this set of only feasible solutions by proposing an entropy-based sampling technique and the optimization goal to maximize accuracy. We demonstrate that our approach is able to find feasible solutions with similar accuracy to other hardware-aware NAS techniques, but already after a much lower number of model evaluations, with examples taken from the MLPerf Tiny Benchmark suite. Christian Heidorn, Frank Hannig, Dominik Riedelbauch, Christoph Strohmeyer, Jürgen Teich |
DATE | 1 |
| 2024 | ALPACA: An Accelerator Chip for Nested Loop ProgramsabstractALPACA is an ASIC implementing an array of 8×8 programmable processing elements for accelerating nested loop programs. Each of them supports 32-bit as well as 8-bit floating point formats. The array is surrounded by 128 memory banks and respective control units to scan loops automatically and perform load/stores without affecting the execution time of the processed loop nest. The chip has been manufactured in 22 nm on a 10 mm2die. It achieves a peak performance of 537.6 GFLOPS @ 700 MHz and a peak energy efficiency of 270 GFLOPS/W @ 50 MHz. Dominik Walter, Marcel Brand, Christian Heidorn, Michael Witterauf, Frank Hannig, Jürgen Teich |
ISCAS | 3 |
| 2024 | Efficient Deployment of Neural Networks for Thermal Monitoring on AURIX TC3xx Microcontrollers
Christian Heidorn, Frank Hannig, Dominik Riedelbauch, Christoph Strohmeyer, Jürgen Teich |
VEHITS | 1 |
| 2021 | Hand Sign Recognition via Deep Learning on Tightly Coupled Processor ArraysabstractThe advent of deep learning has revolutionized the domain of computer vision. Convolutional neural networks (CNNs) became state-of-the-art for solving complex tasks thanks to technological advances of high-end accelerators, such as GPUs and FPGAs, combined in clusters or cloud solutions. In embedded systems, CNNs are also of great interest. However, often these devices cannot afford to offload computational-intensive workloads to the cloud due to strict energy or real-time constraints. Tightly Coupled Processor Arrays (TCPAs) are ideal architectures for accelerating nested loop programs at high energy efficiency. In this demonstrator, we show how TCPAs can meet these requirements at the edge of computing. For illustration, we designed a CNN-based hand sign recognition which is accelerated on a TCPA, implemented the TCPA prototypically as an overlay on a Xilinx Zynq System-on-a-Chip (SoC), and showcase tremendous speedups compared with the integrated ARM Cortex-A9 processor. Christian Heidorn, Dominik Walter, Yunus Emre Candir, Frank Hannig, Jürgen Teich |
FPL | 1 |
| 2020 | Design space exploration for layer-parallel execution of convolutional neural networks on CGRAsabstractIn this work, we systematically explore the design space of throughput, energy, and hardware costs for layer-parallel mappings of Convolutional Neural Networks (CNNs) onto coarse-grained reconfigurable arrays (CGRAs). We derive an analytical model that computes the required resources (processing elements) and buffer memory and thus hardware cost C to sustain a given throughput T as well as the resulting overall energy consumption E for inference. Further, we propose an efficient design space exploration (DSE) to determine the fronts of Pareto-optimal (T,E,C) solutions. This exploration helps to determine the limits of scalability of the presented tiled CGRA accelerator architectures in terms of throughput, the number of parallel layers that can be simultaneously processed, and memory requirements. Finally, we provide an evaluation of energy savings achievable on our architecture in comparison to implementations that execute sequentially a CNN layer-by-layer. In experiments, it is shown that layer-parallel processing is able to reduce energy consumption E by 3.6X, hardware cost C by 1.2X, and increase the achievable throughput T by 6.2X for MobileNet. Christian Heidorn, Frank Hannig, Jürgen Teich |
SCOPES | 1 |
| 2019 | Compilation of Dataflow Applications for Multi-Cores using Adaptive Multi-Objective OptimizationabstractState-of-the-art system synthesis techniques employ meta-heuristic optimization techniques for Design Space Exploration (DSE) to tailor application execution, e.g., defined by a dataflow graph, for a given target platform. Unfortunately, the performance evaluation of each implementation candidate is computationally very expensive, in particular on recent multi-core platforms, as this involves compilation to and extensive evaluation on the target hardware. Applying heuristics for performance evaluation on the one hand allows for a reduction of the exploration time but on the other hand may deteriorate the convergence of the optimization technique toward performance-optimal solutions with respect to the target platform. To address this problem, we propose DSE strategies that are able to dynamically trade off between (i) approximating heuristics to guide the exploration and (ii) accurate performance evaluation, i.e., compilation of the application and subsequent performance measurement on the target platform. Technically, this is achieved by introducing a set of additional, but easily computable guiding objective functions, and varying the set of objective functions that are evaluated during the DSE adaptively. One major advantage of these guiding objectives is that they are generically applicable for dataflow models without having to apply any configuration techniques to tailor their parameters to the specific use case. We show this for synthetic benchmarks as well as a real-world control application. Moreover, the experimental results demonstrate that our proposed adaptive DSE strategies clearly outperform a state-of-the-art DSE approach known from literature in terms of the quality of the gained implementations as well as exploration times. Amongst others, we show a case for a two-core implementation where after about 3 hours of exploration time one of our proposed adaptive DSE strategies already obtains a 60% higher performance value than obtained by the state-of-the-art approach. Even when the state-of-the-art approach is given a total exploration time of more than 2 weeks to optimize this value, the proposed adaptive DSE strategy features a 20% higher performance value after a total exploration time of about 4 days. Tobias Schwarzer, Joachim Falk, Martín Letras, Christian Heidorn, Stefan Wildermann, Jürgen Teich |
ACM Trans. Design Autom. Electr. Syst. | 5 |