Marc Reichenbach

dblp:49/7217 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-9687-6247ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Combining HLS and Computation Coding for Balanced Matrix-Vector Engines on FPGA
abstract
Limited number of DSP48 blocks versus abundant LUTs constrains the throughput of FPGA-based constant matrix–vector multiplications (MVM). We exploit Linear Computation-Coding (LCC) to build multiplier-less, deeply pipelined adder graphs on LUTs and combine it with a DSP-mapped HLS datapath via a pragmatic column-split partition. This hybrid mapping achieves balanced DSP and LUT utilization without sacrificing throughput, resulting in a scalable, high-performance MVM on FPGAs.
Sayanti Pal, Alexander Lehnert, Marc Reichenbach
FCCM3
2026 Integrating an open-source soft-GPU overlay with RISC-V control and high-bandwidth memory
abstract
Image and signal processing workloads are widely deployed on Graphics Processing Units (GPUs) for high throughput and on Field-Programmable Gate Arrays (FPGAs) for hardware specialization and energy efficiency. Soft GPU overlays on FPGAs aim to combine these advantages, yet existing solutions often depend on fixed hard processors or impose platform constraints that limit portability. This work extends a popular open-source soft GPGPU overlay to integrate a soft RISC-V control plane and enable compatibility with High-Bandwidth Memory (HBM2). The resulting system can be instantiated on FPGA boards without a hard ARM processor, improving portability, simplifying system integration, and broadening deployability. Across representative image and signal processing kernels, the soft GPGPU achieves geometric-mean speedups of 114.60 × over a scalar soft RISC-V core and 19.72 × over a hard ARM core, demonstrating substantial performance benefits while retaining FPGA reconfigurability. HBM2 integration further benefits bandwidth-sensitive workloads by increasing sustained throughput and reducing the performance bottlenecks associated with off-chip memory access. Collectively, these results indicate that GPU-like programmability and performance can be delivered on reconfigurable platforms without reliance on hard CPU subsystems, providing a portable and scalable foundation for embedded vision and DSP acceleration.
Hector Gerardo Muñoz Hernandez, Mahdi Taheri, Muhammad Ali 0010, Keyvan Shahin, Alireza Syavashi, Diana Göhringer, Marc Reichenbach, Christian Herglotz, Michael Hübner 0001
J. Syst. Archit.7
2025 High-Performance, Area-Efficient and Predictable Matrix Multiplication for ASIC
abstract
With Deep Learning dominating modern applications, area-efficient and simultaneously high through-put acceleration of inference has become a crucial application. Constant Matrix Vector Multiplication (CMVM) dominates the computational complexity of these workloads, and its efficient acceleration is essential. Our proposed architecture introduces a groundbreaking extension to Computation Coding for dataflow architectures, achieving a$2.3 \times$performance boost for small matrices and paving the way for more efficient Deep Learning inference at a performance of up to$93.52 \frac{\text{GOP}}{\mathrm{s}}$using a single core.
Liliia Almeeva, Alexander Lehnert, Ralf R. Müller, Marc Reichenbach
ASAP4
2025 HeterogeneousRTOS: A CPU-FPGA Real-Time OS for Fault Tolerance on COTS at Near-Zero Timing Cost
abstract
Ionizing particles in the atmosphere may strike circuits causing Single Event Upsets (SEU), affecting the output correctness. Critical real-time systems are traditionally custom-designed, featuring redundancy for guaranteeing fault resilience. The downsides of such custom systems are typically weight, power, energy, space, and cost, compared to Commercial Off-the-Shelf (COTS) solutions. We explored the use of COTS in critical real-time environments by designing a CPU-FPGA heterogeneous system, which features an ARM CPU, running a modified version of FreeRTOS and an FPGA, on which the fault-detector and the scheduler are synthesized, in a redundant configuration for increasing fault resiliency. Moving the scheduler to the FPGA increases its fault resiliency while removing the periodic scheduler execution overhead from the CPU, making the scheduler overhead negligible and allowing for an elevated time resolution: the tasks can almost completely utilize the CPU time. Similarly, synthesizing the fault detector on the FPGA allows the execution of the fault detection in a fault-tolerant way without wasting CPU time. Transient fault resiliency in application tasks is achieved via fault detection and the subsequent fault recovery via re-execution. The fault detector implemented on FPGA uses a machine learning technique to model the behavior of tasks (offline and possibly online) and analyses it during their execution. Regarding fault recovery, the scheduler on the FPGA features a novel mixed-criticality scheduling algorithm that manages re-executions, ensuring the meeting of tasks’ timing constraints. The fault detection showed noticeable results while providing a lower overhead than general-purpose software techniques for improving fault resiliency. To the best of our knowledge, the integrated CPU-FPGA version of the system, featuring fault-tolerance and real-time scheduling, is a novel contribution that may enable the use of low-cost and fast COTS components in critical real-time environments. The source code for both hardware and software was released as open source.
Francesco Ratti, Johannes Knödtel, Marc Reichenbach
ACM Trans. Embed. Comput. Syst.3
2025 RISC-V CPU Design Using RRAM-CMOS Standard Cells
abstract
The breakdown of Dennard scaling has been the driver for many innovations such as multicore CPUs and has fueled the research into novel devices such as resistive random access memory (RRAM). These devices might be a means to extend the scalability of integrated circuits since they allow for fast and nonvolatile operation. Unfortunately, large analog circuits need to be designed and integrated in order to benefit from these cells, hindering the implementation of large systems. This work elaborates on a novel solution, namely, creating digital standard cells utilizing RRAM devices. Albeit this approach can be used both for small gates and large macroblocks, we illustrate it for a 2T2R-cell. Since RRAM devices can be vertically stacked with transistors, this enables us to construct anandstandard cell, which merely consumes the area of two transistors. This leads to a 25% area reduction compared to an equivalent CMOSnandgate. We illustrate achievable area savings with a half-adder circuit and integrate this novel cell into a digital standard cell library. A synthesized RISC-V core using RRAM-based cells results in a 10.7% smaller area than the equivalent design using standard CMOS gates.
Markus Fritscher, Max Uhlmann, Philip Ostrovskyy, Daniel Reiser, Junchao Chen 0001, Jianan Wen, Carsten Schulze, Gerhard Kahmen, Dietmar Fey, Marc Reichenbach, Milos Krstic, Christian Wenger
IEEE Trans. Very Large Scale Integr. Syst.10
2024 Towards an Embedded System for Failure Diagnosis in Drones Using AI and SAC-DM on FPGA
abstract
We present a way of failure detection in real-time unmanned aerial vehicles (UAVs) by integrating Chaos Theory and AI techniques on an FPGA board. The Signal Analysis based on Chaos using the Density of Maxima (SAC-DM) validates the input of the Machine Learning (ML) model due to the relation between the density of maxima and autocorrelation length. While the accuracies achieved solely by SAC-DM are not remarkably high, the ML model demonstrates an accuracy of 92.46% when utilizing sac-dm results as inputs. The unprecedented integration of SAC-DM on FPGA board serves as a solution for high-speed onboard processing, parallel integrated data synchronization and fusion, and an enhanced low-power architecture.
Rafael Batista, Matthias Nickel, Alexander Lehnert, Sergio A. Pertuz 0001, Marc Reichenbach, Diana Göhringer, Alisson Brito
DATE5
2023 Increasing the Robustness of TERO-TRNGs Against Process Variation
abstract
The transition effect ring oscillator is a popular design for building entropy sources because it is compact, built from digital elements only, and is very well suited for FPGAs. However, it is known to be quite sensitive to process variation. Although the latter is useful for building physical unclonable functions, it is interfering with the application as an entropy source. In this article, we investigate an approach to increase reliability. We show that adding a third stage eliminates much of the susceptibility to process variation and how a resulting gigahertz oscillation can be evaluated on an FPGA. The design is supported by physical and stochastic modeling. The physical model is validated using an experiment with dynamically reconfigurable look-up tables.
Christian Skubich, Peter Reichel, Marc Reichenbach
ACM Trans. Reconfigurable Technol. Syst.3
2022 Suitability of ISAs for Data Paths Based on Redundant Number Systems: Is RISC-V the best?
abstract
It has been known for a long time that in processor design, delay in arithmetic circuits can be reduced by using redundant number representations (RNS). This advantage is currently only exploited to a limited extent since various aspects complicate its use. For this reason the redundant representation is abandoned at the boundaries of the AL U and the values are reconverted back to the traditional binary representation. In particular, some operations that are traditionally considered fast are now subject to a higher delay. Among other concerns, this complicates comparison operations (e.g. equal to, greater than) and thus affects the timing behavior of conditional jumps. There is some initial research promising speedups using RNS in register files and the data path, but there are still some open questions. In particular it is important to evaluate how the instruction set is designed. Such a study is necessary to estimate whether it is worthwhile to develop the entire data path beyond the AL U in redundant representation. If no reconversion from redundant to traditional binary number system takes place, then e.g. the evaluation of condition codes or flags is problematic, since all speed advantages are lost again. In this work a qualitative and quantitative analysis of common ISAs in processors with redundant data paths is presented. All relevant properties of an ISA are identified and an evaluation of several common ISAs according to these criteria. A performance comparison of three common RISC ISAs (MIPS, A64 (ARM), RISC- V) is given based on a simulation of the Embench bench-mark suite using an adapted version of QEMU. This comparison estimates the speedup of processors with redundant versus binary data paths. It was found, that RISC- V was overall outperforming the other ISA with a maximum speedup of 1.41.
Johannes Knödtel, Sebastian Rachuj, Marc Reichenbach
DSD3
2021 A Case for Function-as-a-Service with Disaggregated FPGAs
abstract
The slowdown of Moore's law and the end of Dennard scaling created a demand for specialized accelerators, including Field Programmable Gate Arrays (FPGAs), in cloud data centers. At the same time, compute resources are increasingly consumed via public and private clouds and traditional applications are modernized using scalable microservices and Function-as-a-Service (FaaS) offerings. Nonetheless, true FaaS based on FPGAs or other accelerators is virtually absent from the offering catalogs of all major cloud providers. In addition, FPGA applications are typically coded in a monolithic fashion, due to device and vendor specific dependencies, which reduces the portability and usability of FPGA cloud offerings further. However, FPGA-based FaaS can improve execution efficiency and minimize (tail-) latencies while decreasing costs. We propose a novel system architecture, called Mantle, that uses disaggregated FPGAs to enable scalable, usable, portable and efficient FaaS offerings for FPGAs. Our experimental results demonstrate a significant reduction of end-to-end service provisioning time to below 7 seconds and an increase in execution efficiency by a factor of 4 with negligible overhead.
Burkhard Ringlein, François Abel, Dionysios Diamantopoulos, Beat Weiss, Christoph Hagleitner, Marc Reichenbach, Dietmar Fey
CLOUD6
2021 Taming Non-Deterministic Low-Level I/O: Predictable Multi-Core Real-Time Systems by SoC Co-Design
abstract
Predictable and analyzable I/O is one of the considerable challenges in the design of multi-core real-time systems. A common approach to tackle this issue is to partition and schedule I/O transactions such that interference between tasks is minimized. While this works for packet-oriented interfaces with deterministic blocking times, such as ethernet, these techniques are inapplicable to a whole range of I/O devices with nondeterministic behavior that is commonly found in embedded applications. Interfaces, such as SPI, do not allow for fine-grained scheduling and thus exhibit uncontrolled blocking times. Even worse, their configuration and use must be considered as independent transactions requiring costly synchronization between tasks. The resulting detrimental effects are, in particular, pronounced in settings with mixed task requirements on predictability and determinism. All this makes the temporal analysis of such systems cumbersome and overly pessimistic. To solve these issues, we present LOWI/O, an approach to eliminate the interference of low-level non-deterministic I/O interfaces for real-time tasks with high predictability demands (i.e., critical task) while preserving flexibility for tasks with lower requirements (i.e., uncritical tasks). Therefore, we leverage knowledge about the application-specific I/O usage patterns, obtained by static analysis, to derive a tailored hardware architecture. Its key feature is the anticipatory reservation of individual time slots for critical tasks and to mimic preemptivity of I/O units for the remaining system. We have implemented our approach as a toolchain for OSEK-based real-time systems that automatically generates an application-specific SoC design along with a hardware and timing model for subsequent WCET analysis. Our experimental results prove predictable timing for critical tasks with limited impact on uncritical tasks.
Steffen Vaas, Peter Ulbrich, Christian Eichler, Peter Wägemann, Marc Reichenbach, Dietmar Fey
ISORC5
2018 Modeling the Energy Consumption of the HEVC Decoding Process
abstract
In this paper, we present a bit stream feature-based energy model that accurately estimates the energy required to decode a given High Efficiency Video Coding-coded bit stream. Therefore, we take a model from literature and extend it by explicitly modeling the in-loop filters, which was not done before. Furthermore, to prove its superior estimation performance, it is compared with seven different energy models from the literature. By using a unified evaluation framework, we show how accurately the required decoding energy for different decoding systems can be approximated. We give thorough explanations on the model parameters and explain how the model variables are derived. To show the modeling capabilities in general, we test the estimation performance for different decoding software and hardware solutions, where we find that the proposed model outperforms the models from the literature by reaching framewise mean estimation errors of less than 7% for software and less than 15% for hardware-based systems.
Christian Herglotz, Dominic Springer, Marc Reichenbach, Benno Stabernack, André Kaup
IEEE Trans. Circuits Syst. Video Technol.3
2017 Fast heterogeneous computing architectures for smart antennas
Marc Reichenbach, Maximilian Kasparek, Konrad Häublein, Jan Niklas Bauer, Mohammad Alawieh, Dietmar Fey
J. Syst. Archit.1
2016 Low-power analog smart camera sensor for edge detection
abstract
This work presents an intelligent analog image sensor system for smart camera applications with the need of edge or marker detection. The system consists of a 3×3 read-out CMOS image sensor, an analog Sobel stage and additional circuitry like operational amplifiers and comparators to compute a 1 bit image with the edges present in the taken photo. This information can then be further processed digitally to detect specific shapes in order to control robot routines, for example. The architecture of the proposed system is highly desirable as dedicated analog hardware has significant advantages in terms of power and speed compared to digital implementations. The overall system is simulated with the help of a 3×3 CMOS image sensor IC as well as Cadence Virtuoso for analog circuit simulation and MATLAB to convert the sequential information back to an image, and compared to other state of the art CMOS image sensors with edge detection capability. The analog Sobel circuit runs with a clock of 10 MHz and consumes less than 0.79 mW average power for the computation of the example image, and the whole 200×200 pixel image sensor consumes only 5.5 mW at a frame rate of 75 fps.
Christopher Soell, Lan Shi, Jürgen Röber, Marc Reichenbach, Robert Weigel, Amelie Hagelauer
ICIP4
2015 Synthesis and optimization of image processing accelerators using domain knowledge
Oliver Reiche, Konrad Häublein, Marc Reichenbach, Moritz Schmid, Frank Hannig, Jürgen Teich, Dietmar Fey
J. Syst. Archit.3
2012 Realizing real-time centroid detection of multiple objects with marching pixels algorithms on programmable customizing hardware
abstract
SUMMARY In this paper, we present a class of emergent algorithms called Marching Pixels and a corresponding programmable parallel chip architecture. Marching Pixels can be used for real‐time image processing in smart camera chips. They are based on hardware agents, which are virtually crawling in a pixel grid image to find attributes like centroid, rotation, and size of an arbitrary number of objects given in an image. Because of the distributed and local processing scheme of Marching Pixels, reply times in milliseconds can be fulfilled. This means that time is determined where pre‐known objects are located and how they are oriented to the main axes of the image. We present an example Marching Pixels algorithm and corresponding application‐specific and programmable parallel architectures. The latter contains a specific instruction set that allows not only the execution of Marching Pixels algorithms but also of arbitrary Cellular Automata algorithms as an embedded parallel processor. The strengths and weaknesses of this architecture concerning the realization as field‐programmable gate arrays and application‐specific integrated circuits are discussed by means of hardware synthesis results. These results are compared with the solution achievable on a real hardware like the Atom processor. Copyright © 2011 John Wiley & Sons, Ltd.
Dietmar Fey, Marc Reichenbach, Marcus Komann, Ralf Seidler
Concurr. Comput. Pract. Exp.2
2009 Distributed vision with smart pixels
abstract
We study a problem related to computer vision: How can a field of sensors compute higher-level properties of observed objects deterministically in sublinear time, without accessing a central authority? This issue is not only important for real-time processing of images, but lies at the very heart of understanding how a brain may be able to function. In particular, we consider a quadratic field of n "smart pixels" on a video chip that observe a B/W image. Each pixel can exchange low-level information with its immediate neighbors. We show that it is possible to compute the centers of gravity along with a principal component analysis of all connected components of the black grid graph in time O(sqrt(n)), by developing appropriate distributed protocols that are modeled after sweepline methods. Our method is not only interesting from a philosophical and theoretical point of view, it is also useful for actual applications for controling a robot arm that has to seize objects on a moving belt. We describe details of an implementation on an FPGA; the code has also been turned into a hardware design for an application-specific integrated circuit (ASIC).
Sándor P. Fekete, Dietmar Fey, Marcus Komann, Alexander Kröller, Marc Reichenbach, Christiane Schmidt 0001
SCG5