EDBT 2026 Demo / reviewers in the wild / expert
Benno Stabernack
dblp:27/2776
· DBLP profile ↗
13ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-6654-1606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UNNIP: Memory Optimized Universal Neural Network Intra-Prediction for Custom On-Chip Hardware-ArchitecturesabstractNeural network-based intra prediction has demonstrated substantial improvements over traditional prediction modes from block based hybrid video coding frameworks such as HEVC and VVC. On the downside, the used neural networks introduce enormous complexity that obstructs efficient hardware. Due to the inherent sequential structure of intra prediction, this category of tools cannot be accelerated with off-chip NPUs, while on-chip integration entails high area costs induced by the memory-footprint of the method. This paper proposes a universal neural network intra prediction model that is co-designed with an experimental on-chip architecture. The method leverages tasklevel redundancies to achieve a significantly improved tradeoff in terms of coding-gain to memory-footprint. Compared to related work, our method reduces the overall weight-parameters by × 4.38 - near orthogonal to traditional model optimization like pruning and quantization - while the BD-rate also slightly improves by -0. 1 0 %(Y),-0. 1 4 %(U) and -0. 1 0 %(V). This optimization enables the deployment to a prototype FPGA implementation. Viktor Herrmann, Justin Knapheide, Benno Stabernack |
PCS | 3 |
| 2023 | Demonstrating NADA: A Workflow for Distributed CNN Training on FPGA ClustersabstractWe introduce our Network Attached Deep learning Accelerator called NADA, which consists of a novel and flexible HW/SW framework for efficient training of deep neural networks on FPGA clusters. NADA is centered around layer parallelism, instantiating a specific implementation for each layer. These implementations are placed across the desired number of network attached FPGAs in the cluster. The NADA hardware architecture relies on a high-speed UDP/IP pure hardware network stack. We demonstrate the usability of our approach by training a couple of demo networks on an Arria 10 FPGA cluster. Justin Knapheide, Philipp Kreowsky, Benno Stabernack |
FPL | 3 |
| 2023 | Challenges Using FPGA Clusters for Distributed CNN TrainingabstractWhile FPGAs are well-established in the field of CNN inference, there is a lack of research regarding CNN training. To mitigate this, we develop an easy-to-use framework for CNN training on network-attached FPGA clusters. Tightly pipelined layer parallelism promises to facilitate a high utilization across a whole cluster. This comes with challenges, which we discuss in this paper. Philipp Kreowsky, Justin Knapheide, Benno Stabernack |
FPL | 3 |
| 2023 | A High-Throughput, Resource-Efficient Implementation of the RoCEv2 Remote DMA Protocol and its ApplicationabstractThe use of application-specific accelerators in data centers has been the state of the art for at least a decade, starting with the availability of General Purpose GPUs achieving higher performance either overall or per watt. In most cases, these accelerators are coupled via PCIe interfaces to the corresponding hosts, which leads to disadvantages in interoperability, scalability and power consumption. As a viable alternative to PCIe-attached FPGA accelerators this paper proposes standalone FPGAs as Network-attached Accelerators (NAAs) . To enable reliable communication for decoupled FPGAs we present an RDMA over Converged Ethernet v2 (RoCEv2) communication stack for high-speed and low-latency data transfer integrated into a hardware framework. For NAAs to be used instead of PCIe coupled FPGAs the framework must provide similar throughput and latency with low resource usage. We show that our RoCEv2 stack is capable of achieving 100 Gb/s throughput with latencies of less than 4μs while using about 10% of the available resources on a mid-range FPGA. To evaluate the energy efficiency of our NAA architecture, we built a demonstrator with 8 NAAs for machine learning based image classification. Based on our measurements, network-attached FPGAs are a great alternative to the more energy-demanding PCIe-attached FPGA accelerators. Niklas Schelten, Fritjof Steinert, Justin Knapheide, Anton Schulte, Benno Stabernack |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2022 | A YOLO v3-tiny FPGA Architecture using a Reconfigurable Hardware Accelerator for Real-time Region of Interest DetectionabstractWith the recent advances in the fields of machine learning, neural networks and deep-learning algorithms have become a prevalent subject of computer vision. Especially for tasks like object classification and detection Convolutional Neu-ronal Networks (CNNs) have surpassed the previous traditional approaches. In addition to these applications, CNNs can recently also be found in other applications. For example the parametrization of video encoding algorithms as used in our example is quite a new application domain. Especially CNN's high recognition rate makes them particularly suitable for finding Regions of Interest (ROIs) in video sequences, which can be used for adapting the data rate of the compressed video stream accordingly. On the downside, these CNN require an immense amount of processing power and memory bandwidth. Object detection networks such as You Only Look Once (YOLO) try to balance processing speed and accuracy but still rely on power-hungry GPUs to meet real-time requirements. Specialized hardware like Field Programmable Gate Array (FPGA) implementations proved to strongly reduce this problem while still providing sufficient computational power. In this paper we propose a flexible architecture for object detection hardware acceleration based on the YOLO v3-tiny model. The reconfigurable accelerator comprises a high throughput convolution engine, custom blocks for all additional CNN operations and a programmable control unit to manage on-chip execution. The model can be deployed without significant changes based on 32-bit floating point values and without further methods that would reduce the model accuracy. Experimental results show a high capability of the design to accelerate the object detection task with a processing time of 27.5 ms per frame. It is thus real-time-capable for 30 FPS applications at frequency of 200 MHz. Viktor Herrmann, Justin Knapheide, Fritjof Steinert, Benno Stabernack |
DSD | 4 |
| 2021 | Demonstration of a Distributed Accelerator Framework for Energy-efficient ML ProcessingabstractThe broad field of machine learning (ML) is experiencing rapid growth from being applied to ever more types of applications. This leads to an increasing demand for computing performance. To satisfy the computational demand, accelerators like GPUs or FPGAs are used especially in data centers. Despite the higher energy efficiency of accelerators compared to CPUs, new architectures for accelerator usage are necessary to achieve a significant further improvement in energy efficiency. Therefore, we demonstrate an architecture for distributed ML accelerators. We operate FPGA accelerators directly coupled to the network without a host system, which leads to significant power savings. Fritjof Steinert, Justin Knapheide, Benno Stabernack |
FPL | 3 |
| 2020 | Hardware and Software Components towards the Integration of Network-Attached Accelerators into Data CentersabstractThe usage of Field Programmable Gate Arrays(FPGAs) as application-specific accelerators in data centers has grown rapidly in recent years. The main reasons for this are the higher throughput of application-specific accelerators, an improved latency behavior as well as higher energy efficiency compared to conventional computing systems with CPUs and GPUs. In particular, network-attached FPGA accelerators exhibit excellent latency behavior and very high energy efficiency. The high efficiency not only decreases operating costs, but also results in a lower carbon footprint. A disadvantage compared to conventional CPUs and GPUs is the difficult programmability as well as the integration into data centers, which hinders a widespread adoption. To simplify the integration of FPGA accelerators in heterogeneous data centers, we have designed a flexible hardware and software framework to overcome these problems. We also show which components are necessary to fulfill the requirements within a data center. Fritjof Steinert, Niklas Schelten, Anton Schulte, Benno Stabernack |
DSD | 4 |
| 2020 | A High Throughput MobileNetV2 FPGA Implementation Based on a Flexible Architecture for Depthwise Separable ConvolutionabstractConvolutional Neural Networks are widely applied to various computer vision tasks. For most of these applications, high throughput and energy efficiency are top priorities. MobileNetV2 features very low memory requirements as well as a relatively small model size. On the ILSVRC 2012 classification challenge, it provides a decent prediction accuracy of 71.7 percent at low computational requirements. We present an FPGA based MobileNetV2 accelerator with a high throughput of 1050 frames per second at a power consumption of 34 watt under full load. This equates to a power efficiency of 32 milli-joule per frame. We describe our approach of using stream interfaces and auto-generated control signals to enable fast design of flexible architectures. By using quantization techniques, limiting the accuracy of the used number format to a 16 bit fixed point format, we were able to reduce the memory usage for weights as well as activations by a factor of two. Since the basic building block of MobileNetV2 can be used to build higher performance networks as well, the findings of this paper remain applicable, when higher prediction accuracies are required. Justin Knapheide, Benno Stabernack, Maximilian Kuhnke |
FPL | 2 |
| 2018 | Modeling the Energy Consumption of the HEVC Decoding ProcessabstractIn this paper, we present a bit stream feature-based energy model that accurately estimates the energy required to decode a given High Efficiency Video Coding-coded bit stream. Therefore, we take a model from literature and extend it by explicitly modeling the in-loop filters, which was not done before. Furthermore, to prove its superior estimation performance, it is compared with seven different energy models from the literature. By using a unified evaluation framework, we show how accurately the required decoding energy for different decoding systems can be approximated. We give thorough explanations on the model parameters and explain how the model variables are derived. To show the modeling capabilities in general, we test the estimation performance for different decoding software and hardware solutions, where we find that the proposed model outperforms the models from the literature by reaching framewise mean estimation errors of less than 7% for software and less than 15% for hardware-based systems. Christian Herglotz, Dominic Springer, Marc Reichenbach, Benno Stabernack, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Simulation-based HW/SW co-exploration of the concurrent execution of HEVC intra encoding algorithms for heterogeneous multi-core architecturesabstractThe high efficiency video coding (HEVC) standard shows enhanced video compression efficiency at the cost of high performance requirements. To address these requirements different approaches, like algorithmic optimization, parallelization and hardware acceleration can be used leading to a complex design space. In order to find an efficient solution, early design verification and performance evaluation is crucial. Hereby the prevailing methodology is the simulation of the complex HW/SW architecture. Targeting heterogeneous designs, different simulation models have different performance evaluation capabilities making a combined HW/SW co-analysis of the entire system a cumbersome task. To facilitate this co-analysis, we propose a non-intrusive instrumentation methodology for simulation models, which automatically adapts to the model under observation. With the help of this instrumentation methodology we perform the analysis and exploration of different design aspects of a SystemC-based heterogeneous multi-core model of an HEVC intra encoder. In the course of this HW/SW co-analysis various aspects of the parallelization and hardware acceleration of the video coding algorithms are presented and further improved. Due to its cycle accurate nature the developed model is well suited to facilitate various performance evaluations and to drive HW/SW co-optimizations of the explored system, as discussed in this paper. Jens Brandenburg, Benno Stabernack |
J. Syst. Archit. | 2 |
| 2013 | Architecture of a Low Latency Image Rectification Engine for Stereoscopic 3-D HDTV ProcessingabstractThe emerging market of digital 3-D film productions in HD resolution leads to the need for high-quality equipment in the production chain. The incoming video streams of the two cameras require an image rectification due to unavoidable misalignments within the stereoscopic camera setup. This rectification can either take place in postprocessing of the recorded material or it can be applied in real time during the shooting. Especially in the case of streaming and recording of live events, real-time processing is necessary and, additionally, the system has to provide a very low latency. We present a hardware image rectification engine, which supports the processing of stereo high-definition serial digital interfaces video streams with up to 1080p30 video with a latency below 1 ms. The image rectification engines for the two channels are implemented on two Altera Stratix III EP3SL340 running at 74.25 MHz. They can be controlled by the stereoscopy analysis software, which calculates the parameters required for the image rectification at runtime. Heiko Hübert, Benno Stabernack, Frederik Zilly |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Profiling-Based Hardware/Software Co-Exploration for the Design of Video Coding ArchitecturesabstractThe design of embedded hardware/software systems is often subject to strict requirements concerning its various aspects, including real-time performance, power consumption, and die area. Especially for data intensive applications, the number of memory accesses is a dominant factor for these aspects. In order to meet the requirements and design a well-adapted system, the software parts need to be optimized and an adequate system and processor architecture needs to be designed. In this paper, we focus on finding an optimized memory hierarchy for bus-based architectures. Additionally, useful instruction set extensions for application-specific processor cores are explored. For complex applications, this design space exploration is difficult and requires in-depth analysis of the application and its implementation alternatives. Tools are required which aid the designer in the design, optimization, and scheduling of hardware and software. We present a profiling tool for fast and accurate performance, power, and memory access analysis of embedded systems. This paper shows how the tool can be applied for an efficient hardware/software co-exploration within the design flow of processor-centric architectures. This concept has been proven in the design of a mixed hardware/software system with multiple processing units for video decoding. Heiko Hübert, Benno Stabernack |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | A system for QOS-enabled MPEG-4 video transmission over Bluetooth for mobile applicationsabstractStreaming media distribution is a rapidly growing application. The increasing number of broadband Internet subscribers, the growing market for smart in-home networks and the upcoming broadband mobile phone networks encourage this trend. In addition, the number of multimedia enabled mobile devices such as PDAs, smart embedded devices and notebooks is rapidly growing. MPEG-4 is one of the favorite video coding schemes, and Bluetooth is a preferred wireless technology for small and smart devices because of its small size, low power consumption and cost efficiency. In this paper we present a complete quality of service (QoS) enabled system that integrates a MPEG-4 codec and a Bluetooth device for mobile video applications. Corina Scheiter, Rainer Steffen, Markus Zeller, Rudi Knorr, Benno Stabernack, Kai-Immo Wels |
ICME | 5 |