EDBT 2026 Demo / reviewers in the wild / expert
Justin Knapheide
dblp:276/3894
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-3931-4288ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UNNIP: Memory Optimized Universal Neural Network Intra-Prediction for Custom On-Chip Hardware-ArchitecturesabstractNeural network-based intra prediction has demonstrated substantial improvements over traditional prediction modes from block based hybrid video coding frameworks such as HEVC and VVC. On the downside, the used neural networks introduce enormous complexity that obstructs efficient hardware. Due to the inherent sequential structure of intra prediction, this category of tools cannot be accelerated with off-chip NPUs, while on-chip integration entails high area costs induced by the memory-footprint of the method. This paper proposes a universal neural network intra prediction model that is co-designed with an experimental on-chip architecture. The method leverages tasklevel redundancies to achieve a significantly improved tradeoff in terms of coding-gain to memory-footprint. Compared to related work, our method reduces the overall weight-parameters by × 4.38 - near orthogonal to traditional model optimization like pruning and quantization - while the BD-rate also slightly improves by -0. 1 0 %(Y),-0. 1 4 %(U) and -0. 1 0 %(V). This optimization enables the deployment to a prototype FPGA implementation. Viktor Herrmann, Justin Knapheide, Benno Stabernack |
PCS | 2 |
| 2023 | Demonstrating NADA: A Workflow for Distributed CNN Training on FPGA ClustersabstractWe introduce our Network Attached Deep learning Accelerator called NADA, which consists of a novel and flexible HW/SW framework for efficient training of deep neural networks on FPGA clusters. NADA is centered around layer parallelism, instantiating a specific implementation for each layer. These implementations are placed across the desired number of network attached FPGAs in the cluster. The NADA hardware architecture relies on a high-speed UDP/IP pure hardware network stack. We demonstrate the usability of our approach by training a couple of demo networks on an Arria 10 FPGA cluster. Justin Knapheide, Philipp Kreowsky, Benno Stabernack |
FPL | 1 |
| 2023 | Challenges Using FPGA Clusters for Distributed CNN TrainingabstractWhile FPGAs are well-established in the field of CNN inference, there is a lack of research regarding CNN training. To mitigate this, we develop an easy-to-use framework for CNN training on network-attached FPGA clusters. Tightly pipelined layer parallelism promises to facilitate a high utilization across a whole cluster. This comes with challenges, which we discuss in this paper. Philipp Kreowsky, Justin Knapheide, Benno Stabernack |
FPL | 2 |
| 2023 | A High-Throughput, Resource-Efficient Implementation of the RoCEv2 Remote DMA Protocol and its ApplicationabstractThe use of application-specific accelerators in data centers has been the state of the art for at least a decade, starting with the availability of General Purpose GPUs achieving higher performance either overall or per watt. In most cases, these accelerators are coupled via PCIe interfaces to the corresponding hosts, which leads to disadvantages in interoperability, scalability and power consumption. As a viable alternative to PCIe-attached FPGA accelerators this paper proposes standalone FPGAs as Network-attached Accelerators (NAAs) . To enable reliable communication for decoupled FPGAs we present an RDMA over Converged Ethernet v2 (RoCEv2) communication stack for high-speed and low-latency data transfer integrated into a hardware framework. For NAAs to be used instead of PCIe coupled FPGAs the framework must provide similar throughput and latency with low resource usage. We show that our RoCEv2 stack is capable of achieving 100 Gb/s throughput with latencies of less than 4μs while using about 10% of the available resources on a mid-range FPGA. To evaluate the energy efficiency of our NAA architecture, we built a demonstrator with 8 NAAs for machine learning based image classification. Based on our measurements, network-attached FPGAs are a great alternative to the more energy-demanding PCIe-attached FPGA accelerators. Niklas Schelten, Fritjof Steinert, Justin Knapheide, Anton Schulte, Benno Stabernack |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | A YOLO v3-tiny FPGA Architecture using a Reconfigurable Hardware Accelerator for Real-time Region of Interest DetectionabstractWith the recent advances in the fields of machine learning, neural networks and deep-learning algorithms have become a prevalent subject of computer vision. Especially for tasks like object classification and detection Convolutional Neu-ronal Networks (CNNs) have surpassed the previous traditional approaches. In addition to these applications, CNNs can recently also be found in other applications. For example the parametrization of video encoding algorithms as used in our example is quite a new application domain. Especially CNN's high recognition rate makes them particularly suitable for finding Regions of Interest (ROIs) in video sequences, which can be used for adapting the data rate of the compressed video stream accordingly. On the downside, these CNN require an immense amount of processing power and memory bandwidth. Object detection networks such as You Only Look Once (YOLO) try to balance processing speed and accuracy but still rely on power-hungry GPUs to meet real-time requirements. Specialized hardware like Field Programmable Gate Array (FPGA) implementations proved to strongly reduce this problem while still providing sufficient computational power. In this paper we propose a flexible architecture for object detection hardware acceleration based on the YOLO v3-tiny model. The reconfigurable accelerator comprises a high throughput convolution engine, custom blocks for all additional CNN operations and a programmable control unit to manage on-chip execution. The model can be deployed without significant changes based on 32-bit floating point values and without further methods that would reduce the model accuracy. Experimental results show a high capability of the design to accelerate the object detection task with a processing time of 27.5 ms per frame. It is thus real-time-capable for 30 FPS applications at frequency of 200 MHz. Viktor Herrmann, Justin Knapheide, Fritjof Steinert, Benno Stabernack |
DSD | 2 |
| 2021 | Demonstration of a Distributed Accelerator Framework for Energy-efficient ML ProcessingabstractThe broad field of machine learning (ML) is experiencing rapid growth from being applied to ever more types of applications. This leads to an increasing demand for computing performance. To satisfy the computational demand, accelerators like GPUs or FPGAs are used especially in data centers. Despite the higher energy efficiency of accelerators compared to CPUs, new architectures for accelerator usage are necessary to achieve a significant further improvement in energy efficiency. Therefore, we demonstrate an architecture for distributed ML accelerators. We operate FPGA accelerators directly coupled to the network without a host system, which leads to significant power savings. Fritjof Steinert, Justin Knapheide, Benno Stabernack |
FPL | 2 |
| 2020 | A High Throughput MobileNetV2 FPGA Implementation Based on a Flexible Architecture for Depthwise Separable ConvolutionabstractConvolutional Neural Networks are widely applied to various computer vision tasks. For most of these applications, high throughput and energy efficiency are top priorities. MobileNetV2 features very low memory requirements as well as a relatively small model size. On the ILSVRC 2012 classification challenge, it provides a decent prediction accuracy of 71.7 percent at low computational requirements. We present an FPGA based MobileNetV2 accelerator with a high throughput of 1050 frames per second at a power consumption of 34 watt under full load. This equates to a power efficiency of 32 milli-joule per frame. We describe our approach of using stream interfaces and auto-generated control signals to enable fast design of flexible architectures. By using quantization techniques, limiting the accuracy of the used number format to a 16 bit fixed point format, we were able to reduce the memory usage for weights as well as activations by a factor of two. Since the basic building block of MobileNetV2 can be used to build higher performance networks as well, the findings of this paper remain applicable, when higher prediction accuracies are required. Justin Knapheide, Benno Stabernack, Maximilian Kuhnke |
FPL | 1 |