Fritjof Steinert

dblp:276/2259 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0002-8733-3064ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2023 A High-Throughput, Resource-Efficient Implementation of the RoCEv2 Remote DMA Protocol and its Application
abstract
The use of application-specific accelerators in data centers has been the state of the art for at least a decade, starting with the availability of General Purpose GPUs achieving higher performance either overall or per watt. In most cases, these accelerators are coupled via PCIe interfaces to the corresponding hosts, which leads to disadvantages in interoperability, scalability and power consumption. As a viable alternative to PCIe-attached FPGA accelerators this paper proposes standalone FPGAs as Network-attached Accelerators (NAAs) . To enable reliable communication for decoupled FPGAs we present an RDMA over Converged Ethernet v2 (RoCEv2) communication stack for high-speed and low-latency data transfer integrated into a hardware framework. For NAAs to be used instead of PCIe coupled FPGAs the framework must provide similar throughput and latency with low resource usage. We show that our RoCEv2 stack is capable of achieving 100 Gb/s throughput with latencies of less than 4μs while using about 10% of the available resources on a mid-range FPGA. To evaluate the energy efficiency of our NAA architecture, we built a demonstrator with 8 NAAs for machine learning based image classification. Based on our measurements, network-attached FPGAs are a great alternative to the more energy-demanding PCIe-attached FPGA accelerators.
Niklas Schelten, Fritjof Steinert, Justin Knapheide, Anton Schulte, Benno Stabernack
ACM Trans. Reconfigurable Technol. Syst.2
2022 A YOLO v3-tiny FPGA Architecture using a Reconfigurable Hardware Accelerator for Real-time Region of Interest Detection
abstract
With the recent advances in the fields of machine learning, neural networks and deep-learning algorithms have become a prevalent subject of computer vision. Especially for tasks like object classification and detection Convolutional Neu-ronal Networks (CNNs) have surpassed the previous traditional approaches. In addition to these applications, CNNs can recently also be found in other applications. For example the parametrization of video encoding algorithms as used in our example is quite a new application domain. Especially CNN's high recognition rate makes them particularly suitable for finding Regions of Interest (ROIs) in video sequences, which can be used for adapting the data rate of the compressed video stream accordingly. On the downside, these CNN require an immense amount of processing power and memory bandwidth. Object detection networks such as You Only Look Once (YOLO) try to balance processing speed and accuracy but still rely on power-hungry GPUs to meet real-time requirements. Specialized hardware like Field Programmable Gate Array (FPGA) implementations proved to strongly reduce this problem while still providing sufficient computational power. In this paper we propose a flexible architecture for object detection hardware acceleration based on the YOLO v3-tiny model. The reconfigurable accelerator comprises a high throughput convolution engine, custom blocks for all additional CNN operations and a programmable control unit to manage on-chip execution. The model can be deployed without significant changes based on 32-bit floating point values and without further methods that would reduce the model accuracy. Experimental results show a high capability of the design to accelerate the object detection task with a processing time of 27.5 ms per frame. It is thus real-time-capable for 30 FPS applications at frequency of 200 MHz.
Viktor Herrmann, Justin Knapheide, Fritjof Steinert, Benno Stabernack
DSD3
2021 Demonstration of a Distributed Accelerator Framework for Energy-efficient ML Processing
abstract
The broad field of machine learning (ML) is experiencing rapid growth from being applied to ever more types of applications. This leads to an increasing demand for computing performance. To satisfy the computational demand, accelerators like GPUs or FPGAs are used especially in data centers. Despite the higher energy efficiency of accelerators compared to CPUs, new architectures for accelerator usage are necessary to achieve a significant further improvement in energy efficiency. Therefore, we demonstrate an architecture for distributed ML accelerators. We operate FPGA accelerators directly coupled to the network without a host system, which leads to significant power savings.
Fritjof Steinert, Justin Knapheide, Benno Stabernack
FPL1
2020 Hardware and Software Components towards the Integration of Network-Attached Accelerators into Data Centers
abstract
The usage of Field Programmable Gate Arrays(FPGAs) as application-specific accelerators in data centers has grown rapidly in recent years. The main reasons for this are the higher throughput of application-specific accelerators, an improved latency behavior as well as higher energy efficiency compared to conventional computing systems with CPUs and GPUs. In particular, network-attached FPGA accelerators exhibit excellent latency behavior and very high energy efficiency. The high efficiency not only decreases operating costs, but also results in a lower carbon footprint. A disadvantage compared to conventional CPUs and GPUs is the difficult programmability as well as the integration into data centers, which hinders a widespread adoption. To simplify the integration of FPGA accelerators in heterogeneous data centers, we have designed a flexible hardware and software framework to overcome these problems. We also show which components are necessary to fulfill the requirements within a data center.
Fritjof Steinert, Niklas Schelten, Anton Schulte, Benno Stabernack
DSD1