Sven Gesper

dblp:246/5377 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-3570-1638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector Processor
abstract
State-of-the-art applications, such as convolutional neural networks, demand specialized hardware accelerators that address performance and efficiency constraints. An efficient memory hierarchy is mandatory for such hardware systems. While the memory architectures of general-purpose processors (e.g., CPU or GPUs) are based on cache systems, dedicated accelerators have mostly adopted the DMA (Direct Memory Access) concept due to the application field of image processing. DMA features like 2D data transfers or data padding can optimize the memory accesses of image processing. However, DMA lacks the capability to exploit temporal and spatial data reuse, a feature common in cache systems, particularly when multiple DMAs operate in parallel. This article proposes a novel Direct Cached Memory Access (DCMA) architecture, combining both DMA and cache methodologies and their respective advantages. Optimized for image-based AI algorithms, the DCMA architecture facilitates enhanced memory access by integrating multiple, parallel DMA ports with caching capabilities. This design allows for efficient data reuse and parallel memory access. Optimal parameters for the DCMA are determined through a comprehensive design space exploration. The DCMA is evaluated on a state-of-the-art Xilinx UltraScale+ FPGA board coupled with a massive-parallel vertical vector co-processor, called V 2 PRO. The results show the mitigation of the vector processor’s memory bottleneck. By using the proposed DCMA, speedups of up to ×17 for the ResNet-50 CNN can be achieved.
Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá
ACM Trans. Archit. Code Optim.2
2024 Multi-Level Prototyping of a Vertical Vector AI Processing System
abstract
Modern embedded systems must be designed carefully to cope with the complexity and real-time requirements of modern AI (Artificial Intelligence) driven automotive applications, such as Advanced Driver-Assistance Systems (ADAS). Despite increasing complexity, the time to market is decreasing. In this work, a SystemC-based Virtual Prototype of a neural network processing platform is exploited to bypass the limitations of standalone instruction set simulators (ISS) and FPGA prototyping. The processing platform under test is based on a novel massive parallel vector processor architecture coupled with a RISC- V control core that runs widely used convolutional neural networks (CNNs) for object detection. The paper discusses the variations and appropriateness of the three prototyping methods outlined, demonstrating how the Virtual Prototype can address the aforementioned constraints, resulting in a 2.07x increase in accuracy, 16x greater configurations, and more profound insights into the system compared to standalone and FPGA prototyping.
Frederik Kautz, Sven Gesper, Gia Bao Thieu, Hans-Martin Blüthgen, Holger Blume, Guillermo Payá-Vayá
ASAP2
2023 Exploiting Subword Permutations to Maximize CNN Compute Performance and Efficiency
abstract
Neural networks (NNs) are quantized to decrease their computational demands and reduce their memory foot-print. However, specialized hardware is required that supports computations with low bit widths to take advantage of such optimizations. In this work, we propose permutations on subword level that build on top of multi-bit-width multiply-accumulate operations to effectively support low bit width computations of quantized NNs. By applying this technique, we extend the data reuse and further improve compute performance for convolution operations compared to simple vectorization using SIMD (single-instruction-multiple-data). We perform a design space exploration using a cycle accurate simulation with MobileNet and VGG16 on a vector-based processor. The results show a speedup of up to$3.7\times$and a reduction of up to$1.9\times$for required data transfers. Additionally, the control overhead for orchestrating the computation is decreased by up to$3.9\times$.
Michael Beyer, Sven Gesper, Andre Guntoro, Guillermo Payá-Vayá, Holger Blume
ASAP2
2023 ZuSE Ki-Avf: Application-Specific AI Processor for Intelligent Sensor Signal Processing in Autonomous Driving
abstract
Modern and future AI-based automotive applications, such as autonomous driving, require the efficient real-time processing of huge amounts of data from different sensors, like camera, radar, and LiDAR. In the ZuSE-KI-AVF project, multiple university, and industry partners collaborate to develop a novel massive parallel processor architecture, based on a cus-tomized RISC-V host processor, and an efficient high-performance vertical vector coprocessor. In addition, a software development framework is also provided to efficiently program AI-based sensor processing applications. The proposed processor system was verified and evaluated on a state-of-the-art UltraScale+ FPGA board, reaching a processing performance of up to 126.9 FPS, while executing the YOLO-LITE CNN on 224x224 input images. Further optimizations of the FPGA design and the realization of the processor system on a 22nm FDSOI CMOS technology are planned.
Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá, Christoph Riggers, Oliver Renke, Till Fiedler, Jakob Marten, Tobias Stuckenberg, Holger Blume, Christian Weis, Lukas Steiner, Chirag Sudarshan, Norbert Wehn, Lennart M. Reimann, Rainer Leupers, Michael Beyer, Daniel Köhler, Alisa Jauch, Jan Micha Borrmann, Setareh Jaberansari, Tim Berthold, Meinolf Blawat, Markus Kock, Gregor Schewior, Jens Benndorf, Frederik Kautz, Hans-Martin Blüthgen, Christian Sauer 0001
DATE2