Marco Stolba

dblp:267/8210 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2024 A Low-footprint FFT Accelerator for a RISC-V-based Multi-core DSP in FMCW Radars
abstract
Multi-core systems are required by digital signal processors (DSP) to support the revolutionary Multiple-Input Multiple-Output (MIMO) imaging radars in the automotive industry. Such multi-core processors for Frequency Modulated Continuous Wave (FMCW) radars require the use of low- footprint accelerators that would reduce the overhead as the system scales up with the antenna density. In this paper, we propose an FFT accelerator, named RbFFT, optimized for the MIMO radar processing chain. The architecture of RbFFT reduces the overhead by re-using existing memory in the processing element (PE), and employs a dual-radix butterfly engine with mixed bit resolution to optimize resources in dense radars. RbFFT reduces area by implementing for the first time ultra-low compression in its dual twiddle factor ROM. RbFFT also innovates with custom fetching and buffering strategies to improve memory-based FFTs while reusing logic to integrate reverse bit ordering, windowing and inverse FFT (IFFT) within the same accelerator passes. The proposed accelerator is implemented in a 25-Core Smart MPSoC in 22FDX using Adaptive Body Biasing (ABB) at 0.6V. Besides RbFFT being pioneer in specialized FFT accelerators for dense MIMO systems, the results also show state-of-the-art improvements via 11% reduction in the normalized energy consumption, 4% reduction in latency, and 11 times area reduction with relation to previous silicon implementations.
Hector A. Gonzalez, Marco Stolba, Bernhard Vogginger, Tim Rosmeisl, Chen Liu 0031, Christian Mayr 0001
ISCAS2
2021 Ultra-High Compression of Twiddle Factor ROMs in Multi-Core DSP for FMCW Radars
abstract
The increasing density of Multiple-Input Multiple-Output (MIMO) arrays in imaging radars for the automotive industry demands highly parallel systems with low-footprint accelerators, which would enable the concurrent processing of a high number of virtual channels with a low-latency, and without a high area overhead. In this paper, we design, implement, and test multiple handcrafted compression schemes for Twiddle Factor (TF) Read-Only Memories (ROM), aiming to reduce the footprint of a variable-length and dual-radix Fast Fourier Transform (FFT) accelerator in a Multi-core Digital Signal Processor (DSP) for Frequency Modulated Continuous Wave (FMCW) radars. The compression schemes proposed in this paper involve double delta encoding, Radix-specific address optimizations per port, symmetry inclusion, and exploitation of the bit resolution changes within the radar processing chain. All schemes are verified in an FPGA in terms of logic utilization and quantization using a 77-GHz radar, and implemented in a RISCV-based Processing Element (PE) of a Multi-core DSP with an Adaptive Body Bias (ABB) approach in 22FDX technology for assessing area, leakage, and relative latency savings when compared with a dual-ROM equivalent in the state-of-the-art.
Hector A. Gonzalez, Florian Kelber, Marco Stolba, Chen Liu 0031, Bernhard Vogginger, Stefan Hänzsche, Stefan Scholze, Sebastian Höppner, Christian Mayr 0001
ISCAS3
2021 Hardware Implementation of an OPC UA Server for Industrial Field Devices
abstract
Industrial plants suffer from a high degree of complexity and incompatibility in their communication infrastructure, caused by a wild mix of proprietary technologies. This prevents transformation toward Industry 4.0 and the Industrial Internet of Things. Open platform communications unified architecture (OPC UA) is a standardized protocol that addresses these problems with uniform and semantic communication across all levels of the hierarchy. However, its adoption in embedded field devices, such as sensors and actuators, is still lacking due to prohibitive memory and power requirements of software implementations. We have developed a dedicated hardware engine that offloads processing of the OPC UA protocol and enables the realization of compact and low-power field devices with OPC UA support. As part of a proof-of-concept embedded system, we have implemented this engine in a 22-nm FDSOI technology, representing the first ASIC implementation of an OPC UA server. We measured performance, power consumption, and memory footprint of our test chip and compared it with a software implementation based on open62541 and a Raspberry Pi 2B. Our OPC UA hardware engine is 50 times more energy efficient and only requires 36 KiB of memory. The complete system consumes only 24 mW under full load, making it suitable for low-power embedded applications.
Heiner Bauer, Sebastian Höppner, Chris Paul Iatrou, Zohra Charania, Stephan Hartmann 0002, Saif-Ur Rehman, Andreas Dixius, Georg Ellguth, Dennis Walter, Johannes Uhlig, Felix Neumärker, Marc Berthel, Marco Stolba, Florian Kelber, Leon Urbas, Christian Mayr 0001
IEEE Trans. Very Large Scale Integr. Syst.13