Bernhard Vogginger

dblp:86/8924 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0001-9042-5405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2024 A Low-footprint FFT Accelerator for a RISC-V-based Multi-core DSP in FMCW Radars
abstract
Multi-core systems are required by digital signal processors (DSP) to support the revolutionary Multiple-Input Multiple-Output (MIMO) imaging radars in the automotive industry. Such multi-core processors for Frequency Modulated Continuous Wave (FMCW) radars require the use of low- footprint accelerators that would reduce the overhead as the system scales up with the antenna density. In this paper, we propose an FFT accelerator, named RbFFT, optimized for the MIMO radar processing chain. The architecture of RbFFT reduces the overhead by re-using existing memory in the processing element (PE), and employs a dual-radix butterfly engine with mixed bit resolution to optimize resources in dense radars. RbFFT reduces area by implementing for the first time ultra-low compression in its dual twiddle factor ROM. RbFFT also innovates with custom fetching and buffering strategies to improve memory-based FFTs while reusing logic to integrate reverse bit ordering, windowing and inverse FFT (IFFT) within the same accelerator passes. The proposed accelerator is implemented in a 25-Core Smart MPSoC in 22FDX using Adaptive Body Biasing (ABB) at 0.6V. Besides RbFFT being pioneer in specialized FFT accelerators for dense MIMO systems, the results also show state-of-the-art improvements via 11% reduction in the normalized energy consumption, 4% reduction in latency, and 11 times area reduction with relation to previous silicon implementations.
Hector A. Gonzalez, Marco Stolba, Bernhard Vogginger, Tim Rosmeisl, Chen Liu 0031, Christian Mayr 0001
ISCAS3
2022 The operating system of the neuromorphic BrainScaleS-1 system
Eric Müller 0001, Christian Mauch, Sebastian Billaudelle, Andreas Grübl, Maurice Güttler, Dan Husmann de Oliveira, Joscha Ilmberger, Sebastian Jeltsch, Jakob Kaiser, Johann Klähn, Mitja Kleider, Christoph Koke, José Montes, Paul Müller 0002, Johannes Partzsch, Felix Passenberg, Hartmut Schmidt, Bernhard Vogginger, Jonas Weidner, Christian Mayr 0001, Johannes Schemmel
Neurocomputing19
2022 Time-Coded Spiking Fourier Transform in Neuromorphic Hardware
abstract
After several decades of continuously optimizing computing systems, the Moore's law is reaching its end. However, there is an increasing demand for fast and efficient processing systems that can handle large streams of data while decreasing system footprints. Neuromorphic computing answers this need by creating decentralized architectures that communicate with binary events over time. Despite its rapid growth in the last few years, novel algorithms are needed that can leverage the potential of this emerging computing paradigm and can stimulate the design of advanced neuromorphic chips. In this work, we propose a time-based spiking neural network that is mathematically equivalent to the Fourier transform. We implemented the network in the neuromorphic chip Loihi and conducted experiments on five different real scenarios with an automotive frequency modulated continuous wave radar. Experimental results validate the algorithm, and we hope they prompt the design of ad hoc neuromorphic chips that can improve the efficiency of state-of-the-art digital signal processors and encourage research on neuromorphic computing for signal processing.
Javier López-Randulfe, Nico Reeb, Negin Karimi, Chen Liu 0031, Hector A. Gonzalez, Robin Dietrich, Bernhard Vogginger, Christian Mayr 0001, Alois C. Knoll
IEEE Trans. Computers7
2021 Ultra-High Compression of Twiddle Factor ROMs in Multi-Core DSP for FMCW Radars
abstract
The increasing density of Multiple-Input Multiple-Output (MIMO) arrays in imaging radars for the automotive industry demands highly parallel systems with low-footprint accelerators, which would enable the concurrent processing of a high number of virtual channels with a low-latency, and without a high area overhead. In this paper, we design, implement, and test multiple handcrafted compression schemes for Twiddle Factor (TF) Read-Only Memories (ROM), aiming to reduce the footprint of a variable-length and dual-radix Fast Fourier Transform (FFT) accelerator in a Multi-core Digital Signal Processor (DSP) for Frequency Modulated Continuous Wave (FMCW) radars. The compression schemes proposed in this paper involve double delta encoding, Radix-specific address optimizations per port, symmetry inclusion, and exploitation of the bit resolution changes within the radar processing chain. All schemes are verified in an FPGA in terms of logic utilization and quantization using a 77-GHz radar, and implemented in a RISCV-based Processing Element (PE) of a Multi-core DSP with an Adaptive Body Bias (ABB) approach in 22FDX technology for assessing area, leakage, and relative latency savings when compared with a dual-ROM equivalent in the state-of-the-art.
Hector A. Gonzalez, Florian Kelber, Marco Stolba, Chen Liu 0031, Bernhard Vogginger, Stefan Hänzsche, Stefan Scholze, Sebastian Höppner, Christian Mayr 0001
ISCAS5
2017 Neuromorphic hardware in the loop: Training a deep spiking network on the BrainScaleS wafer-scale system
abstract
Emulating spiking neural networks on analog neuromorphic hardware offers several advantages over simulating them on conventional computers, particularly in terms of speed and energy consumption. However, this usually comes at the cost of reduced control over the dynamics of the emulated networks. In this paper, we demonstrate how iterative training of a hardware-emulated network can compensate for anomalies induced by the analog substrate. We first convert a deep neural network trained in software to a spiking network on the BrainScaleS wafer-scale neuromorphic system, thereby enabling an acceleration factor of 10000 compared to the biological time domain. This mapping is followed by the in-the-loop training, where in each training step, the network activity is first recorded in hardware and then used to compute the parameter updates in software via backpropagation. An essential finding is that the parameter updates do not have to be precise, but only need to approximately follow the correct gradient, which simplifies the computation of updates. Using this approach, after only several tens of iterations, the spiking network shows an accuracy close to the ideal software-emulated prototype. The presented techniques show that deep spiking networks emulated on analog neuromorphic devices can attain good computational performance despite the inherent variations of the analog substrate.
Johann Klähn, Guillaume Bellec, Andreas Grübl, Maurice Güttler, Andreas Hartel, Stephan Hartmann 0002, Dan Husmann de Oliveira, Kai Husmann, Sebastian Jeltsch, Vitali Karasenko, Mitja Kleider, Christoph Koke, Alexander Kononov, Christian Mauch, Eric Müller 0001, Paul Müller 0002, Johannes Partzsch, Mihai A. Petrovici, Stefan Schiefer, Stefan Scholze, Vasilis N. Thanasoulis, Bernhard Vogginger, Robert Legenstein, Wolfgang Maass 0001, Christian Mayr 0001, René Schüffny, Johannes Schemmel, Karlheinz Meier
IJCNN23
2017 Live demonstration: Dynamic voltage and frequency scaling for neuromorphic many-core systems
abstract
We present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation.
Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Prateek Joshi, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber
ISCAS3
2017 Dynamic voltage and frequency scaling for neuromorphic many-core systems
abstract
We present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; the effectiveness of the power management technique is demonstrated using a standard benchmark from the application domain. The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation.
Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber
ISCAS3
2017 Pattern representation and recognition with accelerated analog neuromorphic systems
abstract
Despite being originally inspired by the central nervous system, artificial neural networks have diverged from their biological archetypes as they have been remodeled to fit, particular tasks. In this paper, we review several possibilites to reverse map these architectures to biologically more realistic spiking networks with the aim of emulating them on fast, low-power neuromorphic hardware. Since many of these devices employ analog components, which cannot, be perfectly controlled, finding ways to compensate for the resulting effects represents a key challenge. Here, we discuss three different, strategies to address this problem: the addition of auxiliary network components for stabilizing activity, the utilization of inherently robust, architectures and a training method for hardware-emulated networks that, functions without, perfect, knowledge of the system's dynamics and parameters. For all three scenarios, we corroborate our theoretical considerations with experimental results on accelerated analog neuromorphic platforms.
Mihai A. Petrovici, Johann Klähn, Robert D. St. Louis, Anna Schroeder, Guillaume Bellec, Johannes Bill, Oliver Breitwieser, Ilja Bytschok, Andreas Grübl, Maurice Güttler, Andreas Hartel, Stephan Hartmann 0002, Dan Husmann de Oliveira, Kai Husmann, Sebastian Jeltsch, Vitali Karasenko, Mitja Kleider, Christoph Koke, Alexander Kononov, Christian Mauch, Eric Müller 0001, Paul Müller 0002, Johannes Partzsch, Thomas Pfeil, Stefan Schiefer, Stefan Scholze, Anand Subramoney, Vasilis N. Thanasoulis, Bernhard Vogginger, Robert Legenstein, Wolfgang Maass 0001, René Schüffny, Christian Mayr 0001, Johannes Schemmel, Karlheinz Meier
ISCAS30
2014 A pulse communication flow ready for accelerated neuromorphic experiments
abstract
Large-scale neuromorphic systems demand a sophisticated communication infrastructure to support several functionalities like configuration of neuron-and-synapse blocks, pulse stimulation, routing and tracing of the neural activity. This infrastructure is usually implemented around custom-designed FPGA systems. Performance requirements for these systems significantly increase when emulating the biological counterpart at an accelerated timescale. In this paper we characterize the capability of such a system to provide accurate long-term stimulation for emulated spiking networks and tracing of their activity. The design has a capacity of 1 billion timestamped events for both stimulation and tracing with a time resolution up to 48 ns. We characterize the communication flow in terms of throughput, transmission delay, packet loss and jitter. The results show that the implementation meets the needs of learning experiments, which is an important issue for state-of-the-art neuromorphic systems.
Vasilis N. Thanasoulis, Bernhard Vogginger, Johannes Partzsch, René Schüffny
ISCAS2