Stefan Scholze

dblp:35/9749 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 77% Interconnection networks and networks-on-chip · 16% Memory systems · 7%
Computer networks
1 paper
Physical-layer communications · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
database accelerator
0.212016
An MPSoC for energy-efficient database query processing · DAC 2016
Hardware accelerators and domain-specific architectures › query processing
energy-efficient query processing
0.212016
An MPSoC for energy-efficient database query processing · DAC 2016
Physical-layer communications
MIMO
0.112017
A Heterogeneous SDR MPSoC in 28 nm CMOS for Low-Latency Wireless Applications · DAC 2017
Memory systems
processing-in-memory
0.112016
An MPSoC for energy-efficient database query processing · DAC 2016

Methods — techniques the papers use, named apart from their topics

heterogeneous MPSoC design · 0.6dynamic data flow mapping · 0.6runtime task scheduling · 0.2instruction set extension · 0.2dynamic voltage and frequency scaling · 0.2
YearPublicationVenuePosition
2021 Ultra-High Compression of Twiddle Factor ROMs in Multi-Core DSP for FMCW Radars
abstract
The increasing density of Multiple-Input Multiple-Output (MIMO) arrays in imaging radars for the automotive industry demands highly parallel systems with low-footprint accelerators, which would enable the concurrent processing of a high number of virtual channels with a low-latency, and without a high area overhead. In this paper, we design, implement, and test multiple handcrafted compression schemes for Twiddle Factor (TF) Read-Only Memories (ROM), aiming to reduce the footprint of a variable-length and dual-radix Fast Fourier Transform (FFT) accelerator in a Multi-core Digital Signal Processor (DSP) for Frequency Modulated Continuous Wave (FMCW) radars. The compression schemes proposed in this paper involve double delta encoding, Radix-specific address optimizations per port, symmetry inclusion, and exploitation of the bit resolution changes within the radar processing chain. All schemes are verified in an FPGA in terms of logic utilization and quantization using a 77-GHz radar, and implemented in a RISCV-based Processing Element (PE) of a Multi-core DSP with an Adaptive Body Bias (ABB) approach in 22FDX technology for assessing area, leakage, and relative latency savings when compared with a dual-ROM equivalent in the state-of-the-art.
Hector A. Gonzalez, Florian Kelber, Marco Stolba, Chen Liu 0031, Bernhard Vogginger, Stefan Hänzsche, Stefan Scholze, Sebastian Höppner, Christian Mayr 0001
ISCAS7
2019 Performance Analysis of a Comparator Based Mixed-Signal Control Loop in 28 nm CMOS
abstract
In differential signaling systems using copper wires common mode signals are the cause of emission of electromagnetic energy. Especially in Automotive Ethernet systems this is a challenging problem. Beside classical passive components like common mode chokes active circuits can help to reduce the emission. This allows inexpensive and resource-conserving unshielded twisted pair cables to be used. This paper shows the approach of using a mixed-signal control loop based on a comparator and a 8 bit DAC for regulating the common mode voltage of an Automotive Ethernet DAC in 28 nm CMOS. An attenuation for interferers with frequencies up to 500 kHz is achieved and reaches up to 15 dB at maximum. The control loop utilizes the successive approximation algorithm commonly used for delay locked loops and DC trimming in mixed-signal circuits. In contrast to known applications the performance and usability at higher frequencies is considered in this paper. Being a nonlinear, time-variant system an analytical design of the control loop is very difficult. Therefore parametrical measurements show the dependency of frequency, amplitude and signal form of an applied common mode interferer source.
Florian Protze, Martin Kreißig, Frank Ellinger, Sebastian Höppner, Stephan Hartmann 0002, Stefan Hänzsche, Stefan Scholze, Georg Ellguth, Christian Mayr 0001
VLSI-SoC7
2018 Approximate Fixed-Point Elementary Function Accelerator for the SpiNNaker-2 Neuromorphic Chip
abstract
Neuromorphic chips are used to model biologically inspired Spiking-Neural-Networks (SNNs) where most models are based on differential equations. Equations for most SNN algorithms usually contain variables with one or more excomponents. SpiNNaker is a digital neuromorphic chip that has so far been using pre-calculated look-up tables for exponential function. However this approach is limited because the memory requirements grow as more complex neural models are developed. To save already limited memory resources in the next generation SpiNNaker chip, we are including a fast exponential function in the silicon. In this paper we analyse iterative algorithms for elementary functions and show how to build a single hardware accelerator for exp and natural log, for a neuromorphic chip prototype, to be manufactured in a 22 nm FDSOI process. We present the accelerator that has algorithmic level approximation control, allowing it to trade precision for latency and energy efficiency. As an addition to neuromorphic chip application, we provide analysis of a parameterized elementary function unit that can be tailored for other systems with different power, area, accuracy and latency constraints.
Mantas Mikaitis, David R. Lester, Delong Shang, Steve Furber, Gengting Liu, Jim D. Garside, Stefan Scholze, Sebastian Höppner, Andreas Dixius
ARITH7
2017 A Heterogeneous SDR MPSoC in 28 nm CMOS for Low-Latency Wireless Applications
abstract
Current and future applications impose high demands on software-defined radio (SDR) platforms in terms of latency, reliability, and flexibility. This paper presents a heterogeneous SDR MPSoC with a hexagonal network-on-chip to address these issues. It features four data processing modules and a baseband processing engine for iterative multiple-input multiple-output (MIMO) receiving. Integrated memory controllers enable dynamic data flow mapping and application isolation. In a 4 x 4 MIMO application scenario, the MPSoC achieves a throughput of 232 Mbit/s with a latency of 20 μs while consuming 414 mW. It outperforms state-of-the-art platforms in terms of throughput by a factor of 4.
Sebastian Haas, Tobias Seifert, Benedikt Noethen, Stefan Scholze, Sebastian Höppner, Andreas Dixius, Esther P. Adeva, Thomas R. Augustin, Friedrich Pauls, Sadia Moriam, Mattis Hasler, Erik Fischer, Yong Chen 0014, Emil Matús, Georg Ellguth, Stephan Hartmann 0002, Stefan Schiefer, Love Cederstroem, Dennis Walter, Stephan Henker, Stefan Hänzsche, Johannes Uhlig, Holger Eisenreich, Stefan Weithoffer, Norbert Wehn, René Schüffny, Christian Mayr 0001, Gerhard P. Fettweis
DAC4
2017 Neuromorphic hardware in the loop: Training a deep spiking network on the BrainScaleS wafer-scale system
abstract
Emulating spiking neural networks on analog neuromorphic hardware offers several advantages over simulating them on conventional computers, particularly in terms of speed and energy consumption. However, this usually comes at the cost of reduced control over the dynamics of the emulated networks. In this paper, we demonstrate how iterative training of a hardware-emulated network can compensate for anomalies induced by the analog substrate. We first convert a deep neural network trained in software to a spiking network on the BrainScaleS wafer-scale neuromorphic system, thereby enabling an acceleration factor of 10000 compared to the biological time domain. This mapping is followed by the in-the-loop training, where in each training step, the network activity is first recorded in hardware and then used to compute the parameter updates in software via backpropagation. An essential finding is that the parameter updates do not have to be precise, but only need to approximately follow the correct gradient, which simplifies the computation of updates. Using this approach, after only several tens of iterations, the spiking network shows an accuracy close to the ideal software-emulated prototype. The presented techniques show that deep spiking networks emulated on analog neuromorphic devices can attain good computational performance despite the inherent variations of the analog substrate.
Johann Klähn, Guillaume Bellec, Andreas Grübl, Maurice Güttler, Andreas Hartel, Stephan Hartmann 0002, Dan Husmann de Oliveira, Kai Husmann, Sebastian Jeltsch, Vitali Karasenko, Mitja Kleider, Christoph Koke, Alexander Kononov, Christian Mauch, Eric Müller 0001, Paul Müller 0002, Johannes Partzsch, Mihai A. Petrovici, Stefan Schiefer, Stefan Scholze, Vasilis N. Thanasoulis, Bernhard Vogginger, Robert Legenstein, Wolfgang Maass 0001, Christian Mayr 0001, René Schüffny, Johannes Schemmel, Karlheinz Meier
IJCNN21
2017 Live demonstration: Dynamic voltage and frequency scaling for neuromorphic many-core systems
abstract
We present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation.
Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Prateek Joshi, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber
ISCAS10
2017 Dynamic voltage and frequency scaling for neuromorphic many-core systems
abstract
We present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; the effectiveness of the power management technique is demonstrated using a standard benchmark from the application domain. The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation.
Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber
ISCAS9
2017 Pattern representation and recognition with accelerated analog neuromorphic systems
abstract
Despite being originally inspired by the central nervous system, artificial neural networks have diverged from their biological archetypes as they have been remodeled to fit, particular tasks. In this paper, we review several possibilites to reverse map these architectures to biologically more realistic spiking networks with the aim of emulating them on fast, low-power neuromorphic hardware. Since many of these devices employ analog components, which cannot, be perfectly controlled, finding ways to compensate for the resulting effects represents a key challenge. Here, we discuss three different, strategies to address this problem: the addition of auxiliary network components for stabilizing activity, the utilization of inherently robust, architectures and a training method for hardware-emulated networks that, functions without, perfect, knowledge of the system's dynamics and parameters. For all three scenarios, we corroborate our theoretical considerations with experimental results on accelerated analog neuromorphic platforms.
Mihai A. Petrovici, Johann Klähn, Robert D. St. Louis, Anna Schroeder, Guillaume Bellec, Johannes Bill, Oliver Breitwieser, Ilja Bytschok, Andreas Grübl, Maurice Güttler, Andreas Hartel, Stephan Hartmann 0002, Dan Husmann de Oliveira, Kai Husmann, Sebastian Jeltsch, Vitali Karasenko, Mitja Kleider, Christoph Koke, Alexander Kononov, Christian Mauch, Eric Müller 0001, Paul Müller 0002, Johannes Partzsch, Thomas Pfeil, Stefan Schiefer, Stefan Scholze, Anand Subramoney, Vasilis N. Thanasoulis, Bernhard Vogginger, Robert Legenstein, Wolfgang Maass 0001, René Schüffny, Christian Mayr 0001, Johannes Schemmel, Karlheinz Meier
ISCAS27
2016 An MPSoC for energy-efficient database query processing
abstract
This paper presents a heterogeneous database hardware accelerator MPSoC manufactured in 28 nm SLP CMOS. The 18 mm2 chip integrates a runtime task scheduling unit for energy-efficient query processing and hierarchical power management supported by an ultra-fast dynamic voltage and frequency scaling. Four processing elements, connected by a star-mesh network-on-chip, are accelerated by an instruction set extension tailored to fundamental data-intensive applications. We evaluate the MPSoC with typical database benchmarks focusing on scans and bitmap operations. When the processing elements operate on data stored in local memories, the chip consumes 250 mW and shows a 96x energy efficiency improvement compared to state-of-the-art platforms.
Sebastian Haas, Oliver Arnold, Benedikt Noethen, Stefan Scholze, Georg Ellguth, Andreas Dixius, Sebastian Höppner, Stefan Schiefer, Stephan Hartmann 0002, Stephan Henker, Thomas Hocker, Jörg Schreiter, Holger Eisenreich, Jens-Uwe Schluessler, Dennis Walter, Tobias Seifert, Friedrich Pauls, Mattis Hasler, Yong Chen 0014, Hermann Hensel, Sadia Moriam, Emil Matús, Christian Mayr 0001, René Schüffny, Gerhard P. Fettweis
DAC4
2015 An all-digital PWM generator with 62.5ps resolution in 28nm CMOS technology
abstract
This paper presents an all-digital pulse width modulator (PWM) for application in integrated DC-DC converters. Based on a multi-phase clock signal a PWM resolution of 62.5ps is achieved, resulting in up to 16-Bit PWM resolution at 4096ns period. The PWM signal duty cycle and period can be arbitrarily changed within a single cycle which allow efficient alldigital implementation of spread spectrum clocking schemes. The circuit has been implemented in 28nm SLP CMOS technology. It consumes 0.3mW when operating from a 1.0V supply.
Sebastian Höppner, Stefan Hänzsche, Stefan Scholze, René Schüffny
ISCAS3
2012 Live demonstration: A scaled-down version of the BrainScaleS wafer-scale neuromorphic system
abstract
This demonstration is based on the wafer-scale neuromophic system presented in the previous papers by Schemmel et. al. (20120), Scholze et. al. (2011) and Millner et. al. (2010). The demonstration setup will allow the visitors to monitor and partially manipulate the neural events at every level. They will get an insight into the complex interplay between packet-based and realtime communication necessary to combine continuous-time mixed-signal neural networks with a packet-based transport network. Several network experiments implemented on the setup will be accessible for user interaction.
Johannes Schemmel, Andreas Grübl, Stephan Hartmann 0002, Alexander Kononov, Christian Mayr 0001, Karlheinz Meier, Sebastian Millner, Johannes Partzsch, Stefan Schiefer, Stefan Scholze, René Schüffny, Marc-Olivier Schwartz
ISCAS10
2012 A 32 GBit/s communication SoC for a waferscale neuromorphic system
Stefan Scholze, Holger Eisenreich, Sebastian Höppner, Georg Ellguth, Stephan Henker, Mario Ander, Stefan Hänzsche, Johannes Partzsch, Christian Mayr 0001, René Schüffny
Integr.1
2011 Live demonstration: Packet-based AER with 3Gevent/s cumulative throughput
abstract
Traditionally, the communication in neuromorphic VLSI systems has been done via parallel asynchronous transmission of Address-Event-Representations (AER) of neuron pulses. Recently, there has been a move towards greater event transmission speed via a serialization of the AER protocols. We give a live demonstration of a packet based synchronous serial AER infrastructure presented in a recent paper, which handles the complete off-wafer communication and configuration for a newly developed waferscale neuromorphic system, operating at a factor of 104faster than biological real-time. Pulse packets are routed from the host PC via Gbit Ethernet to an FPGA board, which forwards them to 4 purpose designed Digital Network ASICs (DNCs) on the same board. The DNCs buffer and sort the pulses, implementing 32 2GBit/s Low Voltage Differential Signaling (LVDS) interfaces to the neuromorphic circuits on the wafer. Pulse communication to other wafers is done via an FPGA- FPGA communication using 10 Gbit/s Aurora links.
Stefan Schiefer, Stephan Hartmann 0002, Stefan Scholze, Johannes Partzsch, Christian Mayr 0001, Stephan Henker, René Schüffny
ISCAS3