EDBT 2026 Demo / reviewers in the wild / expert
Christian Mayr 0001
dblp:44/6754 · also Christian Georg Mayr
· DBLP profile ↗
46ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0003-3502-0872ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A variational framework for local learning with probabilistic latent representationsabstractWe introduce a novel method for distributed learning by dividing deep neural networks into blocks and incorporating feedback networks to propagate target information backwards, enabling auxiliary local losses.Forward and backward propagation operate in parallel with independent weights, addressing locking and weight transport problems.Our approach is rooted in a statistical view of training, treating block output activations as parameters of probability distributions to measure alignment between forward and backward passes.Error backpropagation is then performed locally within blocks, hence block-local learning.Preliminary results across tasks and architectures showcase state-of-the-art performance, establishing a principled framework for asynchronous distributed learning.* CTF and KKN are Cabrel Teguemne Fokam, Khaleelulla Khan Nazeer, Christian Mayr 0001, Anand Subramoney, David Kappel |
ESANN | 3 |
| 2024 | Capturing Uncertainty over Time for Spiking Neural Networks by Exploiting Conformal Prediction SetsabstractThere is a great interest in harnessing the advantages of spiking neural networks. An increasing portion of research is focusing on the deployment of such models. The problem of safe decision making has similarities with classical networks. We apply spiking neural networks to time-series classification tasks where their stateful nature is beneficial. We show that the well-known method of Conformal Prediction (CP) is capable of distinguishing between wrong and correct decisions in this setting similar to but while being less expensive than Evidential Deep Learning and Neural Network Ensembles. In this work we argue that classification uncertainty in time should additionally be considered but is not captured by the length of prediction sets output from CP. Our main contribution addresses the issue that existing CP methods for classification do not consider the aforementioned problem. Our method takes as input the prediction sets which can be output from present conformal prediction and then extends these methods by a smoothed length and combined set algorithm. We apply our method to spiking neural network-based classifiers trained on four different time-series datasets. We show that our method outputs a more suitable uncertainty metric at a given point in time than just the unmodified set length of CP for classification. Daniel Scholz 0002, Oliver Emonds, Felix Kreutz, Pascal Gerhards, Jiaxin Huang 0007, Klaus Knobloch, Alois C. Knoll, Christian Mayr 0001 |
ICMLA | 8 |
| 2024 | A Low-footprint FFT Accelerator for a RISC-V-based Multi-core DSP in FMCW RadarsabstractMulti-core systems are required by digital signal processors (DSP) to support the revolutionary Multiple-Input Multiple-Output (MIMO) imaging radars in the automotive industry. Such multi-core processors for Frequency Modulated Continuous Wave (FMCW) radars require the use of low- footprint accelerators that would reduce the overhead as the system scales up with the antenna density. In this paper, we propose an FFT accelerator, named RbFFT, optimized for the MIMO radar processing chain. The architecture of RbFFT reduces the overhead by re-using existing memory in the processing element (PE), and employs a dual-radix butterfly engine with mixed bit resolution to optimize resources in dense radars. RbFFT reduces area by implementing for the first time ultra-low compression in its dual twiddle factor ROM. RbFFT also innovates with custom fetching and buffering strategies to improve memory-based FFTs while reusing logic to integrate reverse bit ordering, windowing and inverse FFT (IFFT) within the same accelerator passes. The proposed accelerator is implemented in a 25-Core Smart MPSoC in 22FDX using Adaptive Body Biasing (ABB) at 0.6V. Besides RbFFT being pioneer in specialized FFT accelerators for dense MIMO systems, the results also show state-of-the-art improvements via 11% reduction in the normalized energy consumption, 4% reduction in latency, and 11 times area reduction with relation to previous silicon implementations. Hector A. Gonzalez, Marco Stolba, Bernhard Vogginger, Tim Rosmeisl, Chen Liu 0031, Christian Mayr 0001 |
ISCAS | 6 |
| 2024 | Automating application-driven customization of ASIPs: A survey
Eslam Hussein, Bernd Waschneck, Christian Mayr 0001 |
J. Syst. Archit. | 3 |
| 2023 | Efficient recurrent architectures through activity sparsity and sparse back-propagation through time
Anand Subramoney, Khaleelulla Khan Nazeer, Mark Schöne, Christian Mayr 0001, David Kappel |
ICLR | 4 |
| 2022 | Industry-track: Towards Agile Design of Neural Processing UnitabstractMore and more specialized processors, known as Neural Processing Units (NPUs), have been or are being built for deep neural network inference. Design and optimization of this kind of processor are inseparable from the deep learning ecosystem and corresponding underlying software. This HW/SW co-design requirement poses challenges for designers. Therefore, in this work, we experiment with an agile development method to shorten the development cycles of NPUs. We utilize Chisel for hardware design and develop a custom Chisel backend for generating cycle-accurate simulators with C++/Python APIs. On top of the simulator, we built a Python software stack for software development, performance evaluation, and simulation-based verification. The proposed method is purely software and does not involve real hardware, thus allowing the integration of software agile development methods into digital designs. In the experiments, we show how it helps us identify inherent hardware limitations and how it shortens our development cycles. Binyi Wu, Wolfgang Furtner, Bernd Waschneck, Christian Mayr 0001 |
CODES+ISSS | 4 |
| 2022 | The Scale4Edge RISC-V EcosystemabstractThis paper introduces the project Scale4Edge. The project is focused on enabling an effective RISC-V ecosystem for optimization of edge applications. We describe the basic components of this ecosystem and introduce the envisioned demonstrators, which will be used in their evaluation. Wolfgang Ecker, Peer Adelt, Wolfgang Müller 0003, Reinhold Heckmann, Milos Krstic, Vladimir Herdt, Rolf Drechsler, Gerhard Angst, Ralf Wimmer 0001, Andreas Mauderer, Rafael Stahl, Karsten Emrich, Daniel Mueller-Gritschneder, Bernd Becker 0001, Philipp M. Scholl, Eyck Jentzsch, Jan Schlamelcher, Kim Grüttner, Paul Palomero Bernardo, Oliver Bringmann 0001, Brindusa Mihaela Damian-Kosterhon, Julian Oppermann, Andreas Koch 0001, Jörg Bormann, Johannes Partzsch, Christian Mayr 0001, Wolfgang Kunz |
DATE | 26 |
| 2022 | Neural Architecture Search for Low-Precision Neural Networks
Binyi Wu, Bernd Waschneck, Christian Mayr 0001 |
ICANN (4) | 3 |
| 2022 | Prototyping of Low-Cost Configurable Sparse Neural Processing Unit with Buffer and Mixed-Precision Reshapeable MAC ArrayabstractMore recently, it has become possible to run deep learning algorithms on edge devices such as microcontrollers due to continuous improvements in neural network optimization algorithms such as quantization and neural architecture search. Nonetheless, most of the embedded hardware available today still falls short of the requirements of running deep neural networks. As a result, specialized processors have emerged to improve the inference efficiency of deep learning algorithms. However, most are not for edge applications that require efficient and low-cost hardware. Therefore, we design and prototype a low-cost configurable sparse Neural Processing Unit (NPU). The NPU has a built-in buffer and a reshapable mixed-precision multiply-accumulator (MAC) array. The computing and memory resources of the NPU are parameterized, and different NPUs can be derived. Besides, users can also conFigure the NPU at runtime to fully utilize the resources. In our experiments, the 200MHz NPU with only 32 MACs is more than 32 times faster than the 400MHzSTM32H7 when inferring MobileNet-Vl. Besides, the yielded NPUs can achieve roofline or even beyond roofline performance. The buffer and reshapeable MAC array push the NPU’s attainable performance to the roofline, while the feature of supporting sparsity allows the NPU to obtain performance beyond the roofline. Binyi Wu, Wolfgang Furtner, Bernd Waschneck, Christian Mayr 0001 |
ICPADS | 4 |
| 2022 | Hardware-Efficient Ultrasonic Entrance Counting: Comparing Different Machine Learning ApproachesabstractIn this work, the classification of walking direction based on ultrasonic signals has been examined for entrance counting. Feed-forward and recurrent neural network architectures as well as simpler machine learning techniques have been investigated and compared with classical signal processing techniques.Using only a single ultrasonic receiver, the focus was set on the development of a hardware-efficient system concept. Different ultrasonic measurement methods in time and frequency domain have been compared with the perspective of a holistic energy optimization. The analysis of the system’s hardware efficiency was completed by an estimation of algorithmic latency, energy and storage consumption based on the arithmetic of the classification algorithms. All algorithms showed an estimated energy consumption of less than 10 μJ for a single inference on a state-of-the-art implementation of an ARM® Cortex® M4F micro-controller, which was found to be negligible compared to the energy of the measurement principle. Compared to other sensor types and multi-sensor systems, a state-of-the-art test accuracy of 99.72% could be achieved for differentiating between the two entrance directions of a present person and the absence of a person. Tim Langer, Bernd Waschneck, Johannes Partzsch, Florian Kelber, Christian Mayr 0001 |
ICPR | 5 |
| 2022 | Convolutional Neural Networks Quantization with Double-Stage Squeeze-and-ThresholdabstractIt has been proven that, compared to using 32-bit floating-point numbers in the training phase, Deep Convolutional Neural Networks (DCNNs) can operate with low-precision during inference, thereby saving memory footprint and power consumption. However, neural network quantization is always accompanied by accuracy degradation. Here, we propose a quantization method called double-stage Squeeze-and-Threshold (double-stage ST) to close the accuracy gap with full-precision models. While accurate colors in pictures can be pleasing to the viewer, they are not necessary for distinguishing objects. The era of black and white television proves this idea. As long as the limited colors are filled reasonably for different objects, the objects can be well identified and distinguished. Our method utilizes the attention mechanism to adjust the activations and learn the thresholds to distinguish objects (features). We then divide the numerically rich activations into intervals (a limited variety of numerical values) by the learned thresholds. The proposed method supports both binarization and multi-bit quantization. Our method achieves state-of-the-art results. In binarization, ReActNet [Z. Liu, Z. Shen, S. Li, K. Helwegen, D. Huang and K. Cheng, arXiv:abs/2106.11309 ] trained with our method outperforms the previous state-of-the-art result by 0.2 percentage points. Whereas in multi-bit quantization, the top-1 accuracy of the 3-bit ResNet-18 [K. He, X. Zhang, S. Ren and J. Sun, Deep residual learning for image recognition, 2016 IEEE Conf. Computer Vision and Pattern Recognition, CVPR 2016, 27–30 June 2016, Las Vegas, NV, USA (IEEE Computer Society, 2016), pp. 770–778] model exceeds the top-1 accuracy of its full-precision baseline model by 0.4 percentage points. The double-stage ST activation quantization method is easy to apply by inserting it before the convolution. Besides, the double-stage ST is detachable after training and introducing no computational cost in inference. Binyi Wu, Bernd Waschneck, Christian Mayr 0001 |
Int. J. Neural Syst. | 3 |
| 2022 | The operating system of the neuromorphic BrainScaleS-1 system
Eric Müller 0001, Christian Mauch, Sebastian Billaudelle, Andreas Grübl, Maurice Güttler, Dan Husmann de Oliveira, Joscha Ilmberger, Sebastian Jeltsch, Jakob Kaiser, Johann Klähn, Mitja Kleider, Christoph Koke, José Montes, Paul Müller 0002, Johannes Partzsch, Felix Passenberg, Hartmut Schmidt, Bernhard Vogginger, Jonas Weidner, Christian Mayr 0001, Johannes Schemmel |
Neurocomputing | 21 |
| 2022 | Time-Coded Spiking Fourier Transform in Neuromorphic HardwareabstractAfter several decades of continuously optimizing computing systems, the Moore's law is reaching its end. However, there is an increasing demand for fast and efficient processing systems that can handle large streams of data while decreasing system footprints. Neuromorphic computing answers this need by creating decentralized architectures that communicate with binary events over time. Despite its rapid growth in the last few years, novel algorithms are needed that can leverage the potential of this emerging computing paradigm and can stimulate the design of advanced neuromorphic chips. In this work, we propose a time-based spiking neural network that is mathematically equivalent to the Fourier transform. We implemented the network in the neuromorphic chip Loihi and conducted experiments on five different real scenarios with an automotive frequency modulated continuous wave radar. Experimental results validate the algorithm, and we hope they prompt the design of ad hoc neuromorphic chips that can improve the efficiency of state-of-the-art digital signal processors and encourage research on neuromorphic computing for signal processing. Javier López-Randulfe, Nico Reeb, Negin Karimi, Chen Liu 0031, Hector A. Gonzalez, Robin Dietrich, Bernhard Vogginger, Christian Mayr 0001, Alois C. Knoll |
IEEE Trans. Computers | 8 |
| 2021 | Analyzing ARM CoreSight ETMv4.x Data Trace Stream with a Real-time Hardware AcceleratorabstractDebugging and verification of modern SoCs is a vital step in realizing complex systems consisting of various components. Monitoring memory operations such as data transfer address and value is an essential debugging and verification feature. ARM CoreSight technology generates a specific debug trace stream standard to monitor the memory without affecting the normal execution of the system. This paper proposes a hardware architecture to analyze the debug trace stream in realtime. It is implemented on the Xilinx Virtex xc6vcx75t-2ff784 FPGA device and can operate at 125 MHz and occupies less than 8 % of the FPGA resources. Seyed Mohammad Ali Zeinolabedin, Johannes Partzsch, Christian Mayr 0001 |
DATE | 3 |
| 2021 | Squeeze-and-Threshold Based Quantization for Low-Precision Neural Networks
Binyi Wu, Bernd Waschneck, Christian Mayr 0001 |
EANN | 3 |
| 2021 | Opportunities For A Hardware-Based OPC UA Server Implementation In Industry 4.0abstractWith the advent of the fourth industrial revolution i.e. Industry 4.0, plants and factories are becoming smarter and interconnected. The transitions demand vertical integration and seamless connectivity. For this purpose, there is a need for semantic communication between various devices including the heavily resource-constrained field devices. To address this, a real-time capable hardware-based implementation of a well-established semantic communication protocol, i.e. OPC Unified Architecture was designed and developed. This chip-based implementation is power-efficient and compact, making it suitable for the field level. The chip was analyzed and incorporated in a demonstrator as a proof of concept of its integration at field level in a plant module of the process industry. Various opportunities are also examined where the chip could be utilized to deliver benefits to existing and future technologies. Zohra Charania, Chris Paul Iatrou, Valentin Khaydarov, Richard Jacob, Robert Wittig, Heiner Bauer, Sebastian Höppner, René Bachmann, Philipp Bauer, Hendrik Deckert, Christian Mayr 0001, Gerhard P. Fettweis, Leon Urbas |
IECON | 11 |
| 2021 | Ultra-High Compression of Twiddle Factor ROMs in Multi-Core DSP for FMCW RadarsabstractThe increasing density of Multiple-Input Multiple-Output (MIMO) arrays in imaging radars for the automotive industry demands highly parallel systems with low-footprint accelerators, which would enable the concurrent processing of a high number of virtual channels with a low-latency, and without a high area overhead. In this paper, we design, implement, and test multiple handcrafted compression schemes for Twiddle Factor (TF) Read-Only Memories (ROM), aiming to reduce the footprint of a variable-length and dual-radix Fast Fourier Transform (FFT) accelerator in a Multi-core Digital Signal Processor (DSP) for Frequency Modulated Continuous Wave (FMCW) radars. The compression schemes proposed in this paper involve double delta encoding, Radix-specific address optimizations per port, symmetry inclusion, and exploitation of the bit resolution changes within the radar processing chain. All schemes are verified in an FPGA in terms of logic utilization and quantization using a 77-GHz radar, and implemented in a RISCV-based Processing Element (PE) of a Multi-core DSP with an Adaptive Body Bias (ABB) approach in 22FDX technology for assessing area, leakage, and relative latency savings when compared with a dual-ROM equivalent in the state-of-the-art. Hector A. Gonzalez, Florian Kelber, Marco Stolba, Chen Liu 0031, Bernhard Vogginger, Stefan Hänzsche, Stefan Scholze, Sebastian Höppner, Christian Mayr 0001 |
ISCAS | 9 |
| 2021 | Hardware Implementation of an OPC UA Server for Industrial Field DevicesabstractIndustrial plants suffer from a high degree of complexity and incompatibility in their communication infrastructure, caused by a wild mix of proprietary technologies. This prevents transformation toward Industry 4.0 and the Industrial Internet of Things. Open platform communications unified architecture (OPC UA) is a standardized protocol that addresses these problems with uniform and semantic communication across all levels of the hierarchy. However, its adoption in embedded field devices, such as sensors and actuators, is still lacking due to prohibitive memory and power requirements of software implementations. We have developed a dedicated hardware engine that offloads processing of the OPC UA protocol and enables the realization of compact and low-power field devices with OPC UA support. As part of a proof-of-concept embedded system, we have implemented this engine in a 22-nm FDSOI technology, representing the first ASIC implementation of an OPC UA server. We measured performance, power consumption, and memory footprint of our test chip and compared it with a software implementation based on open62541 and a Raspberry Pi 2B. Our OPC UA hardware engine is 50 times more energy efficient and only requires 36 KiB of memory. The complete system consumes only 24 mW under full load, making it suitable for low-power embedded applications. Heiner Bauer, Sebastian Höppner, Chris Paul Iatrou, Zohra Charania, Stephan Hartmann 0002, Saif-Ur Rehman, Andreas Dixius, Georg Ellguth, Dennis Walter, Johannes Uhlig, Felix Neumärker, Marc Berthel, Marco Stolba, Florian Kelber, Leon Urbas, Christian Mayr 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 16 |
| 2020 | Method for the Computer-Aided Schematic Design and Simulation of Hydrogel-Based Microfluidic SystemsabstractWe present an approach to the schematic design and time-dependent simulation of microfluidic systems that are composed of basic components, such as pumps, channels, mixing junctions, and valves. For the schematic capture and the simulation, the Cadence Virtuoso design framework is employed, with the components modeled by Verilog-AMS. The pressures and flow rates are calculated with the help of the hydraulic analogy between electronics and fluidics. Signal flow quantities are used to express the flow of chemical concentration. Our focus is on systems employing chemofluidic transistors, i.e., hydrogel-based microvalves that facilitate autonomous flow control without the need for an external control unit. Our approach also covers layout editing by the use of programmed cells. The utility of the approach is exemplified by the successful top-down design of a chemofluidic oscillator circuit and by the functionally correct simulation of a chemofluidic SR flip-flop. Andreas Voigt, Jörg Schreiter, Philipp Frank, Cesare Pini, Christian Mayr 0001, Andreas Richter 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | A Fast Lock-In Ultra Low-Voltage ADPLL Clock Generator with Adaptive Body Biasing in 22nm FDSOI TechnologyabstractSystems on Chip for the Internet of Things require fast-locking and robust clock generators to maximize the effectiveness of power management techniques such as Dynamic Voltage and Frequency Scaling and duty-cycling. We present an ADPLL clock generator based on a fully digital DCO architecture with an inherently linear and offset-free tuning characteristic that allows fast lock-in within three and frequency changes during operation within two reference cycles. Measurements from a testchip in 22nm FDSOI CMOS technology show operation from 0.4 to 0.8 V and 20 to 790 MHz. At 0.5 V, only 107 μW are consumed to generate a 100 MHz clock with 59 ps RMS period jitter. Adaptive Body Biasing improves the jitter performance by up to 40% through compensation of PVT variation. Florian Schraut, Holger Eisenreich, Sebastian Höppner, Christian Mayr 0001 |
ISCAS | 4 |
| 2019 | A Multi-Bit PFD Architecture for ADPLLs with Built-In Jitter Self-CalibrationabstractThis paper presents a multi-bit phase-frequency detector (PFD) architecture with self-calibration scheme to reduce the jitter in all-digital phase-locked loops. A standard bang-bang PFD is extended by two additional PFDs which allow the measurement of the ADPLL jitter distribution width. A built-in self-calibration algorithm can utilise this feature to optimise the loop filter gain for minimised overall jitter. The proposed technique is demonstrated in hardware within an LC-ADPLL in 28 nm SLP CMOS technology with 7.5 GHz output from a reference frequency of 100 MHz. By using the proposed PFD technique the total accumulated jitter is reduced by 20%. Franz Marcus Schüffny, Sebastian Höppner, Alexander Oefelein, Christian Mayr 0001 |
ISCAS | 4 |
| 2019 | Performance Analysis of a Comparator Based Mixed-Signal Control Loop in 28 nm CMOSabstractIn differential signaling systems using copper wires common mode signals are the cause of emission of electromagnetic energy. Especially in Automotive Ethernet systems this is a challenging problem. Beside classical passive components like common mode chokes active circuits can help to reduce the emission. This allows inexpensive and resource-conserving unshielded twisted pair cables to be used. This paper shows the approach of using a mixed-signal control loop based on a comparator and a 8 bit DAC for regulating the common mode voltage of an Automotive Ethernet DAC in 28 nm CMOS. An attenuation for interferers with frequencies up to 500 kHz is achieved and reaches up to 15 dB at maximum. The control loop utilizes the successive approximation algorithm commonly used for delay locked loops and DC trimming in mixed-signal circuits. In contrast to known applications the performance and usability at higher frequencies is considered in this paper. Being a nonlinear, time-variant system an analytical design of the control loop is very difficult. Therefore parametrical measurements show the dependency of frequency, amplitude and signal form of an applied common mode interferer source. Florian Protze, Martin Kreißig, Frank Ellinger, Sebastian Höppner, Stephan Hartmann 0002, Stefan Hänzsche, Stefan Scholze, Georg Ellguth, Christian Mayr 0001 |
VLSI-SoC | 9 |
| 2018 | A physical synthesis flow for early technology evaluation of silicon nanowire based reconfigurable FETsabstractSilicon Nanowire (SiNW) based reconfigurable field-effect transistors (RFETs) provide an additional gate terminal called the program gate which gives the freedom of programming p-type or n-type functionality for the same device at runtime. This enables the circuit designers to pack more functionality per computational unit. This saves processing costs as only one device type is required, and no doping and associated lithography steps are needed for this technology. In this paper, we present a complete design flow including both logic and physical synthesis for circuits based on SiNW RFETs. We propose layouts of logic gates, Liberty and LEF (Library Exchange Format) files to enable further research in the domain of these novel, functionally enhanced transistors. We show that in the first of its kind comparison, for these fully symmetrical reconfigurable transistors, the area after placement and routing for SiNW based circuits is 17% more than that of CMOS for MCNC benchmarks. Further, we discuss areas of improvement for obtaining better area results from the SiNW based RFETs from a fabrication and technology point of view. The future use of self-aligned techniques to structure two independent gates within a smaller pitch holds the promise of substantial area reduction. Shubham Rai, Ansh Rupani, Dennis Walter, Michael Raitza, Andre Heinzig, Tim Baldauf, Jens Trommer, Christian Mayr 0001, Walter M. Weber, Akash Kumar 0001 |
DATE | 8 |
| 2017 | A Heterogeneous SDR MPSoC in 28 nm CMOS for Low-Latency Wireless ApplicationsabstractCurrent and future applications impose high demands on software-defined radio (SDR) platforms in terms of latency, reliability, and flexibility. This paper presents a heterogeneous SDR MPSoC with a hexagonal network-on-chip to address these issues. It features four data processing modules and a baseband processing engine for iterative multiple-input multiple-output (MIMO) receiving. Integrated memory controllers enable dynamic data flow mapping and application isolation. In a 4 x 4 MIMO application scenario, the MPSoC achieves a throughput of 232 Mbit/s with a latency of 20 μs while consuming 414 mW. It outperforms state-of-the-art platforms in terms of throughput by a factor of 4. Sebastian Haas, Tobias Seifert, Benedikt Noethen, Stefan Scholze, Sebastian Höppner, Andreas Dixius, Esther P. Adeva, Thomas R. Augustin, Friedrich Pauls, Sadia Moriam, Mattis Hasler, Erik Fischer, Yong Chen 0014, Emil Matús, Georg Ellguth, Stephan Hartmann 0002, Stefan Schiefer, Love Cederstroem, Dennis Walter, Stephan Henker, Stefan Hänzsche, Johannes Uhlig, Holger Eisenreich, Stefan Weithoffer, Norbert Wehn, René Schüffny, Christian Mayr 0001, Gerhard P. Fettweis |
DAC | 27 |
| 2017 | Neuromorphic hardware in the loop: Training a deep spiking network on the BrainScaleS wafer-scale systemabstractEmulating spiking neural networks on analog neuromorphic hardware offers several advantages over simulating them on conventional computers, particularly in terms of speed and energy consumption. However, this usually comes at the cost of reduced control over the dynamics of the emulated networks. In this paper, we demonstrate how iterative training of a hardware-emulated network can compensate for anomalies induced by the analog substrate. We first convert a deep neural network trained in software to a spiking network on the BrainScaleS wafer-scale neuromorphic system, thereby enabling an acceleration factor of 10000 compared to the biological time domain. This mapping is followed by the in-the-loop training, where in each training step, the network activity is first recorded in hardware and then used to compute the parameter updates in software via backpropagation. An essential finding is that the parameter updates do not have to be precise, but only need to approximately follow the correct gradient, which simplifies the computation of updates. Using this approach, after only several tens of iterations, the spiking network shows an accuracy close to the ideal software-emulated prototype. The presented techniques show that deep spiking networks emulated on analog neuromorphic devices can attain good computational performance despite the inherent variations of the analog substrate. Johann Klähn, Guillaume Bellec, Andreas Grübl, Maurice Güttler, Andreas Hartel, Stephan Hartmann 0002, Dan Husmann de Oliveira, Kai Husmann, Sebastian Jeltsch, Vitali Karasenko, Mitja Kleider, Christoph Koke, Alexander Kononov, Christian Mauch, Eric Müller 0001, Paul Müller 0002, Johannes Partzsch, Mihai A. Petrovici, Stefan Schiefer, Stefan Scholze, Vasilis N. Thanasoulis, Bernhard Vogginger, Robert Legenstein, Wolfgang Maass 0001, Christian Mayr 0001, René Schüffny, Johannes Schemmel, Karlheinz Meier |
IJCNN | 26 |
| 2017 | Live demonstration: Dynamic voltage and frequency scaling for neuromorphic many-core systemsabstractWe present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation. Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Prateek Joshi, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber |
ISCAS | 14 |
| 2017 | Dynamic voltage and frequency scaling for neuromorphic many-core systemsabstractWe present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; the effectiveness of the power management technique is demonstrated using a standard benchmark from the application domain. The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation. Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber |
ISCAS | 13 |
| 2017 | A fixed point exponential function accelerator for a neuromorphic many-core systemabstractMany models of spiking neural networks heavily rely on exponential waveforms. On neuromorphic multiprocessor systems like SpiNNaker, they have to be approximated by dedicated algorithms, often dominating the processing load. Here we present a processor extension for fast calculation of exponentials, aimed at integration in the next-generation SpiNNaker system. Our implementation achieves single-LSB precision in a 32bit fixed-point format and 250Mexp/s throughput at 0.44nJ/exp for nominal supply (1.0V), or 0.21nJ/exp at 0.7V supply and 77Mexp/s, demonstrating a throughput multiplication of almost 50 and 98% energy reduction at 2% area overhead per processor on a 28nm CMOS chip. Johannes Partzsch, Sebastian Höppner, Matthias Eberlein, René Schüffny, Christian Mayr 0001, David R. Lester, Steve Furber |
ISCAS | 5 |
| 2017 | Pattern representation and recognition with accelerated analog neuromorphic systemsabstractDespite being originally inspired by the central nervous system, artificial neural networks have diverged from their biological archetypes as they have been remodeled to fit, particular tasks. In this paper, we review several possibilites to reverse map these architectures to biologically more realistic spiking networks with the aim of emulating them on fast, low-power neuromorphic hardware. Since many of these devices employ analog components, which cannot, be perfectly controlled, finding ways to compensate for the resulting effects represents a key challenge. Here, we discuss three different, strategies to address this problem: the addition of auxiliary network components for stabilizing activity, the utilization of inherently robust, architectures and a training method for hardware-emulated networks that, functions without, perfect, knowledge of the system's dynamics and parameters. For all three scenarios, we corroborate our theoretical considerations with experimental results on accelerated analog neuromorphic platforms. Mihai A. Petrovici, Johann Klähn, Robert D. St. Louis, Anna Schroeder, Guillaume Bellec, Johannes Bill, Oliver Breitwieser, Ilja Bytschok, Andreas Grübl, Maurice Güttler, Andreas Hartel, Stephan Hartmann 0002, Dan Husmann de Oliveira, Kai Husmann, Sebastian Jeltsch, Vitali Karasenko, Mitja Kleider, Christoph Koke, Alexander Kononov, Christian Mauch, Eric Müller 0001, Paul Müller 0002, Johannes Partzsch, Thomas Pfeil, Stefan Schiefer, Stefan Scholze, Anand Subramoney, Vasilis N. Thanasoulis, Bernhard Vogginger, Robert Legenstein, Wolfgang Maass 0001, René Schüffny, Christian Mayr 0001, Johannes Schemmel, Karlheinz Meier |
ISCAS | 34 |
| 2016 | An MPSoC for energy-efficient database query processingabstractThis paper presents a heterogeneous database hardware accelerator MPSoC manufactured in 28 nm SLP CMOS. The 18 mm2 chip integrates a runtime task scheduling unit for energy-efficient query processing and hierarchical power management supported by an ultra-fast dynamic voltage and frequency scaling. Four processing elements, connected by a star-mesh network-on-chip, are accelerated by an instruction set extension tailored to fundamental data-intensive applications. We evaluate the MPSoC with typical database benchmarks focusing on scans and bitmap operations. When the processing elements operate on data stored in local memories, the chip consumes 250 mW and shows a 96x energy efficiency improvement compared to state-of-the-art platforms. Sebastian Haas, Oliver Arnold, Benedikt Noethen, Stefan Scholze, Georg Ellguth, Andreas Dixius, Sebastian Höppner, Stefan Schiefer, Stephan Hartmann 0002, Stephan Henker, Thomas Hocker, Jörg Schreiter, Holger Eisenreich, Jens-Uwe Schluessler, Dennis Walter, Tobias Seifert, Friedrich Pauls, Mattis Hasler, Yong Chen 0014, Hermann Hensel, Sadia Moriam, Emil Matús, Christian Mayr 0001, René Schüffny, Gerhard P. Fettweis |
DAC | 23 |
| 2016 | Beyond spike-timing dependent plasticity in memristor crossbar arraysabstractMemristors have emerged as promising, area-efficient, nano-scale devices for implementing models of synaptic plasticity in hybrid CMOS-memristor neuromorphic architectures. These architectures aim at reproducing the learning capabilities of biological networks by emulating the complex dynamics of biological neurons and synapses. However, to maximize the density of these elements in crossbar arrays, learning circuits have often been limited to the implementation of simple spike timing-dependent plasticity (STDP) mechanisms. We propose novel hybrid CMOS-memristor circuits that reproduce more effective and realistic plasticity rules which depend on the timing of the pre-synaptic input spike and on the state of the post-synaptic neuron, and which allow the integration of dense crossbar memristor arrays. To implement these plasticity rules in memristor crossbar arrays, the circuits driving the memristors' post-synaptic terminals actively sense the activity on the pre-synaptic terminals to apply the appropriate stimulation waveforms across the memristors. We illustrate the advantages of this scheme by using it to implement a spike-based perceptron plasticity r ule. Hesham Mostafa, Christian Mayr 0001, Giacomo Indiveri |
ISCAS | 2 |
| 2016 | A Calibration Technique for Bang-Bang ADPLLs Using Jitter Distribution MonitoringabstractThis brief presents a built-in self-calibration (BISC) technique for minimization of the total jitter in bang-bang all-digital phase-locked loops (ADPLLs). It is based on the addition of a monitoring phase-frequency detector (PFD) with tunable delay cells for the reference clock and the divider clock and a counter for this PFD output signal. This allows for on-chip binary comparison of the jitter distribution widths at the ADPLL PFD input, when ADPLL filter parameters are altered. Since only a relative comparison is performed, no accurate delay calibration is required. The statistical properties of this comparison of two random distributions are analyzed theoretically, and guidelines for circuit dimensioning are derived. The proposed method is used for BISC by adaption of the ADPLL filter coefficients. This allows for jitter minimization under process, voltage and temperature variations as well as gain and period jitter of the digitally controlled oscillator. The proposed calibration technique is verified by system simulations and measurements of a silicon prototype implementation in 28-nm CMOS technology. Sebastian Höppner, Johannes Partzsch, Johannes Neumann, René Schüffny, Christian Mayr 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2014 | VLSI implementation of a conductance-based multi-synapse using switched-capacitor circuitsabstractFor neuromorphic ICs, the implemented synaptic dynamics play an important role in the complexity achievable when running networks on the overall IC. One of these ingredients for realistic dynamics are conductance-based synapses, which in contrast to current-based synapses let a neuron adapt in various ways to its input characteristics. Another ingredient is classical neuronal spike-frequency adaptation. Both are usually realized in fully-analog subthreshold circuits, making them hard to port to modern sub-100nm technologies. In contrast, we present a compact switched-capacitor (SC) model of a conductance-based synapse that can be widely configured to accurately depict e.g. NMDA, GABA or AMPA type synapses. The SC approach is inherently easy to port between technologies and its digital part benefits fully from technology scaling. We show how this synapse circuit can also be utilized to endow a neuron with spike-frequency adaptation (SFA). Marko Noack, Markus Krause, Christian Mayr 0001, Johannes Partzsch, René Schüffny |
ISCAS | 3 |
| 2013 | A location-independent direct link neuromorphic interfaceabstractWith neuromorphic hardware rapidly moving towards large-scale, possibly immovable systems capable of implementing brain-scale neural models in hardware, there is an emerging need to be able to integrate multi-system combinations of sensors and cortical processors over distributed, multisite configurations. If there were a standard, direct interface allowing large systems to communicate using native signalling, it would be possible to use heterogeneous resources efficiently according to their task suitability. We propose a UDP-based AER spiking interface that permits direct bidirectional spike communications over standard networks, and demonstrate a practical implementation with two large-scale neuromorphic systems, BrainScaleS and SpiNNaker. Internally, the interfaces at either end appear as interceptors which decode and encode spikes in a standardised AER address format onto UDP frames. The system is able to run a spiking neural network distributed over the two systems, in both a side-by-side setup with a direct cable link and over the Internet between 2 widely spaced sites. Such a model not only realises a solution for connecting remote sensors or processors to a large, central neuromorphic simulation platform, but also opens possibilities for interesting automated remote neural control, such as parameter tuning, for large, complex neural systems, and suggests methods to overcome differences in timescale and simulation model between different platforms. With its entirely standard protocol and physical layer, the interface makes large neuromorphic systems a distributed, accessible resource available to all. Alex Rast, Johannes Partzsch, Christian Mayr 0001, Johannes Schemmel, Luis A. Plana, Steve Temple, David R. Lester, René Schüffny, Steve Furber |
IJCNN | 3 |
| 2013 | A model based comparison of BiFeO3 device applicability in neuromorphic hardwareabstractTwo terminal devices with switchable resistance have been of interest to electrical engineers for a long time, but only in the last few years has this attracted widespread attention. Recently a BiFeOe (BFO) capacitor-like metal-insulator-metal (MIM) structure was proposed as a synthetic synapse in neuromorphic systems, implementing voltage waveform driven spike timing dependent plasticity (STDP). Using a new device model that faithfully reproduces measurements of BFO-MIM structures we analyze how the switching characteristic affects the STDP learning window. Our simulations indicate that the gradual increase in the resistance change of BFO MIM structures result in a robust STDP with a biologically realistic learning window, whereas a distinct threshold followed by a steep hysteresis curve produce a narrow learning window and inflict strict operating conditions. Therefore we conclude that the steepness of the current voltage hysteresis curve is a fundamental characteristic to consider when designing synthetic synapses for neuromorphic hardware. Love Cederstroem, Paul Stärke, Christian Mayr 0001, Yao Shuai, Heidemarie Schmidt, René Schüffny |
ISCAS | 3 |
| 2013 | Live demonstration: Multiple-timescale plasticity in a neuromorphic systemabstractI. Demo Description Traditionally, neuromorphic ICs have integrated only reduced subsets of the rich repertoire of plasticity seen in biological preparations [1], [2]. The focus with respect to long term plasticity has been mostly on Spike-Time-Dependent Plasticity (STDP) [1]. Several ICs have also implemented forms of presynaptic short term dynamics, which filter synaptic pulse input, but have no influence on other timescales of plasticity. Here, we demonstrate an IC that implements short-term-, long-term-, and metaplasticity in an integrated way following [3], where these three different timescales interact to form the overall weight at the synapse. Fig. 1 shows an example presynaptic pattern with depression and the membrane trace as input for learning [3]. The resulting analog weight state shows the influence of presynaptic depression in the step increases, comparable to [1]. Also, different settings for the learning threshold exhibit a bias towards weight increase/decrease on a metaplastic (i.e. slow) timescale similar to [2]. The overall setup features several Maple-ICs of each 16 neurons and 512 of the above synapses, interlinked via FPGA-based pulse transmission. This allows network sizes of up to 200 neurons, sufficient to demonstrate the necessity for this type of learning for a range of computational neuroscience models. Christian Mayr 0001, Johannes Partzsch, Marko Noack, René Schüffny |
ISCAS | 1 |
| 2012 | Live demonstration: A scaled-down version of the BrainScaleS wafer-scale neuromorphic systemabstractThis demonstration is based on the wafer-scale neuromophic system presented in the previous papers by Schemmel et. al. (20120), Scholze et. al. (2011) and Millner et. al. (2010). The demonstration setup will allow the visitors to monitor and partially manipulate the neural events at every level. They will get an insight into the complex interplay between packet-based and realtime communication necessary to combine continuous-time mixed-signal neural networks with a packet-based transport network. Several network experiments implemented on the setup will be accessible for user interaction. Johannes Schemmel, Andreas Grübl, Stephan Hartmann 0002, Alexander Kononov, Christian Mayr 0001, Karlheinz Meier, Sebastian Millner, Johannes Partzsch, Stefan Schiefer, Stefan Scholze, René Schüffny, Marc-Olivier Schwartz |
ISCAS | 5 |
| 2012 | Waveform Driven Plasticity in BiFeO3 Memristive Devices: Model and ImplementationabstractMemristive devices have recently been proposed as efficient implementations of plastic synapses in neuromorphic systems. The plasticity in these memristive devices, i.e. their resistance change, is defined by the applied waveforms. This behavior resembles biological synapses, whose plasticity is also triggered by mechanisms that are determined by local waveforms. However, learning in memristive devices has so far been approached mostly on a pragmatic technological level. The focus seems to be on finding any waveform that achieves spike-timing-dependent plasticity (STDP), without regard to the biological veracity of said waveforms or to further important forms of plasticity. Bridging this gap, we make use of a plasticity model driven by neuron waveforms that explains a large number of experimental observations and adapt it to the characteristics of the recently introduced BiFeO$_3$ memristive material. Based on this approach, we show STDP for the first time for this material, with learning window replication superior to previous memristor-based STDP implementations. We also demonstrate in measurements that it is possible to overlay short and long term plasticity at a memristive device in the form of the well-known triplet plasticity. To the best of our knowledge, this is the first implementations of triplet plasticity on any physical memristive device. Christian Mayr 0001, Paul Stärke, Johannes Partzsch, René Schüffny, Love Cederstroem, Yao Shuai, Nan Du 0004, Heidemarie Schmidt |
NIPS | 1 |
| 2012 | A 32 GBit/s communication SoC for a waferscale neuromorphic system
Stefan Scholze, Holger Eisenreich, Sebastian Höppner, Georg Ellguth, Stephan Henker, Mario Ander, Stefan Hänzsche, Johannes Partzsch, Christian Mayr 0001, René Schüffny |
Integr. | 9 |
| 2011 | Live demonstration: Packet-based AER with 3Gevent/s cumulative throughputabstractTraditionally, the communication in neuromorphic VLSI systems has been done via parallel asynchronous transmission of Address-Event-Representations (AER) of neuron pulses. Recently, there has been a move towards greater event transmission speed via a serialization of the AER protocols. We give a live demonstration of a packet based synchronous serial AER infrastructure presented in a recent paper, which handles the complete off-wafer communication and configuration for a newly developed waferscale neuromorphic system, operating at a factor of 104faster than biological real-time. Pulse packets are routed from the host PC via Gbit Ethernet to an FPGA board, which forwards them to 4 purpose designed Digital Network ASICs (DNCs) on the same board. The DNCs buffer and sort the pulses, implementing 32 2GBit/s Low Voltage Differential Signaling (LVDS) interfaces to the neuromorphic circuits on the wafer. Pulse communication to other wafers is done via an FPGA- FPGA communication using 10 Gbit/s Aurora links. Stefan Schiefer, Stephan Hartmann 0002, Stefan Scholze, Johannes Partzsch, Christian Mayr 0001, Stephan Henker, René Schüffny |
ISCAS | 5 |
| 2010 | A critique of BCM behavior verification for STDP-type plasticity models
Christian Mayr 0001, Johannes Partzsch, René Schüffny |
ESANN | 1 |
| 2010 | Replicating experimental spike and rate based neural learning in CMOSabstractThe computational function of neural networks is thought to depend primarily on the learning/plasticity function carried out at the synapse. Neuromorphic circuit realizations have taken this into account by implementing a variety of synaptical processing functions, with most recent synapse circuits replicating some form of Spike Time Dependent Plasticity (STDP). However, STDP is being challenged by older rate-dependent learning rules as well as by biological experiments exhibiting more complex timing rules (e.g. spike triplets) as well as simultaneous rate- and timing dependent plasticity. In this paper, we present a circuit realization of a plasticity rule based on the postsynaptic neuron potential as well as the transmission profile of the presynaptic spike. To the best of our knowledge, this is the first circuit realization of synaptical behaviour which moves significantly beyond STDP, replicating the triplet experiments of Froemke and Dan, the combined timing and rate experiments of Sjoestroem et al., as well as conventional BCM behaviour. Christian Mayr 0001, Marko Noack, Johannes Partzsch, René Schüffny |
ISCAS | 1 |
| 2009 | Transient responses of activity-dependent synapses to modulated pulse trains
Christian Mayr 0001, Johannes Partzsch, René Schüffny |
Neurocomputing | 1 |
| 2008 | BCM and Membrane Potential: Alternative Ways to Timing Dependent Plasticity
Johannes Partzsch, Christian Mayr 0001, René Schüffny |
ICONIP (1) | 2 |
| 2007 | Neighborhood Rank Order Coding for Robust Texture Analysis and Feature ExtractionabstractResearch into the visual cortex and general neural information processing has led to various attempts to integrate pulse computation schemes in image analysis systems. Of interest is especially the robustness of representing an analogue signal in the phase or duration of a pulsed, quasi-digital signal, as well as the possibility of direct digital interaction, i.e. computation, among these signals. Such a computation can also achieve information compaction for subsequent processing stages. By using a pulse order encoding scheme motivated by dendritic pulse interaction, we will show that a powerful low-level feature and texture extraction operator, called pulsed local orientation coding (PLOC), can be implemented. Feature extraction results are being presented, and a possible VLSI implementation is detailed. Christian Mayr 0001, René Schüffny |
HIS | 1 |
| 2007 | Gabor-Like Image Filtering Using a Neural MicrocircuitabstractIn this letter, we present an implementation of a neural microcircuit for image processing employing Hebbian-adaptive learning. The neuronal circuit utilizes only excitatory synapses to correlate action potentials, extracting the uncorrelated ones, which contain significant image information. This circuit is capable of approximating Gabor-like image filtering and other image processing functions. Christian Mayr 0001, Arne Heittmann, René Schüffny |
IEEE Trans. Neural Networks | 1 |