EDBT 2026 Demo / reviewers in the wild / expert
Sebastian Höppner
dblp:19/9430
· DBLP profile ↗
22ranked-venue papers
7as first author
3since 2021 · last 2021
0000-0002-9938-2736ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Opportunities For A Hardware-Based OPC UA Server Implementation In Industry 4.0abstractWith the advent of the fourth industrial revolution i.e. Industry 4.0, plants and factories are becoming smarter and interconnected. The transitions demand vertical integration and seamless connectivity. For this purpose, there is a need for semantic communication between various devices including the heavily resource-constrained field devices. To address this, a real-time capable hardware-based implementation of a well-established semantic communication protocol, i.e. OPC Unified Architecture was designed and developed. This chip-based implementation is power-efficient and compact, making it suitable for the field level. The chip was analyzed and incorporated in a demonstrator as a proof of concept of its integration at field level in a plant module of the process industry. Various opportunities are also examined where the chip could be utilized to deliver benefits to existing and future technologies. Zohra Charania, Chris Paul Iatrou, Valentin Khaydarov, Richard Jacob, Robert Wittig, Heiner Bauer, Sebastian Höppner, René Bachmann, Philipp Bauer, Hendrik Deckert, Christian Mayr 0001, Gerhard P. Fettweis, Leon Urbas |
IECON | 7 |
| 2021 | Ultra-High Compression of Twiddle Factor ROMs in Multi-Core DSP for FMCW RadarsabstractThe increasing density of Multiple-Input Multiple-Output (MIMO) arrays in imaging radars for the automotive industry demands highly parallel systems with low-footprint accelerators, which would enable the concurrent processing of a high number of virtual channels with a low-latency, and without a high area overhead. In this paper, we design, implement, and test multiple handcrafted compression schemes for Twiddle Factor (TF) Read-Only Memories (ROM), aiming to reduce the footprint of a variable-length and dual-radix Fast Fourier Transform (FFT) accelerator in a Multi-core Digital Signal Processor (DSP) for Frequency Modulated Continuous Wave (FMCW) radars. The compression schemes proposed in this paper involve double delta encoding, Radix-specific address optimizations per port, symmetry inclusion, and exploitation of the bit resolution changes within the radar processing chain. All schemes are verified in an FPGA in terms of logic utilization and quantization using a 77-GHz radar, and implemented in a RISCV-based Processing Element (PE) of a Multi-core DSP with an Adaptive Body Bias (ABB) approach in 22FDX technology for assessing area, leakage, and relative latency savings when compared with a dual-ROM equivalent in the state-of-the-art. Hector A. Gonzalez, Florian Kelber, Marco Stolba, Chen Liu 0031, Bernhard Vogginger, Stefan Hänzsche, Stefan Scholze, Sebastian Höppner, Christian Mayr 0001 |
ISCAS | 8 |
| 2021 | Hardware Implementation of an OPC UA Server for Industrial Field DevicesabstractIndustrial plants suffer from a high degree of complexity and incompatibility in their communication infrastructure, caused by a wild mix of proprietary technologies. This prevents transformation toward Industry 4.0 and the Industrial Internet of Things. Open platform communications unified architecture (OPC UA) is a standardized protocol that addresses these problems with uniform and semantic communication across all levels of the hierarchy. However, its adoption in embedded field devices, such as sensors and actuators, is still lacking due to prohibitive memory and power requirements of software implementations. We have developed a dedicated hardware engine that offloads processing of the OPC UA protocol and enables the realization of compact and low-power field devices with OPC UA support. As part of a proof-of-concept embedded system, we have implemented this engine in a 22-nm FDSOI technology, representing the first ASIC implementation of an OPC UA server. We measured performance, power consumption, and memory footprint of our test chip and compared it with a software implementation based on open62541 and a Raspberry Pi 2B. Our OPC UA hardware engine is 50 times more energy efficient and only requires 36 KiB of memory. The complete system consumes only 24 mW under full load, making it suitable for low-power embedded applications. Heiner Bauer, Sebastian Höppner, Chris Paul Iatrou, Zohra Charania, Stephan Hartmann 0002, Saif-Ur Rehman, Andreas Dixius, Georg Ellguth, Dennis Walter, Johannes Uhlig, Felix Neumärker, Marc Berthel, Marco Stolba, Florian Kelber, Leon Urbas, Christian Mayr 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Hard Real-Time Capable OPC UA Server as Hardware Peripheral for Single Chip IoT SystemsabstractThe fast semantics project examines the use of the OPC Unified Architecture (OPC UA) in embedded industrial systems and proposes the design of a customizable, hard real time capable OPC UA Intellectual Property Core (IP Core) for single chip computing plattforms. This allows using OPC UA in both novel energy efficient sensor applications and in state of the art field devices. These single chip OPC UA servers form the semantic data sources for future applications such as cloud based added value services or machine learning applications. This article presents the design alternatives and first synthesis results for the implementation of OPC UA servers in embedded systems. Chris Paul Iatrou, Heiner Bauer, Markus Graube, Sebastian Höppner, Julian Rahm, Leon Urbas |
ETFA | 4 |
| 2019 | A Fast Lock-In Ultra Low-Voltage ADPLL Clock Generator with Adaptive Body Biasing in 22nm FDSOI TechnologyabstractSystems on Chip for the Internet of Things require fast-locking and robust clock generators to maximize the effectiveness of power management techniques such as Dynamic Voltage and Frequency Scaling and duty-cycling. We present an ADPLL clock generator based on a fully digital DCO architecture with an inherently linear and offset-free tuning characteristic that allows fast lock-in within three and frequency changes during operation within two reference cycles. Measurements from a testchip in 22nm FDSOI CMOS technology show operation from 0.4 to 0.8 V and 20 to 790 MHz. At 0.5 V, only 107 μW are consumed to generate a 100 MHz clock with 59 ps RMS period jitter. Adaptive Body Biasing improves the jitter performance by up to 40% through compensation of PVT variation. Florian Schraut, Holger Eisenreich, Sebastian Höppner, Christian Mayr 0001 |
ISCAS | 3 |
| 2019 | A Multi-Bit PFD Architecture for ADPLLs with Built-In Jitter Self-CalibrationabstractThis paper presents a multi-bit phase-frequency detector (PFD) architecture with self-calibration scheme to reduce the jitter in all-digital phase-locked loops. A standard bang-bang PFD is extended by two additional PFDs which allow the measurement of the ADPLL jitter distribution width. A built-in self-calibration algorithm can utilise this feature to optimise the loop filter gain for minimised overall jitter. The proposed technique is demonstrated in hardware within an LC-ADPLL in 28 nm SLP CMOS technology with 7.5 GHz output from a reference frequency of 100 MHz. By using the proposed PFD technique the total accumulated jitter is reduced by 20%. Franz Marcus Schüffny, Sebastian Höppner, Alexander Oefelein, Christian Mayr 0001 |
ISCAS | 2 |
| 2019 | Performance Analysis of a Comparator Based Mixed-Signal Control Loop in 28 nm CMOSabstractIn differential signaling systems using copper wires common mode signals are the cause of emission of electromagnetic energy. Especially in Automotive Ethernet systems this is a challenging problem. Beside classical passive components like common mode chokes active circuits can help to reduce the emission. This allows inexpensive and resource-conserving unshielded twisted pair cables to be used. This paper shows the approach of using a mixed-signal control loop based on a comparator and a 8 bit DAC for regulating the common mode voltage of an Automotive Ethernet DAC in 28 nm CMOS. An attenuation for interferers with frequencies up to 500 kHz is achieved and reaches up to 15 dB at maximum. The control loop utilizes the successive approximation algorithm commonly used for delay locked loops and DC trimming in mixed-signal circuits. In contrast to known applications the performance and usability at higher frequencies is considered in this paper. Being a nonlinear, time-variant system an analytical design of the control loop is very difficult. Therefore parametrical measurements show the dependency of frequency, amplitude and signal form of an applied common mode interferer source. Florian Protze, Martin Kreißig, Frank Ellinger, Sebastian Höppner, Stephan Hartmann 0002, Stefan Hänzsche, Stefan Scholze, Georg Ellguth, Christian Mayr 0001 |
VLSI-SoC | 4 |
| 2018 | Approximate Fixed-Point Elementary Function Accelerator for the SpiNNaker-2 Neuromorphic ChipabstractNeuromorphic chips are used to model biologically inspired Spiking-Neural-Networks (SNNs) where most models are based on differential equations. Equations for most SNN algorithms usually contain variables with one or more excomponents. SpiNNaker is a digital neuromorphic chip that has so far been using pre-calculated look-up tables for exponential function. However this approach is limited because the memory requirements grow as more complex neural models are developed. To save already limited memory resources in the next generation SpiNNaker chip, we are including a fast exponential function in the silicon. In this paper we analyse iterative algorithms for elementary functions and show how to build a single hardware accelerator for exp and natural log, for a neuromorphic chip prototype, to be manufactured in a 22 nm FDSOI process. We present the accelerator that has algorithmic level approximation control, allowing it to trade precision for latency and energy efficiency. As an addition to neuromorphic chip application, we provide analysis of a parameterized elementary function unit that can be tailored for other systems with different power, area, accuracy and latency constraints. Mantas Mikaitis, David R. Lester, Delong Shang, Steve Furber, Gengting Liu, Jim D. Garside, Stefan Scholze, Sebastian Höppner, Andreas Dixius |
ARITH | 8 |
| 2017 | A Heterogeneous SDR MPSoC in 28 nm CMOS for Low-Latency Wireless ApplicationsabstractCurrent and future applications impose high demands on software-defined radio (SDR) platforms in terms of latency, reliability, and flexibility. This paper presents a heterogeneous SDR MPSoC with a hexagonal network-on-chip to address these issues. It features four data processing modules and a baseband processing engine for iterative multiple-input multiple-output (MIMO) receiving. Integrated memory controllers enable dynamic data flow mapping and application isolation. In a 4 x 4 MIMO application scenario, the MPSoC achieves a throughput of 232 Mbit/s with a latency of 20 μs while consuming 414 mW. It outperforms state-of-the-art platforms in terms of throughput by a factor of 4. Sebastian Haas, Tobias Seifert, Benedikt Noethen, Stefan Scholze, Sebastian Höppner, Andreas Dixius, Esther P. Adeva, Thomas R. Augustin, Friedrich Pauls, Sadia Moriam, Mattis Hasler, Erik Fischer, Yong Chen 0014, Emil Matús, Georg Ellguth, Stephan Hartmann 0002, Stefan Schiefer, Love Cederstroem, Dennis Walter, Stephan Henker, Stefan Hänzsche, Johannes Uhlig, Holger Eisenreich, Stefan Weithoffer, Norbert Wehn, René Schüffny, Christian Mayr 0001, Gerhard P. Fettweis |
DAC | 5 |
| 2017 | Live demonstration: Dynamic voltage and frequency scaling for neuromorphic many-core systemsabstractWe present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation. Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Prateek Joshi, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber |
ISCAS | 1 |
| 2017 | Dynamic voltage and frequency scaling for neuromorphic many-core systemsabstractWe present a dynamic voltage and frequency scaling technique within SoCs for per-core power management: the architecture allows for individual, self triggered performance-level scaling of the processing elements (PEs) within less than 100ns. This technique enables each core to adjust its local supply voltage and frequency depending on its current computational load. A test chip has been implemented in 28nm CMOS technology, as prototype of the SpiNNaker2 neuromorphic many core system, containing 4 PEs which are operational within the range of 1.1V down to 0.7V at frequencies from 666MHz down to 100MHz; the effectiveness of the power management technique is demonstrated using a standard benchmark from the application domain. The particular domain area of this application specific processor is real-time neuromorphics. Using a standard benchmark - the synfire chain - we show that the total power consumption can be reduced by 45%, with 85% baseline power reduction and a 30% reduction of energy per neuron and synapse computation, all while maintaining biological real-time operation. Sebastian Höppner, Yexin Yan, Bernhard Vogginger, Andreas Dixius, Johannes Partzsch, Felix Neumärker, Stephan Hartmann 0002, Stefan Schiefer, Stefan Scholze, Georg Ellguth, Love Cederstroem, Matthias Eberlein, Christian Mayr 0001, Steve Temple, Luis A. Plana, Jim D. Garside, Simon Davidson, David R. Lester, Steve Furber |
ISCAS | 1 |
| 2017 | A fixed point exponential function accelerator for a neuromorphic many-core systemabstractMany models of spiking neural networks heavily rely on exponential waveforms. On neuromorphic multiprocessor systems like SpiNNaker, they have to be approximated by dedicated algorithms, often dominating the processing load. Here we present a processor extension for fast calculation of exponentials, aimed at integration in the next-generation SpiNNaker system. Our implementation achieves single-LSB precision in a 32bit fixed-point format and 250Mexp/s throughput at 0.44nJ/exp for nominal supply (1.0V), or 0.21nJ/exp at 0.7V supply and 77Mexp/s, demonstrating a throughput multiplication of almost 50 and 98% energy reduction at 2% area overhead per processor on a 28nm CMOS chip. Johannes Partzsch, Sebastian Höppner, Matthias Eberlein, René Schüffny, Christian Mayr 0001, David R. Lester, Steve Furber |
ISCAS | 2 |
| 2016 | An MPSoC for energy-efficient database query processingabstractThis paper presents a heterogeneous database hardware accelerator MPSoC manufactured in 28 nm SLP CMOS. The 18 mm2 chip integrates a runtime task scheduling unit for energy-efficient query processing and hierarchical power management supported by an ultra-fast dynamic voltage and frequency scaling. Four processing elements, connected by a star-mesh network-on-chip, are accelerated by an instruction set extension tailored to fundamental data-intensive applications. We evaluate the MPSoC with typical database benchmarks focusing on scans and bitmap operations. When the processing elements operate on data stored in local memories, the chip consumes 250 mW and shows a 96x energy efficiency improvement compared to state-of-the-art platforms. Sebastian Haas, Oliver Arnold, Benedikt Noethen, Stefan Scholze, Georg Ellguth, Andreas Dixius, Sebastian Höppner, Stefan Schiefer, Stephan Hartmann 0002, Stephan Henker, Thomas Hocker, Jörg Schreiter, Holger Eisenreich, Jens-Uwe Schluessler, Dennis Walter, Tobias Seifert, Friedrich Pauls, Mattis Hasler, Yong Chen 0014, Hermann Hensel, Sadia Moriam, Emil Matús, Christian Mayr 0001, René Schüffny, Gerhard P. Fettweis |
DAC | 7 |
| 2016 | A Calibration Technique for Bang-Bang ADPLLs Using Jitter Distribution MonitoringabstractThis brief presents a built-in self-calibration (BISC) technique for minimization of the total jitter in bang-bang all-digital phase-locked loops (ADPLLs). It is based on the addition of a monitoring phase-frequency detector (PFD) with tunable delay cells for the reference clock and the divider clock and a counter for this PFD output signal. This allows for on-chip binary comparison of the jitter distribution widths at the ADPLL PFD input, when ADPLL filter parameters are altered. Since only a relative comparison is performed, no accurate delay calibration is required. The statistical properties of this comparison of two random distributions are analyzed theoretically, and guidelines for circuit dimensioning are derived. The proposed method is used for BISC by adaption of the ADPLL filter coefficients. This allows for jitter minimization under process, voltage and temperature variations as well as gain and period jitter of the digitally controlled oscillator. The proposed calibration technique is verified by system simulations and measurements of a silicon prototype implementation in 28-nm CMOS technology. Sebastian Höppner, Johannes Partzsch, Johannes Neumann, René Schüffny, Christian Mayr 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | An all-digital PWM generator with 62.5ps resolution in 28nm CMOS technologyabstractThis paper presents an all-digital pulse width modulator (PWM) for application in integrated DC-DC converters. Based on a multi-phase clock signal a PWM resolution of 62.5ps is achieved, resulting in up to 16-Bit PWM resolution at 4096ns period. The PWM signal duty cycle and period can be arbitrarily changed within a single cycle which allow efficient alldigital implementation of spread spectrum clocking schemes. The circuit has been implemented in 28nm SLP CMOS technology. It consumes 0.3mW when operating from a 1.0V supply. Sebastian Höppner, Stefan Hänzsche, Stefan Scholze, René Schüffny |
ISCAS | 1 |
| 2014 | A compact on-chip IR-drop measurement system in 28 nm CMOS technologyabstractA sensor system for measuring the power-ground (PG) noise in very large scale integrated circuits is presented. The proposed system utilizes sensor elements with standard cell dimensions enabling high spatial resolution voltage measurements of power and ground rails. Asynchronous sub-sampling is used to directly convert the analog signals into the digital domain inside the sensors to ensure precise waveform acquisition. Timing signals are derived from a all-digital phase-locked-loop (ADPLL) which guarantees accurate low-noise sampling of the supply waveforms. The sensor system has been implemented in a 28nm CMOS test chip. Simultaneous acquisition of voltage drop and ground bounce at 300 probe points within a 120μm × 120μm macro at 62.5 ps time and up to 250μV voltage resolution shows the capabilities of both, high spatial and high temporal resolution measurement of PG noise. Sebastian Dietel, Sebastian Höppner, Holger Eisenreich, Georg Ellguth, Stefan Hänzsche, Stephan Henker, René Schüffny, Tim Brauninger, Ulrich Fiedler |
ISCAS | 2 |
| 2013 | A Compact Clock Generator for Heterogeneous GALS MPSoCs in 65-nm CMOS TechnologyabstractThis paper presents an all-digital phase-locked loop (ADPLL) clock generator for globally asynchronous locally synchronous (GALS) multiprocessor systems-on-chip (MPSoCs). With its low power consumption of 2.7 mW and ultra small chip area of 0.0078 mm2it can be instantiated per core for fine-grained power management like DVFS. It is based on an ADPLL providing a multiphase clock signal from which core frequencies from 83 to 666 MHz with 50% duty cycle are generated by phase rotation and frequency division. The clock meets the specification for DDR2/DDR3 memory interfaces. Additionally, it provides a dedicated high-speed clock up to 4 GHz for serial network-on-chip data links. Core frequencies can be changed arbitrarily within one clock cycle for fast dynamic frequency scaling applications. The performance including statistical analysis of mismatch has been verified by a prototype in 65-nm CMOS technology. Sebastian Höppner, Holger Eisenreich, Stephan Henker, Dennis Walter, Georg Ellguth, René Schüffny |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | A power management architecture for fast per-core DVFS in heterogeneous MPSoCsabstractThis paper presents a power management architecture for MPSoCs that allows fast switching between multiple onchip supply voltage levels per core. Operation is based on distinct scenarios for power-up and supply level change, with individual numbers of pre-charge switches for supply noise reduction. The power management controller is highly configurable for adaption to a wide range of supply network parasitics in heterogeneous MPSoCs. This architecture has been validated by measurements in 65nm CMOS technology. Power-up and DVFS level changes can be performed in less than 20ns with reduced parasitic voltage drop of active cores. Sebastian Höppner, Chenming Shao, Holger Eisenreich, Georg Ellguth, Mario Ander, René Schüffny |
ISCAS | 1 |
| 2012 | A 32 GBit/s communication SoC for a waferscale neuromorphic system
Stefan Scholze, Holger Eisenreich, Sebastian Höppner, Georg Ellguth, Stephan Henker, Mario Ander, Stefan Hänzsche, Johannes Partzsch, Christian Mayr 0001, René Schüffny |
Integr. | 3 |
| 2011 | Strategies for initial sizing and operating point analysis of analog circuitsabstractThis work presents novel analog sizing flows based on analytical techniques. A graph-based operating point driven sizing approach provides operating point voltages and a rough sizing with respect to constraints. A voltage-range analysis method using linearized operating-point models obtains information about feasible voltage ranges. A direct-sizing method solves nonlinear algebraic circuit equations directly to obtain design parameters from specifications. All three methods require no or only minimum simulation effort and can provide quick insight into circuit design space and constraints in an early design stage. They allow flexible inclusion into state-of-the-art simulation-based optimization flows, where they lead to improved results with less optimization effort and prevent unnecessary simulation effort on unfeasible circuit topologies. The sizing flows are enhanced by a commercial optimization tool in order to obtain reliable circuits. Volker Boos, Jacek Nowak, Matthias Sylvester, Stephan Henker, Sebastian Höppner, Heiko Grimm, Dominik Krausse, Ralf Sommer |
DATE | 5 |
| 2010 | Wide swing signal amplification by SC voltage doublingabstractThis paper presents a switched capacitor voltage doubler as analog signal amplifier to increase the swing in low voltage electronics like PLLs. A signal is doubled with respect to a reference voltage which results in a wide signal swing exceeding the supply rails. An analytical expression of the transfer function, from which design equations for circuit parameters can be obtained, is derived including the effect of incomplete voltage settling due to finite switch resistances. The theory is validated by measurements of a voltage doubler as part of a PLL charge pump in 90nm CMOS technology. Sebastian Höppner, René Schüffny, Zuo-Min Tsai, Huei Wang |
ISCAS | 1 |
| 2010 | A low-power cell-based-design multi-port register file in 65nm CMOS technologyabstractThis paper presents the design of a register file with 4 write and 6 read ports for an SDR multiprocessor in 65nm CMOS technology. A cell-based design (CBD) methodology is employed in which the circuit is partitioned into complex sub-cells, optimized on transistor level and layout. Each cell is completely characterized concerning timing and power for seamless integration into a semi-custom design flow. The CBD implementation shows 30% savings of power and 40% of area compared to a conventional semi-custom solution. The average power is 2.7mW from 1.0V supply and 300MHz operating frequency which is superior to previously published designs. Johannes Uhlig, Sebastian Höppner, Georg Ellguth, René Schüffny |
ISCAS | 2 |