EDBT 2026 Demo / reviewers in the wild / expert
Norbert Wehn
dblp:48/6980
· DBLP profile ↗
155ranked-venue papers
4as first author
43since 2021 · last 2026
0000-0002-9010-086XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 109 · 4 first-author · 33 since 2021Software engineering, systems software and programming languages · 52 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021Computer networks · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microscaling-Stochastic Computing Based Systolic Arrays for Energy-Efficient Deep Neural Network InferenceabstractDeep neural networks (DNNs) require increasingly high compute and memory resources. Microscaling (MX) data formats improve energy efficiency and preserve accuracy under aggressive bit-width reduction, but further gains from continued bit-width reduction remain challenging. This work proposes a hybrid computation scheme that integrates MX data formats with stochastic computing (SC) to improve the energy efficiency of DNN inference under constrained bit widths. Model parameters are stored in MX format, while multiplications and accumulations are performed in the SC domain. MX reduces memory footprint, while SC improves compute energy efficiency. To address the latency and accuracy challenges of SC, we employ a parallel bitstream-generation scheme and an encoding strategy that reduces random fluctuation error. Experimental results demonstrate up to a 2× improvement in energy efficiency while maintaining inference accuracy within 1–2% of an FP32 baseline. Mohammad Hassani Sadi, Bilal Hammoud, Norbert Wehn |
DATE | 3 |
| 2026 | Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAIabstractArtificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps. Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl |
DATE | 24 |
| 2026 | From RTL to Prompt Coding: Empowering the Next Generation of Chip Designers through LLMs
Lukas Krupp, Matthew Venn, Norbert Wehn |
ISCAS | 3 |
| 2026 | AnaCraft: Duel-Play Probabilistic-Model-Based Reinforcement Learning for Sample-Efficient PVT-Robust Analog Circuit Sizing OptimizationabstractRecent advancements in machine learning offer the potential for finding faster and robust optimization approaches for analog circuit design automation. However, fully automated yet fast and PVT-robust sizing algorithms are still lacking as even the most recent methods continue to require extensive simulations or domain-specific circuit expertise. In this paper, we present a PVT-robust analog circuit sizing method, called AnaCraft, that is the first to introduce an adversarial training scheme of multi-agent reinforcement learning (RL) for robust circuit design automation. We adopt the soft actor-critic (SAC) agent for circuit sizing, which outperforms other actor-critic agents in stability and robustness. Then, we introduce a duel-play scheme to address PVT-robustness, where sizing agents cooperate to find optimal circuit parameters while competing with an adversarial PVT agent. We combine this approach with the model-based policy optimization method: an ensemble of probabilistic models is trained and used to extract many short rollouts of generated data for updating the sizing agents. We test our algorithm on the sizing of operational amplifiers in a 45nm CMOS technology, as well as on a complex data receiver circuit in a predictive 7nm FinFET technology. This demonstrates our approach’s ability to find PVT-robust power-area-optimal sizes for advanced technologies and circuits. Our proposed method achieves a higher figure of merit with up to 3x fewer circuit simulations and 2x less runtime compared to existing state-of-the-art methods. Mohsen Ahmadzadeh, Jan Lappas, Norbert Wehn, Georges Gielen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-SourceabstractChip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators. Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali |
CODES+ISSS | 24 |
| 2025 | Multi-Partner Project: Sustainable Textile Electronics (STELEC)abstractE-textiles are rapidly emerging as an important area of electronic circuit applications. It also facilitates many socially important applications such as personalized health, elderly care, and smart agriculture. However, the environmental impact and sustainability of e-textiles remain very problematic. STELEC, short for Sustainable Textile ELECtronics, is an interdisciplinary research project funded by the European Innovation Council (EIC) under the Pathfinder programme on the responsible elec-tronics topic seeking cutting-edge innovation. STELEC started in September 2024 and is in its initial stage. The project is a multinational collaboration of research institutes, universities and companies across Europe. It aims at developing next-generation textile-based electronics in applications from sensing, processing to AI, with a commitment to full lifecycle sustainability. Bo Zhou 0005, Mengxi Liu 0004, Sizhen Bian, Daniel Geißler, Paul Lukowicz, José Miranda 0001, Jonathan Dan, David Atienza 0001, Mohamed Amine Riahi, Norbert Wehn, Russel N. Torah, Sheng Yong, Stephen P. Beeby, Magdalena Kohler, Berit Greinke, Junchun Yu, Vincent Nierstrasz, Leila Sheldrick, Rebecca Stewart, Tommaso Nieri, Matteo Maccanti, Daniele S. Spinelli |
DATE | 10 |
| 2025 | Improving Chip Design Enablement for Universities in Europe - A Position PaperabstractThe semiconductor industry is pivotal to Europe's economy, especially within the industrial and automotive sectors. However, Europe faces a significant shortfall in chip design capabilities, marked by a severe skilled labor shortage and lagging contributions in the design value chain segment. This paper explores the role of European universities and academic initiatives in enhancing chip design education and research to address these deficits. We provide a comprehensive overview of current European chip design initiatives, analyze major challenges in recruitment, productivity, technology access, and design enablement, and identify strategic opportunities to strengthen chip design capabilities within academic institutions. Our analysis leads to a series of recommendations that highlight the need for coordinated efforts and strategic investments to overcome these challenges. Lukas Krupp, Ian O'Connor, Luca Benini, Christoph Studer, Joachim Neves Rodrigues, Norbert Wehn |
DATE | 6 |
| 2025 | Multi-Partner Project: Open-Source Design Tools for Co-Development of AI Algorithms and AI Chips: (Initial Stage)abstractChip technologies are crucial for the digital transformation of industry and society. Artificial Intelligence (AI) is playing an increasingly important role in both our daily lives and in industry. The development of advanced AI chip designs, essential for the successful deployment of AI, is of critical importance for innovation and competitiveness. However, challenges arise from the complexity of hardware development, expensive access to state-of-the-art design tools, and a global shortage of hardware experts. In addition to cost optimization, computational power, and energy consumption, security and trustworthiness are becoming increasingly important. This project aims to address these challenges in AI chip design by enabling efficient hardware development. We are developing a seamless transition between software-based AI model development and optimization, and efficient hardware implementation, while considering security, trustworthiness, and energy efficiency. An open-source approach plays a key role, facilitating access for small and medium-sized enterprises (SMEs) and expanding the community involved in AI chip design to help mitigate the shortage of skilled professionals. Mehdi Baradaran Tahoori, Jürgen Becker 0001, Jörg Henkel, Wolfgang Kunz, Ulf Schlichtmann, Georg Sigl, Jürgen Teich, Norbert Wehn |
DATE | 8 |
| 2025 | Design of a Low-Power 4.3 Gb/s Transceiver Using Pre-computed Lookup TablesabstractHigh-speed memory interfaces require the design of optimised low-power and robust analog circuits. This poses significant challenges in sizing due to the inability of using complex modern transistor models, such as the Berkeley Shortchannel IGFET Model (BSIM), to calculate the sizing of the circuits based on hand-analysis. This forces designers to perform iterative simulations, which is time-intensive and error-prone. To address these issues, this work presents an automated sizing approach using pre-computed lookup tables (LUTs) for a 4.3 Gb/s LPDDR4X transceiver in 12nm FinFET technology. The receiver is designed based on gm/IDsizing methodology where look-up tables are used to compute a set of matrices representing the possible design points based on the circuit topology. The design space is constrained by the biasing level, gain, and bandwidth to find an optimised design point in terms of operation region, speed and power. The driver circuit design is automated based on a new algorithm which computes the estimated ON-resistance from the look up table and finds an optimum sizing which fulfills the required impedance range. The calculated sizing of the proposed design approach was used as an input for spectre simulator to compare the simulation results with the input specifications, showing an error margin of less than 1%. Furthermore, the receiver power consumption was evaluated to be 70% less than the work in literature. The proposed driver topology, which uses low voltage swing terminated logic (LVSTL) and near-ground signaling (NGS), provides a38−120Ωimpedance range at all process, voltage and temperature variations (PVT) for the postlayout results and consumes relatively low power when compared to literature. Hussien Abdo, Jan Lappas, Mohammadreza Esmaeilpour, Christian Weis, Norbert Wehn |
DDECS | 5 |
| 2025 | Analysis and Mitigation of Radiation Effects in SRAM-based Register Files
Surendra Hemaram, Mahta Mayahinia, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif, Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian |
ETS | 5 |
| 2025 | Machine Learning for Improving Timing AccuracyabstractTiming analysis belongs to the cornerstones of digital integrated circuit design. But with increasing technology complexity and lower power supply voltages, vigilance is needed to ensure that the timing paradigm continues to be sufficient for its purpose. State-Of-the-art timing models utilize current or voltage waveforms that are generated from circuit simulation, but they require large amounts of storage -especially when power supply, temperature, and process variations are included. Also, Miller coupling between the input and output of a gate, which can cause an up to a 70% error in propagation delay, is poorly modeled.In this paper, we propose a novel machine learning-based approach to waveform modeling in digital circuits. By leveraging machine learning techniques, we are able to improve the accuracy and predictive capabilities of timing models while simultaneously reducing model size by more than 5×. Mohamed Amine Riahi, Sani R. Nassif, Norbert Wehn |
ISCAS | 3 |
| 2025 | eFAirWrite: Bringing energy efficient text entry to next generation smart devicesabstractText entry tasks for emerging wireless Augmented Reality (AR) and Virtual Reality (VR) devices can be realized in many ways, one of the most promising methods is based on an Inertial Measurement Unit (IMU) sensor, which resembles a human writing style and is called air-writing. An existing air-writing Deep Neural Network (DNN) based algorithm called FAirWrite achieves state-of-the-art accuracy. However, this algorithm is optimized only for accuracy without considering the implementation constraints. State-of-the-art implementation executes the algorithm in a cloud, which is associated with large communication latency; privacy issues with respect to data transfer; and a need for reliable internet connection. A solution that tackles all three challenges is executing the algorithm locally at the edge in the closest proximity to the sensor. However, inference at the edge is challenging due to limited memory and computing resources, which can impact accuracy and increase latency. Additionally, battery-powered edge devices must adhere to strict power and energy consumption limits. All these constraints collectively restrict the model size and computational complexity that can be deployed near the sensor. In this work, we explore various optimizations required to enable state-of-the-art FAirWrite algorithm for a real-world deployment scenario, i.e. executing on an edge device. We perform a multi-layer design-space exploration considering multiple levels of design hierarchy spanning from optimizations applied on algorithmic level down to hardware level, considering various deployment run-times and edge platforms including embedded micro-controller, embedded Central Processing Unit (CPU), embedded Graphics Processing Unit (GPU), Neural Processing Unit (NPU), and Field-Programmable Gate Array (FPGA). The complexity reduction optimizations result in a smaller eFAirWrite model without any accuracy degradation. We propose and implement a custom hardware architecture of the algorithm on an FPGA by utilizing custom data-types and memory hierarchy. We demonstrate that the FPGA implementation achieves , , , and higher energy efficiency as compared to NPU, embedded GPU, embedded CPU, and micro-controller, respectively, which makes it a suitable edge platform for portable and real-time air-writing text entry. Muhammad Mohsin Ghaffar, Ahmad Abdullah, Junaid Younas, Vladimir Rybalkin, Jonas Ney, Paul Lukowicz, Norbert Wehn |
Expert Syst. Appl. | 7 |
| 2025 | A Lightweight PUF-Based Weights Obfuscation Technique for Secure In-Memory AI InferenceabstractIn-Memory Computing (IMC) has introduced a novel computational approach that substantially improves emerging embedded AI accelerators’ latency and power consumption efficiency. Despite the numerous advantages, IMC architectures also introduce new security vulnerabilities that may compromise the confidentiality of the deployed Neural Network (NN) algorithms. In this work, following an analysis of the potential threats, we present a novel lightweight security countermeasure for IMC accelerators. This methodology can be employed to de-obfuscate the pre-trained weights of NN architectures whose bits’ significance has been reordered prior to the deployment phase onto the IMC crossbar. The proposed solution is based on the coordinated action of a Ferroelectric Field-Effect Transistor (FeFET) based Physical Unclonable Function (PUF) design and shifting registers. These components perform custom arithmetic shift operations on the values calculated by the IMC device at runtime to obtain a coherent inference computation. Furthermore, a design-space exploration method is proposed to investigate the trade-off between area overhead and the level of security provided by the implementation. The results show that with less than 3% of area overhead our design is robust against all the tested attack strategies. Luca Parrini, Anirban Kar, Benjamin Hettwer, Taha Soliman, Yogesh Singh Chauhan, Hussam Amrouch, Norbert Wehn |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Row-Merged Polar Codes: Analysis, Design, and Decoder ImplementationabstractRow-merged polar codes are a family of pre-transformed polar codes (PTPCs) with little precoding overhead. Providing an improved distance spectrum over plain polar codes, they are capable to perform close to the finite-length capacity bounds. However, there is still a lack of efficient design procedures for row-merged polar codes. Using novel weight enumeration algorithms with low computational complexity, we propose a design methodology for row-merged polar codes that directly considers their minimum distance properties. The codes significantly outperform state-of-the-art cyclic redundancy check (CRC)-aided polar codes under successive cancellation list (SCL) decoding in error-correction performance. Furthermore, we present fast simplified successive cancellation list (Fast-SSCL) decoding of PTPCs, based on which we derive a high-throughput, unrolled architecture template for fully pipelined decoders. Implementation results of SCL decoders for row-merged polar codes in a 12nm technology additionally demonstrate the superiority of these codes with respect to the implementation costs, compared to state-of-the-art reference decoder implementations. Andreas Zunker, Marvin Geiselhart, Lucas Johannsen, Claus Kestel, Stephan ten Brink, Timo Vogt, Norbert Wehn |
IEEE Trans. Commun. | 7 |
| 2024 | Timing Analysis beyond Complementary CMOS Logic StylesabstractWith scaling unabated, device density continues to increase, but power and thermal budgets prevent the full use of all available devices. This leads to the exploration of alternative circuit styles beyond traditional CMOS, especially dynamic data-dependent styles, but the excessive pessimism inherent in conventional static timing analysis tools presents a barrier to adoption. One such circuit family is Pass-Transistor Logic (PTL), which holds significant promise but behaves differently from CMOS in that traditional CMOS-oriented EDA tools cannot produce sufficiently accurate performance estimates. In this work, we revisit timing analysis and its premises and show a significantly improved methodology of a more generalized dynamic timing engine that accurately predicts timing performance for traditional CMOS as well as PTL with an accuracy of 4.0% compared to SPICE and with a run-time comparable to traditional gate-level simulation. The run-time improvement compared with SPICE is four orders of magnitude. Jan Lappas, Mohamed Amine Riahi, Christian Weis, Norbert Wehn, Sani R. Nassif |
ASPDAC | 4 |
| 2024 | A Mapping of Triangular Block Interleavers to DRAM for Optical Satellite CommunicationabstractCommunication in optical downlinks of low earth orbit (LEO) satellites requires interleaving to enable reliable data transmission. These interleavers are orders of magnitude larger than conventional interleavers utilized for example in wireless communication. Hence, the capacity of on-chip memories (SRAMs) is insufficient to store all symbols and external memories (DRAMs) must be used. Due to the overall requirement for very high data rates beyond 100 Gbit/s, DRAM bandwidth then quickly becomes a critical bottleneck of the communication system. In this paper, we investigate triangular block interleavers for the aforementioned application and show that the standard mapping of symbols used for SRAMs results in low bandwidth utilization for DRAMs, in some cases below 50 %. As a solution, we present a novel mapping approach that combines different optimizations and achieves over 90 % bandwidth utilization in all tested configurations. Further, the mapping can be applied to any JEDEC-compliant DRAM device. Lukas Steiner, Timo Lehnigk-Emden, Markus Fehrenz, Norbert Wehn |
DATE | 4 |
| 2024 | Achieving High Throughput with a Trainable Neural-Network-Based Equalizer for Communications on FPGAabstractThe ever-increasing data rates of modern communication systems lead to severe distortions of the communication signal, imposing great challenges to state-of-the-art signal processing algorithms. In this context, neural network (NN)-based equalizers are a promising concept since they can compensate for impairments introduced by the channel. However, due to the large computational complexity, efficient hardware implementation of NNs is challenging. Especially the backpropagation algorithm, required to adapt the NN's parameters to varying channel conditions, is highly complex, limiting the throughput on resource-constrained devices like field programmable gate arrays (FPGAs). In this work, we present an FPGA architecture of an NN-based equalizer that exploits batch-level parallelism of the convolutional layer to enable a custom mapping scheme of two multiplication to a single digital signal processor (DSP). Our implementation achieves a throughput of up to 20 GBd, which enables the equalization of high-data-rate nonlinear optical fiber channels while providing adaptation capabilities by retraining the NN using backpropagation. As a result, our FPGA implementation outperforms an embedded graphics processing unit (GPU) in terms of throughput by two orders of magnitude. Further, we achieve a higher energy efficiency and throughput as state-of-the-art NN training FPGA implementations. Thus, this work fills the gap of high-throughput NN-based equalization while enabling adaptability by NN training on the edge FPGA. Jonas Ney, Norbert Wehn |
DSD | 2 |
| 2024 | Error Detection and Correction Codes for Safe In-Memory ComputationsabstractIn-Memory Computing (IMC) introduces a new paradigm of computation that offers high efficiency in terms of latency and power consumption for AI accelerators. However, the non-idealities and defects of emerging technologies used in advanced IMC can severely degrade the accuracy of inferred Neural Networks (NN) and lead to malfunctions in safety-critical applications. In this paper, we investigate an architectural-level mitigation technique based on the coordinated action of multiple checksum codes, to detect and correct errors at run-time. This implementation demonstrates higher efficiency in recovering accuracy across different AI algorithms and technologies compared to more traditional methods such as Triple Modular Redundancy (TMR). The results show that several configurations of our implementation recover more than 91% of the original accuracy with less than half of the area required by TMR and less than 40% of latency overhead. Luca Parrini, Taha Soliman, Benjamin Hettwer, Jan Micha Borrmann, Simranjeet Singh, Ankit Bende, Vikas Rana, Farhad Merchant, Norbert Wehn |
ETS | 9 |
| 2024 | FPGA Onboard Processing of Tiny-ML Models Using Radar-Sensing for Oil Spill MonitoringabstractDeep learning models have been widely used recently for environmental monitoring. Unless designed carefully, they are well known for their large computational complexity and long computation time, which limit their direct utilization as onsite embedded solutions. Therefore, to assure their suitability for onboard processing on drones during tactical responses, we propose in this paper energy-efficient and highly accurate tiny machine learning (TinyML) models in the field of radar remote sensing for oil spill monitoring. More precisely, the proposed signal processing models are based on artificial neural networks and use radar signals to detect oil slicks on top of the seawater and estimate their thicknesses using drones in calm ocean conditions. Despite their extremely small size (7-37 Bytes), their accuracy in detection and parameter estimation exceeds 94%. Moreover, the proposed models are characterized by low-power (32-173 mW) and low-latency (0.35-0.59 µs) performance when implemented on the FPGA computing platform. Bilal Hammoud, Jonas Ney, Charbel Bou-Mosleh, Norbert Wehn |
IGARSS | 4 |
| 2024 | Radar Backscattering Sensitivity to Emulsions Using Spectral Analysis from Nadir-Aerial Response SystemsabstractTo locate oil slicks, most state-of-the-art systems use synthetic aperture radar (SAR) imaging techniques at specific incident angles. However, only a few works present a detailed quantitative electromagnetic (EM) modeling of emulsified oil slicks. Therefore, in this paper, we investigate the electromagnetic modeling of selected oil emulsion models by extending their analysis over multiple SAR operating frequency bands and from a different operation point (from the nadir), which is required for the development of better monitoring solutions. The study presents a detailed spectral analysis of oil emulsions in all L-, C-, and X- frequency bands, and shows how their dielectric properties change differently according to the scanning EM wave frequency. We further highlight how the oil thickness introduces a cyclic behavior in the radar measurement, which affects pure oil slick detection at each frequency band. Finally, we investigate the radar backscattering sensitivity to emulsions, against clean water surfaces, for different percentages of water-in-oil and different slick thicknesses. Bilal Hammoud, Norbert Wehn |
IGARSS | 2 |
| 2024 | Do Radiation and Aging Impact DVFS? TCAD-based Analysis on 22 nm FDSOI Latches
Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif |
IOLTS | 3 |
| 2024 | Testing for aging in advanced SRAM: From front end of the line transistors to back end of the line interconnectsabstractThe long-term reliability of Static Random Access Memory (SRAM) is crucial for safety-critical applications, such as those in the automotive industry. In the front-end-of-line (FEoL), the transistor elements are susceptible to negative bias temperature instability (NBTI), while in the back-end-of-line (BEoL) the interconnects are susceptible to electromigration (EM), especially in scaled technology nodes. To meet safety-critical standards, it is essential to investigate the combined aging mechanisms within the SRAM array and to develop effective testing methodologies during the operational lifetime of the system. Such methodologies are also crucial for enabling the early detection of in-field failures. In this paper, a precise aging model is presented that extends the Technology Computer-Aided Design (TCAD) transistor model with a detailed NBTI model and includes physical modeling for EM. This approach provides insights into the combined effects of NBTI and EM on the degradation of SRAM writability, considering the entire SRAM subarray, including the bit-cell array and peripheral circuits in Fin Field-Effect Transistors (FinFET) technology. Mahta Mayahinia, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif, Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian |
ITC | 4 |
| 2024 | A Low-Power Linear Phase Interpolation-Based Delay Line in 12nm FinFET TechnologyabstractA novel low-power high-linear phase interpolation-based delay line in 12nm FinFET technology is detailed in this paper. The proposed delay line exhibits 50% improvement in terms of power consumption compared to the previous work. In addition, the presented architecture to the best of our knowledge is the most efficient delay line for fine tuning in advanced technology nodes due to the low complexity and complete controllability over resolution, delay range and target frequency. The analysis in this paper indicates that the input slew rate plays an indispensable role in the linearity of the delay line. Consequently, two identical resistors are added to the input of the phase interpolator unit to decrease the slew rate. This approach significantly improves the linearity over a wide frequency range. The proposed delay line dissipates 0.56 mW from a 0.8 V supply voltage and 5 GHz operating frequency. Mohammadreza Esmaeilpour, Jan Lappas, Christian Weis, Norbert Wehn |
VLSI-SoC | 4 |
| 2024 | Addressing the Combined Effect of Transistor and Interconnect Aging in SRAM towards Silicon Lifecycle ManagementabstractThe long-term reliability of the Static Random Access Memory (SRAM) module, as an important component of computing architectures, is crucial for safety-critical applications such as automotive. In the front end of the line (FEoL), the transistor elements are vulnerable to negative bias temperature instability (NBTI), while the back end of the line (BEoL) interconnect is prone to electromigration (EM). Complying with safety-critical standards as part of silicon lifecycle management (SLM) infrastructure requires an understanding of the combined aging mechanisms of transistors and interconnects in SRAM. Moreover, a precise aging model is a prerequisite for effective aging testing and mitigation strategies. For this aim, we augment the Technology Computer-Aided Design (TCAD) transistor model with a detailed NBTI model at the FEoL, and use measurement-calibrated physical modeling of EM at the BEoL, to create an integrated analysis that can provide deeper insights into the individual and combined effects of NBTI and EM for SRAM operation. Our findings reveal the mutual acceleration of delay faults and hard stuck-at faults caused by NBTI and EM in SRAM, offering a precise methodology for estimating the time to failure under these conditions. Mahta Mayahinia, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori, Sani R. Nassif, Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian |
VTS | 4 |
| 2024 | Trends in Channel Coding for 6GabstractError correction coding (i.e., channel coding) is a key ingredient of any digital communications system. In mobile wireless communications, channel codes have evolved from simple convolutional codes in Global System for Mobile Communications (GSM) (2G), parallel concatenated (turbo) codes in Universal Mobile Telecommunications Service (UMTS) (3G), and long-term evolution (LTE) (4G), to carefully designed multirate/multilength low-density parity-check (LDPC) codes in 5G, combined with polar codes for short messages on the synchronization channel. Based on this rich history, and by accounting for the technological advances in very large-scale integration, this article will outline some recent trends in channel coding as they may be applied in 6G systems, ranging from novel approaches for short blocklengths such as automorphism ensemble decoding, via ideas of coding for multiple access, to concepts for unified coding schemes that may simplify encoding/decoding hardware at competitive error-correcting performance. Sisi Miao, Claus Kestel, Lucas Johannsen, Marvin Geiselhart, Laurent Schmalen, Alexios Balatsoukas-Stimming, Gianluigi Liva, Norbert Wehn, Stephan ten Brink |
Proc. IEEE | 8 |
| 2023 | ZuSE Ki-Avf: Application-Specific AI Processor for Intelligent Sensor Signal Processing in Autonomous DrivingabstractModern and future AI-based automotive applications, such as autonomous driving, require the efficient real-time processing of huge amounts of data from different sensors, like camera, radar, and LiDAR. In the ZuSE-KI-AVF project, multiple university, and industry partners collaborate to develop a novel massive parallel processor architecture, based on a cus-tomized RISC-V host processor, and an efficient high-performance vertical vector coprocessor. In addition, a software development framework is also provided to efficiently program AI-based sensor processing applications. The proposed processor system was verified and evaluated on a state-of-the-art UltraScale+ FPGA board, reaching a processing performance of up to 126.9 FPS, while executing the YOLO-LITE CNN on 224x224 input images. Further optimizations of the FPGA design and the realization of the processor system on a 22nm FDSOI CMOS technology are planned. Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá, Christoph Riggers, Oliver Renke, Till Fiedler, Jakob Marten, Tobias Stuckenberg, Holger Blume, Christian Weis, Lukas Steiner, Chirag Sudarshan, Norbert Wehn, Lennart M. Reimann, Rainer Leupers, Michael Beyer, Daniel Köhler, Alisa Jauch, Jan Micha Borrmann, Setareh Jaberansari, Tim Berthold, Meinolf Blawat, Markus Kock, Gregor Schewior, Jens Benndorf, Frederik Kautz, Hans-Martin Blüthgen, Christian Sauer 0001 |
DATE | 13 |
| 2023 | Oil Spill Detection in Calm Ocean Conditions: A U-Net Model Novel SolutionabstractOil spills severely damage marine life and coastal environments. To reduce their polluting effect on the ecosystem, it is important to promptly react to potential spills for early detection and monitoring. In this paper, we propose a drone-based solution with a deep-learning U-net model. It processes the radar backscattering dominated by the specular component in calm ocean conditions to detect contaminated sea surfaces with oil spills. Results show that our approach achieves a high detection rate exceeding 90% for thick oil slicks in the range of 1-10 mm. Bilal Hammoud, Charbel Bou-Mosleh, Mohamed Moursi, Norbert Wehn |
IGARSS | 4 |
| 2023 | A Novel Iterative Estimation Technique Using Radar Sensing to Remotely Characterize Oil Slicks During SpillsabstractFor environmental and financial reasons, it is critical to develop new monitoring techniques that reduce the damage to the world’s marine ecosystems from oil spills. Information about the oil spill, such as the distribution of its thickness and its physical characteristics, will help in effective spill containment and tactical countermeasures. In this paper, based on radar sensing, we develop an iterative maximum-likelihood estimation approach to remotely extract both required information: the slick thickness and a physical characteristic of spilled slicks represented by the relative permittivity (dielectric constant). The targeted ranges are 1-10 mm thicknesses for thick oils, and 1.9-3.3 relative permittivities for light and crude oils. Results prove the performance accuracy of the proposed iterative approach with few iterations. Bilal Hammoud, Norbert Wehn |
IGARSS | 2 |
| 2023 | A Learning-Based Approach for Single Event Transient Analysis in Pass Transistor LogicabstractPass transistor logic (PTL) has emerged recently in advanced high-speed optical communication system due to its higher speed and lower power consumption compared to traditional CMOS logic. However, the sensitivity to radiation-induced soft errors of PTL implementations is significant different from CMOS circuitry, which emphasizes the need for understanding the mechanism of soft error propagation in PTL. Due to the non-conventional logic structure in PTL, previous approaches of pulse width modelling in CMOS logic are no more applicable since they are not always measurable. Hence, in this paper, we propose a learning-based structural regression modeling approach to explore the soft error propagation mechanism in PTL at transistor level. Our models can be easily mapped onto higher level to analyze soft error propagation in any complex PTL designs. The experimental results on a 4-bit ripple carry adder demonstrate that our models can achieve high accuracy compared with SPICE simulation. Zhihang Wu, Christian Weis, Norbert Wehn, Mehdi Baradaran Tahoori |
IOLTS | 4 |
| 2022 | Revisiting Pass-Transistor Logic Styles in a 12nm FinFET Technology NodeabstractWith the slow-down of Moore's law and the increasing requirements on energy efficiency, alternative logic styles compared to complementary static CMOS have to be revisited for digital circuit implementations. Pass Transistor Logic (PTL) gained much attention in the '90s, however, only a limited number of recent investigations and publications regarding PTL exist that use advanced technology nodes. This paper compares key performance metrics of 22 different PTL based 1-bit full adder designs to a complementary static CMOS logic reference, using a recent 12nm FinFET technology. The figures of merit are the propagation delay, the energy consumption, and the energy-delay-product (EDP). Our investigations show that PTL based adder circuits can have an up to 49% decreased delay and a 48% and 63% reduced energy consumption and EDP, respectively, compared to a state-of-the-art complementary CMOS logic reference. In addition, we analyzed the impact of PVT variations on the delay for selected PTL full adder designs. Jan Lappas, André Lucas Chinazzo, Christian Weis, Chenyang Xia, Zhihang Wu, Leibin Ni, Norbert Wehn |
DATE | 7 |
| 2022 | Machine learning based soft error rate estimation of pass transistor logic in high-speed communicationabstractRecent advanced high-speed communication systems, such as optical systems, require highest reliability at lowest possible power consumption. Thus, Pass Transistor Logic (PTL) is gaining lots of interest in these communication systems due to its power saving potential compared to traditional CMOS logic. However, due to the non-conventional logic structure, its susceptibility to radiation-induced soft errors is different from CMOS circuitry. Due to the unique generation and propagation of Single Event Transients (SETs) in PTL, different approaches for PTL soft error rate (SER) estimation are required. In this paper we propose a machine learning (ML) approach for SET propagation in PTL logic. Multi-layer feed-forward neural network together with support vector classifier (SVC) are used to build the SET pulse width and pulse amplitude models. Bayesian optimization using Gaussian Processes is utilized to tune the hyperparameters of neural network. The experimental results on full adder (FA), which is the key component in many large cirucits such as ALU, and comparison with Monte Carlo (MC) spectre simulations confirm the accuracy and speed of the proposed method. Jan Lappas, André Lucas Chinazzo, Christian Weis, Zhihang Wu, Leibin Ni, Norbert Wehn, Mehdi Baradaran Tahoori |
ETS | 7 |
| 2022 | FPGA-based Trainable Autoencoder for Communication SystemsabstractIn communication systems, autoencoder refers to a system that replaces parts of the traditional transmitter and receiver of the baseband processing chain with artificial neural networks (ANNs). This allows to jointly train the system for an underlying channel model by reconstructing the input symbols at the output. Since the actual behavior of a real communication channel cannot be perfectly reproduced by an abstract model, it is necessary for the autoencoder to adapt to the changing conditions at runtime. Thus, online fine-tuning, in the form of ANN-retraining is of great importance. A platform able to satisfy the low-latency and low-power requirements of embedded communication systems are Field-programmable gate arrays (FPGAs). In this paper, we present an online-trainable low-power FPGA architecture for the receiver of an autoencoder-based communication chain. The architecture is embedded into an exploration framework that automatically determines the optimal degree of parallelism to minimize latency or power consumption. Our solutions achieve 2000×higher throughput than a high-performance GPU, draw 5×less power than an embedded CPU and are 5800×more energy efficient compared to an embedded GPU, for a batch size of one. To the best of our knowledge, this is the first FPGA-based autoencoder implementation for communication systems. Jonas Ney, Sebastian Dörner, Matthias Herrmann, Mohammad Hassani Sadi, Jannis Clausius, Stephan ten Brink, Norbert Wehn |
FPGA | 7 |
| 2022 | Blind and Channel-agnostic Equalization Using Adversarial NetworksabstractDue to the rapid development of autonomous driving, the Internet of Things and streaming services, modern communication systems have to cope with varying channel conditions and a steadily rising number of users and devices. This, and the still rising bandwidth demands, can only be met by intelligent network automation, which requires highly flexible and blind transceiver algorithms. To tackle those challenges, we propose a novel adaptive equalization scheme, which exploits the prosperous advances in deep learning by training an equalizer with an adversarial network. The learning is only based on the statistics of the transmit signal, so it is blind regarding the actual transmit symbols and agnostic to the channel model. The proposed approach is independent of the equalizer topology and enables the application of powerful neural network based equalizers. In this work, we prove this concept in simulations of different―both linear and nonlinear―transmission channels and demonstrate the capability of the proposed blind learning scheme to approach the performance of non-blind equalizers. Furthermore, we provide a theoretical perspective and highlight the challenges of the approach. Vincent Lauinger, Manuel Dossinger, Jonas Ney, Norbert Wehn, Laurent Schmalen |
GLOBECOM | 4 |
| 2022 | A Maximum A-Posteriori Probabilistic Approach using UAV-Nadir-Looking Wide-Band Radar for Remote Sensing Oil-Spill DetectionabstractIn this paper, we present a maximum a-posteriori probabilistic approach for oil spill detection at very low wind speeds using nadir-looking wide-band radar systems mounted on drones. Such platforms allow for to have radar measurements for calm ocean conditions when the winds' speed is very small challenging current state-of-the-art techniques used for oil spill detection. We study the detection accuracy by exploiting the variation in the distribution of radar power reflectivities from both C- and X-band for different oil thicknesses and electromagnetic wave frequencies. The joint probability density function (pdf)-based detector shows that by optimally combining reflectivity values evaluated at multiple scanning frequencies, the performance is boosted over the whole range of possible slick thicknesses. The probability of detection is further improved by running multiple scans of the scene. Bilal Hammoud, Norbert Wehn |
IGARSS | 2 |
| 2022 | Optimization of DRAM based PIM Architecture for Energy-Efficient Deep Neural Network TrainingabstractDeep Neural Network (DNN) training consumes high-energy. On the other hand, DNNs deployed on edge devices demand very high-energy efficiency. In this context, Processing-in-Memory (PIM) is an emerging compute paradigm that bridges the memory-computation gap to improve the energy-efficiency. DRAMs are one such memory type employed for designing energy-efficient PIM architectures for DNN training. One of the major issues of DRAM-PIM architectures designed for DNN training is the high number of internal data accesses within a bank between the memory arrays and the PIM computation units (e.g. 51% more than inference). These internal data accesses in the state-of-the-art DRAM PIM architectures consume very high energy compared to computation units. Hence, it is important to reduce the internal data access energy within the DRAM bank for further improving the energy efficiency of DRAMPIM architectures. We present three novel optimizations that together reduce the internal data access energy up to 81.54%. Our first optimization modifies the bank data access circuit to enable partial accesses of data instead of the conventional fixed granularity accesses, thereby exploiting the available sparsity during training. The second optimization is to have a dedicated low-energy region within the DRAM bank that has low capacitive load of global wires and shorter data movement. Finally, we propose a 12-bit high dynamic range floating-point format called TinyFloat that reduces the total number of data access energy by 20% compared to IEEE 754 half and single precision. Chirag Sudarshan, Mohammad Hassani Sadi, Christian Weis, Norbert Wehn |
ISCAS | 4 |
| 2022 | FeFET versus DRAM based PIM Architectures: A Comparative StudyabstractThe throughput and energy efficiency of compute-centric architectures for memory intensive Deep Neural Networks (DNN) applications are limited by memory bound issues like high data-access energy, long latencies, and limited bandwidth. Processing-in-Memory (PIM) is a very promising approach to address these challenges and bridge the memory-computation gap. PIM places computational logic inside the memory to exploit minimum data movement and massive internal data parallelism. There are currently two PIM trends: 1) Use of emerging non-volatile memories to perform highly parallel analog computation of MAC operations and implicit storage of weights within the memory arrays, and 2) exploiting mature memory technologies that are enhanced by additional logic to enable efficient computation of MAC operations near the memory arrays. In this paper, we will compare both trends from an architectural perspective. Our study mainly emphasizes on FeFET memories (an emerging memory candidate) and DRAM memories (a mature memory candidate). We will highlight the major architectural constraints of these memory candidates that impact the PIM designs and their overall performance. Finally, we will assess feasible choice of candidate for different computations or DNN task types. Chirag Sudarshan, Taha Soliman, Thomas Kämpfe, Christian Weis, Norbert Wehn |
VLSI-SoC | 5 |
| 2022 | Spatially Coupled Serially Concatenated Codes: Performance Evaluation and VLSI Design TradeoffsabstractSpatially coupled serially concatenated codes (SC-SCCs) are constructed by coupling several classical turbo-like component codes. The resulting spatially coupled codes provide a close-to-capacity performance and low error floor, which have attracted a lot of interest in the past few years. The aim of this paper is to perform a comprehensive design space exploration to reveal different aspects of SC-SCCs, which is missing in the literature. More specifically, we investigate the effect of block length, coupling memory, decoding window size, and number of iterations on the decoding performance, complexity, latency, and throughput of SC-SCCs. To this end, we propose two decoding algorithms for the SC-SCCs:block-wiseandwindow-wisedecoders. For these, we present VLSI architectural templates and explore them based on building blocks implemented in 12nm FinFET technology. Linking architectural templates with the new algorithms, we demonstrate various tradeoffs between throughput, silicon area, latency, and decoding performance. Mojtaba Mahdavi 0001, Stefan Weithoffer, Matthias Herrmann, Liang Liu 0002, Ove Edfors, Norbert Wehn, Michael Lentmaier |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNN AccelerationabstractToday, a large number of applications depend on deep neural networks (DNN) to process data and perform complicated tasks at restricted power and latency specifications. Therefore, processing-in-memory (PIM) platforms are actively explored as a promising approach to improve the throughput and the energy efficiency of DNN computing systems. Several PIM architectures adopt resistive non-volatile memories as their main unit to build crossbar-based accelerators for DNN inference. However, these structures suffer from several drawbacks such as reliability, low accuracy, large ADCs/DACs power consumption and area, high write energy, and so on. In this article, we present a new mixed-signal in-memory architecture based on the bit-decomposition of the multiply and accumulate (MAC) operations. Our in-memory inference architecture uses a single FeFET as a non-volatile memory cell. Compared to the prior work, this system architecture provides a high level of parallelism while using only 3-bit ADCs. Also, it eliminates the need for any DAC. In addition, we provide flexibility and a very high utilization efficiency even for varying tasks and loads. Simulations demonstrate that we outperform state-of-the-art efficiencies with 36.5 TOPS/W and can pack 2.05 TOPS with 8-bit activation and 4-bit weight precision in an area of 4.9 mm 2 using 22 nm FDSOI technology. Employing binary operation, we obtain 1169 TOPS/W and over 261 TOPS/W/mm 2 on system level. Taha Soliman, Nellie Laleni, Tobias Kirchner, Franz Müller 0001, Thomas Kämpfe, Andre Guntoro, Norbert Wehn |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2022 | When Massive GPU Parallelism Ain't Enough: A Novel Hardware Architecture of 2D-LSTM Neural NetworkabstractMultidimensional Long Short-Term Memory (MD-LSTM) neural network is an extension of one-dimensional LSTM for data with more than one dimension. MD-LSTM achieves state-of-the-art results in various applications, including handwritten text recognition, medical imaging, and many more. However, its implementation suffers from the inherently sequential execution that tremendously slows down both training and inference compared to other neural networks. The main goal of the current research is to provide acceleration for inference of MD-LSTM. We advocate that Field-Programmable Gate Array (FPGA) is an alternative platform for deep learning that can offer a solution when the massive parallelism of GPUs does not provide the necessary performance required by the application. In this article, we present the first hardware architecture for MD-LSTM. We conduct a systematic exploration to analyze a tradeoff between precision and accuracy. We use a challenging dataset for semantic segmentation, namely historical document image binarization from the DIBCO 2017 contest and a well-known MNIST dataset for handwritten digit recognition. Based on our new architecture, we implement FPGA-based accelerators that outperform Nvidia Geforce RTX 2080 Ti with respect to throughput by up to 9.9 and Nvidia Jetson AGX Xavier with respect to energy efficiency by up to 48 . Our accelerators achieve higher throughput, energy efficiency, and resource efficiency than FPGA-based implementations of convolutional neural networks (CNNs) for semantic segmentation tasks. For the handwritten digit recognition task, our FPGA implementations provide higher accuracy and can be considered as a solution when accuracy is a priority. Furthermore, they outperform earlier FPGA implementations of one-dimensional LSTMs with respect to throughput, energy efficiency, and resource efficiency. Vladimir Rybalkin, Jonas Ney, Menbere Tekleyohannes, Norbert Wehn |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2021 | A Novel DRAM-Based Process-in-Memory Architecture and its Implementation for CNNsabstractProcessing-in-Memory (PIM) is an emerging approach to bridge the memory-computation gap. One of the key challenges of PIM architectures in the scope of neural network inference is the deployment of traditional area-intensive arithmetic multipliers in memory technology, especially for DRAM-based PIM architectures. Hence, existing DRAM PIM architectures are either confined to binary networks or exploit the analog property of the sub-array bitlines to perform bulk bit-wise logic operations. The former reduces the accuracy of predictions, i.e. Quality-of-results, while the latter increases overall latency and power consumption. Chirag Sudarshan, Taha Soliman, Cecilia De la Parra, Christian Weis, Leonardo Ecco, Matthias Jung 0001, Norbert Wehn, Andre Guntoro |
ASP-DAC | 7 |
| 2021 | HALF: Holistic Auto Machine Learning for FPGAsabstractDeep Neural Networks (DNNs) are capable of solving complex problems in domains related to embedded systems, such as image and natural language processing. To efficiently implement DNNs on a specific FPGA platform for a given cost criterion, e.g., energy efficiency, an enormous amount of design parameters must be considered from the topology down to the final hardware implementation. Interdependencies between the different design layers must be taken into account and explored efficiently, making it hardly possible to find optimized solutions manually. An automatic, holistic design approach can improve the quality of DNN implementations on FPGA significantly. To this end, we present a cross-layer design space exploration methodology. It comprises optimizations starting from a hardware-aware topology search for DNNs down to the final optimized implementation for a given FPGA platform. The methodology is implemented in our Holistic Auto machine Learning for FPGAs (HALF) framework, which combines an evolutionary search algorithm, various optimization steps, and a library of parametrizable hardware DNN modules. HALF automates both the exploration process and the implementation of optimized solutions on a target FPGA platform for various applications. We demonstrate the performance of HALF on a medical use case for arrhythmia detection for three different design goals, i.e., low-energy, low-power, and high-throughput. Our FPGA implementation outperforms a TensorRT optimized model on an Nvidia Jetson platform in both throughput and energy consumption. Jonas Ney, Dominik Marek Loroch, Vladimir Rybalkin, Nico Weber, Jens Krüger 0004, Norbert Wehn |
FPL | 6 |
| 2021 | ADMM-Based ML Decoding: from Theory to PracticeabstractInteger Linear Programming (ILP) is a general method to solve the Maximum-Likelihood (ML) decoding problem for all kinds of binary linear codes. To this end, state-of-the-art techniques use a Branch-and-Bound (B&B) framework to partition the underlying integer linear problem into several relaxed linear problems. These linear problems then have to be solved in reasonable time by an efficient Linear Programming (LP) solver. Recently, the Alternating Direction Method of Multipliers (ADMM) has been proposed for efficient software and hardware LP decoding of sparse codes, hence, an ADMM-based ML decoder seems to be a promising approach. In this paper, we investigate this approach with respect to its algorithmic and implementation-specific challenges. Kira Kraft, Norbert Wehn |
ICASSP | 2 |
| 2021 | Exploiting Resiliency for Kernel-Wise CNN Approximation Enabled by Adaptive Hardware DesignabstractEfficient low-power accelerators for Convolutional Neural Networks (CNNs) largely benefit from quantization and approximation, which are typically applied layer-wise for efficient hardware implementation. In this work, we present a novel strategy for efficient combination of these concepts at a deeper level, which is at each channel or kernel. We first apply layer-wise, low bit-width, linear quantization and truncation-based approximate multipliers to the CNN computation. Then, based on a state-of-the-art resiliency analysis, we are able to apply a kernel-wise approximation and quantization scheme with negligible accuracy losses, without further retraining. Our proposed strategy is implemented in a specialized framework for fast design space exploration. This optimization leads to a boost in estimated power savings of up to 34% in residual CNN architectures for image classification, compared to the base quantized architecture. Cecilia De la Parra, Ahmed El-Yamany, Taha Soliman, Akash Kumar 0001, Norbert Wehn, Andre Guntoro |
ISCAS | 5 |
| 2020 | Efficient FeFET Crossbar Accelerator for Binary Neural NetworksabstractThis paper presents a novel ferroelectric field-effect transistor (FeFET) in-memory computing architecture dedicated to accelerate Binary Neural Networks (BNNs). We present in-memory convolution, batch normalization and dense layer processing through a grid of small crossbars with reduced unit size, which enables multiple bit operation and value accumulation. Additionally, we explore the possible operations parallelization for maximized computational performance. Simulation results show that our new architecture achieves a computing performance up to 2.46 TOPS while achieving a high power efficiency reaching 111.8 TOPS/Watt and an area of 0.026 mm2in 22nm FDSOI technology. Taha Soliman, Ricardo Olivo, Tobias Kirchner, Cecilia De la Parra, Maximilian Lederer, Thomas Kämpfe, Andre Guntoro, Norbert Wehn |
ASAP | 8 |
| 2020 | Analysis and Optimization of TLS-based Security Mechanisms for Low Power IoT SystemsabstractSecurity has been valued as one of the most critical issues in the Internet of Things (IoT). The usage of cryptographic hardware accelerators is state of the art to improve the energy efficiency of security-related protocols like TLS. This way, feasible battery runtimes are ensured for resource-constrained IoT edge devices.However, the enhancement provided by the accelerator is not propagated in the same magnitude to the TLS handshake process.Hence, to further improve the security mechanism at TLS level, a holistic analysis of the IoT system is required to track down bottlenecks that have been revealed due to unilateral cryptographic optimization.In this paper, we present appropriate methods to identify further limiting factors and propose corresponding optimizations. We show how that overall energy demand for ephemeral TLS handshakes can be reduced by 50% and the related latency by about one order of magnitude compared to related work. Frederik Lauer, Carl Christian Rheinländer, Claus Kestel, Norbert Wehn |
CCGRID | 4 |
| 2020 | Fast and Accurate DRAM Simulation: Can we Further Accelerate it?abstractThe simulation of Dynamic Random Access Memories (DRAMs) in a system context requires highly accurate models due to the complex timing and power behavior of DRAMs. However, cycle accurate DRAM models often become the bottleneck regarding the overall simulation time. Therefore, fast but accurate DRAM simulation models are mandatory. This paper proposes two new performance optimized DRAM models that further accelerate the simulation speed with only a negligible degradation in accuracy. The first model is an enhanced Transaction Level Model (TLM), which uses a look-up table to accelerate parts of the simulation that feature a high memory access density for online scenarios. The second model is a neural network based simulator for offline trace analysis. We show a mathematical methodology to generate the inputs for the Look-Up Table (LUT) and an optimized artificial training set for the neural network. The enhanced TLM model is up to 5 times faster compared to a state-of-the-art TLM DRAM simulator. The neural network is able to speed up the simulation up to a factor of 10×, while inferring on a GPU. Both solutions provide only a slight decrease in accuracy of approximately 5%. Johannes Feldmann, Kira Kraft, Lukas Steiner, Norbert Wehn, Matthias Jung 0001 |
DATE | 4 |
| 2020 | TLS-Level Security for Low Power Industrial IoT Network InfrastructuresabstractThe Industrial Internet of Things (IIoT) enables communication services between machinery and cloud to enhance industrial processes e.g. by collecting relevant process parameters or providing predictable maintenance. Since the data is often origin from critical infrastructures, the security of the data channel is the main challenge, and is often weakened due to limited compute power and energy availability of battery-powered sensor nodes. Lightweight alternatives to standard security protocols avoid computationally intensive algorithms, however, they do not provide the same level of trust as established standards such as Transport Layer Security (TLS).In this paper, we propose an IIoT network system that enables a secure end-to-end IP communication between ultra-low-power sensor nodes and cloud servers. It provides full TLS support to ensure perfect forward secrecy by using hardware accelerators to reduce the energy demand of the security algorithms. Our results show that the energy overhead of the TLS handshake can be significantly reduced to enable a secure IIoT infrastructure with a reasonable battery lifetime of the edge devices. Jochen Mades, Gerd Ebelt, Boris Janjic, Frederik Lauer, Carl Christian Rheinländer, Norbert Wehn |
DATE | 6 |
| 2020 | When Massive GPU Parallelism Ain't Enough: A Novel Hardware Architecture of 2D-LSTM Neural NetworkabstractMultidimensional Long Short-Term Memory (MD-LSTM) neural network is an extension of one-dimensional LSTM for data with more than one dimension that allows MD-LSTM to show state-of-the-art results in various applications including handwritten text recognition, medical imaging, and many more. However, efficient implementation suffers from very sequential execution that tremendously slows down both training and inference compared to other neural networks. This is the primary reason that prevents intensive research involving MD-LSTM in the recent years, despite large progress in microelectronics and architectures. The main goal of the current research is to provide acceleration for inference of MD-LSTM, so to open a door for efficient training that can boost application of MD-LSTM. By this research we advocate that FPGA is an alternative platform for deep learning that can offer a solution in cases when a massive parallelism of GPUs does not provide the necessary performance required by the application. In this paper, we present the first hardware architecture for MD-LSTM. We conduct a systematic exploration of precision vs. accuracy trade-off using challenging dataset for historical document image binarization from DIBCO 2017 contest, and well known MNIST dataset for handwritten digits recognition. Based on our new architecture we implement FPGA-based accelerator that outperforms NVIDIA K80 GPU implementation in terms of runtime by up to 50x and energy efficiency by up to 746x. At the same time, our accelerator demonstrates higher accuracy and comparable throughput in comparison with state-of-the-art FPGA-based implementations of multilayer perceptron for MNIST dataset. Vladimir Rybalkin, Norbert Wehn |
FPGA | 2 |
| 2020 | Fully Pipelined Iteration Unrolled Decoders the Road to TB/S Turbo DecodingabstractTurbo codes are a well-known code class used for example in the LTE mobile communications standard. They provide built-in rate flexibility and a low-complexity and fast encoding. However, the serial nature of their decoding algorithm makes high-throughput hardware implementations difficult. In this paper, we present recent findings on the implementation of ultra-high throughput Turbo decoders. We illustrate how functional parallelization at the iteration level can achieve a throughput of several hundred Gb/s in 28 nm technology. Our results show that, by spatially parallelizing the half-iteration stages of fully pipelined iteration unrolled decoders into X-windows of size 32, an area reduction of 40% can be achieved. We further evaluate the area savings through further reduction of the X-window size. Lastly, we show how the area complexity and the throughput of the fully pipelined iteration unrolled architecture scale to larger frame sizes. We consider the same target bit error rate performance for all frame sizes and highlight the direct correlation to area consumption. Stefan Weithoffer, Rami Klaimi, Charbel Abdel Nour, Norbert Wehn, Catherine Douillard |
ICASSP | 4 |
| 2020 | Access-Aware Per-Bank DRAM Refresh for Reduced DRAM Refresh OverheadabstractThe performance and energy penalties of DRAM refresh have increased in successive generations of higher capacity DRAM devices. This trend is likely to continue in future systems where the internal DRAM refresh cycle is opaque to the memory controller and the memory device oblivious of context. This paper presents Access-Aware Per-bank DRAM Refresh, a refresh control method that mitigates the negative impacts of refreshes, and its memory controller architecture. Novel capabilities are introduced in the memory controller. An access-aware refresh control unit analyses the short-term history of memory accesses translating row activations into refresh masks. Refresh masks are used either to skip rows, shortening the refresh cycle, or to completely omit refresh operations. The set of DRAM commands is extended with two new per-bank refresh commands that provide the memory controller not only an ability to omit refreshes, but also a context-rich fine-grained control of refresh operations. A proof of concept model of our architecture is implemented in a virtual platform where a set of applications is used to exercise the memory subsystem. Evaluations show that for the workloads considered the proposed architecture and refresh control method improve, either by reducing the latency or by completely omitting, up to 19% of the refresh operations. Éder Zulian, Christian Weis, Norbert Wehn |
ISCAS | 3 |
| 2020 | A 506Gbit/s Polar Successive Cancellation List Decoder with CRCabstractPolar codes have recently attracted significant attention due to their excellent error-correction capabilities. However, efficient decoding of Polar codes for high throughput is very challenging. Beyond 5G, data rates towards 1Tbit/s are expected. Low complexity decoding algorithms like Successive Cancellation (SC) decoding enable such high throughput but suffer on errorcorrection performance. Polar Successive Cancellation List (SCL) decoders, with and without Cyclic Redundancy Check (CRC), exhibit a much better error-correction but imply higher implementation cost. In this paper we in-depth investigate and quantify various trade-offs of these decoding algorithms with respect to error-correction capability and implementation costs in terms of area, throughput and energy efficiency in a 28nm CMOS FD-SOI technology. We present a framework that automatically generates decoder architectures for throughputs beyond 100Gbit/s. This framework includes various architectural optimizations for SCL decoders that go beyond State-of-the-Art. We demonstrate a 506Gbit/s SCL decoder with CRC that was generated by this framework. Claus Kestel, Lucas Johannsen, Oliver Griebel, Jhon Jimenez, Timo Vogt, Timo Lehnigk-Emden, Norbert Wehn |
PIMRC | 7 |
| 2020 | Low-complexity Computational Units for the Local-SOVA Decoding AlgorithmabstractRecently the Local-SOVA algorithm was suggested as an alternative to the max-Log MAP algorithm commonly used for decoding Turbo codes. In this work, we introduce new complexity reductions to the Local-SOVA algorithm, which allow an efficient implementation at a marginal BER penalty of 0.05 dB. Furthermore, we present the first hardware architectures for the computational units of the Local-SOVA algorithm, namely for the add-compare select unit and the soft output unit, targeting radix orders 2, 4 and 8. We provide place & route implementation results for 22nm technology and demonstrate an area reduction of 46-75% for the soft output unit for radix orders ≥ 4 in comparison with the respective max-Log MAP soft output unit. These area reductions compensate for the overhead in the add compare select unit, resulting in overall area saving of around 27-46% compared to the max-Log-MAP. These savings simplify the design and implementation of high throughput Turbo decoders. Stefan Weithoffer, Rami Klaimi, Charbel Abdel Nour, Norbert Wehn, Catherine Douillard |
PIMRC | 4 |
| 2020 | Advanced Hardware Architectures for Turbo Code Decoding Beyond 100 Gb/sabstractIn this paper, we present two new hardware architectures for Turbo Code decoding that combine functional, spatial and iteration parallelism. Our first architecture is the first fully pipelined iteration unrolled architecture that supports multiple frame sizes. This frame flexibility is achieved by providing a set of interleavers designed to achieve a hardware implementation with a reduced routing overhead. The second architecture efficiently utilizes the dynamics of the error rate distribution for different decoding iterations and is comprised of two stages. First, a fully pipelined iteration unrolled decoder stage applied for a pre-determined number of iterations and a second stage with an iterative afterburner-decoder activated only for frames not successfully decoded by the first stage. We give post place & route results for implementations of both architectures for a maximum frame size of K = 128 and demonstrate a throughput of 102.4 Gb/s in 2S nm FDSOI technology. With an area efficiency of 6.19 and 7.15 Gb/s/m$m^{2}$ our implementations clearly outperform state of the art. Stefan Weithoffer, Oliver Griebel, Rami Klaimi, Charbel Abdel Nour, Norbert Wehn |
WCNC | 5 |
| 2020 | Harvester-aware transient computing: Utilizing the mechanical inertia of kinetic energy harvesters for a proactive frequency-based power loss detection
Carl Christian Rheinländer, Norbert Wehn |
Integr. | 2 |
| 2020 | A Reduced-Complexity Projection Algorithm for ADMM-Based LP DecodingabstractThe alternating direction method of multipliers has recently been adapted for linear programming decoding of low-density parity-check codes. The computation of the projection onto the parity polytope is the core of this algorithm and usually involves a sorting operation, which is the main effort of the projection. In this paper, we present an algorithm with low complexity to compute this projection. The algorithm relies on new findings in the recursive structure of the parity polytope and iteratively fixes selected components. As shown in our realistic simulation setup, it requires up to 37% less arithmetical operations compared with state-of-the-art projections. Additionally, it does not involve a sorting operation, which is needed in all exact state-of-the-art projection algorithms. These two benefits make it appealing for efficient hardware and software implementations. Florian Gensheimer, Tobias Dietz, Kira Kraft, Stefan Ruzika, Norbert Wehn |
IEEE Trans. Inf. Theory | 5 |
| 2019 | Speculative Temporal Decoupling Using fork()abstractTemporal decoupling is a state-of-the-art method to speed up virtual prototypes. In this technique, a process is allowed to run ahead of simulation time for a specific interval called quantum. By using this method, the number of synchronization points, i.e. context switches, in the simulator is reduced and therefore, the simulation speed can be increased significantly. However, using this approach can introduce functional simulation errors due to missed synchronization events. Thus, using temporal decoupling implies a trade-off between speed and accuracy and the size of the quantum must be chosen wisely with respect to the simulated application. In loosely timed simulations most of the functional errors are tolerable for the sake of simulation speed. However, for instance safety critical errors are rare but can lead to fatal results and must be handled carefully. Prior works present mechanisms based on checkpoints (storing/restoring the internal state of the simulation model) in order to rollback in simulation time and correct the occurred errors by forcing synchronization. However, checkpointing approaches are intrusive and require changes to both the source code of all the used simulation models and the kernel of the simulator. In this paper we present a non-intrusive rollback approach for error-free temporal decoupling, which allows the usage of closed source models by using Unix’s fork() system call. Furthermore, we provide a case study based on the IEEE simulation standard SystemC. Matthias Jung 0001, Frank Schnicke, Markus Damm, Thomas Kuhn 0001, Norbert Wehn |
DATE | 5 |
| 2019 | Polar Code Decoder FrameworkabstractPolar codes gained large interest in the last years since they are the first channel codes that are proven to achieve channel capacity. Due to this property, Polar codes were recently adopted for the 5G standard. We present an industrial framework for the generation of Polar code decoders for highest data throughput. The framework automatically generates VHDL models ready for synthesis, placement and routing and corresponding simulation models to assess the communications performance. This framework enables Polar code decoder IP providers to give fast feedback to customers on communications and implementation performance. We demonstrate that this framework outperforms existing manually optimized decoders especially in terms of energy efficiency. Timo Lehnigk-Emden, Matthias Alles, Claus Kestel, Norbert Wehn |
DATE | 4 |
| 2019 | An In-DRAM Neural Network Processing EngineabstractMany advanced neural network inference engines are bounded by the available memory bandwidth. The conventional approach to address this issue is to employ high bandwidth memory devices or to adapt data compression techniques (reduced precision, sparse weight matrices). Alternatively, an emerging approach to bridge the memory-computation gap and to exploit extreme data parallelism is Processing in Memory (PIM). The close proximity of the computation units to the memory cells reduces the amount of external data transactions and it increases the overall energy efficiency of the memory system. In this work, we present a novel PIM based Binary Weighted Network (BWN) inference accelerator design that is inline with the commodity Dynamic Random Access Memory (DRAM) design and process. In order to exploit data parallelism and minimize energy, the proposed architecture integrates the basic BWN computation units at the output of the Primary Sense Amplifiers (PSAs) and the rest of the substantial logic near the Secondary Sense Amplifiers (SSAs). The power and area values are obtained at sub-array (SA) level using exhaustive circuit level simulations and full-custom layout. The proposed architecture results in an area overhead of 25 % compared to a commodity 8 Gb DRAM and delivers a throughput of 63.59 FPS (Frames per Second) for AlexNet. We also demonstrate that our architecture is extremely energy efficient, 7.25× higher FPS/W, as compared to previous works. Chirag Sudarshan, Jan Lappas, Muhammad Mohsin Ghaffar, Vladimir Rybalkin, Christian Weis, Matthias Jung 0001, Norbert Wehn |
ISCAS | 7 |
| 2018 | Improving the error behavior of DRAM by exploiting its Z-channel propertyabstractIn this paper, we present a new communication theoretic channel model for Dynamic Random Access Memory (DRAM) retention errors, that relies on the fully asymmetric retention error behavior of DRAM cells. This new model shows that the traditional approach is over pessimistic and we confirm this with real measurements of DDR3 and DDR4 DRAM devices. Together with an exploitation of the vendor specific true- and anti-cell structure, a low complexity bit-flipping approach is presented, that can largely increase DRAM's reliability with minimum overhead. Kira Kraft, Chirag Sudarshan, Deepak M. Mathew, Christian Weis, Norbert Wehn, Matthias Jung 0001 |
DATE | 5 |
| 2018 | The transprecision computing paradigm: Concept, design, and applicationsabstractGuaranteed numerical precision of each elementary step in a complex computation has been the mainstay of traditional computing systems for many years. This era, fueled by Moore's law and the constant exponential improvement in computing efficiency, is at its twilight: from tiny nodes of the Internet-of-Things, to large HPC computing centers, sub-picoJoule/operation energy efficiency is essential for practical realizations. To overcome the power wall, a shift from traditional computing paradigms is now mandatory. In this paper we present the driving motivations, roadmap, and expected impact of the European project OPRECOMP. OPRECOMP aims to (i) develop the first complete transprecision computing framework, (ii) apply it to a wide range of hardware platforms, from the sub-milliWatt up to the MegaWatt range, and (iii) demonstrate impact in a wide range of computational domains, spanning IoT, Big Data Analytics, Deep Learning, and HPC simulations. By combining together into a seamless design transprecision advances in devices, circuits, software tools, and algorithms, we expect to achieve major energy efficiency improvements, even when there is no freedom to relax end-to-end application quality of results. Indeed, OPRECOMP aims at demolishing the ultra-conservative “precise” computing abstraction, replacing it with a more flexible and efficient one, namely transprecision computing. Cristiano Malossi, Michael Schaffner, Anca Mariana Molnos, Luca Gammaitoni, Giuseppe Tagliavini, Andrew P. J. Emerson, Andrés Tomás, Dimitrios S. Nikolopoulos, Eric Flamand, Norbert Wehn |
DATE | 10 |
| 2018 | An analysis on retention error behavior and power consumption of recent DDR4 DRAMsabstractDRAM technology is scaling aggressively that results in high leakage power, worse data retention time behavior, and large process variations. Due to these process variations, vendors provide large guard bands on various DRAM currents and timing specifications that are over pessimistic. Detailed knowledge on the DRAM retention behavior and currents for the average case allow to improve memory system performance and energy efficiency of specific applications by moving away from worst case behavior. In this paper, we present an advanced measurement platform to investigate off-the-shelf DDR4 DRAMs' retention behavior, and to precisely measure various DRAM currents (IDDs and IPPs) at a wide range of operating temperatures. Error Checking and Correction (ECC) schemes are popular in correcting randomly scattered single bit errors. Since retention failures also occur randomly, ECCs can be used to improve DRAM retention behavior. Therefore, for the first time, we show the influence of ECC on the retention behavior of recent DDR4 DRAMs, and how it varies across various DRAM architectures considering detailed structure of the DRAM (true-cell devices/mixed-cell devices). Deepak M. Mathew, Martin Schultheis, Carl Christian Rheinländer, Chirag Sudarshan, Christian Weis, Norbert Wehn, Matthias Jung 0001 |
DATE | 6 |
| 2018 | iDocChip: A Configurable Hardware Architecture for Historical Document Image Processing: Percentile Based BinarizationabstractEnd-to-end Optical Character Recognition (OCR) systems are heavily used to convert document images into machine-readable text. Commercial and open-source OCR systems (like Abbyy, OCRopus, Tesseract etc.) have traditionally been optimized for contemporary documents like books, letters, memos, and other end-user documents. However, these systems are difficult to use equally well for digitizing historical document images, which contain degradations like non-uniform shading, bleed-through, and irregular layout; such degradations usually do not exist in contemporary document images. Vladimir Rybalkin, Syed Saqib Bukhari, Muhammad Mohsin Ghaffar, Aqib Ghafoor, Norbert Wehn, Andreas Dengel 0001 |
DocEng | 5 |
| 2018 | FINN-L: Library Extensions and Design Trade-Off Analysis for Variable Precision LSTM Networks on FPGAsabstractIt is well known that many types of artificial neural networks, including recurrent networks, can achieve a high classification accuracy even with low-precision weights and activations. The reduction in precision generally yields much more efficient hardware implementations in regards to hardware cost, memory requirements, energy, and achievable throughput. In this paper, we present the first systematic exploration of this design space as a function of precision for Bidirectional Long Short-Term Memory (BiLSTM) neural network. Specifically, we include an in-depth investigation of precision vs. accuracy using a fully hardware-aware training flow, where during training quantization of all aspects of the network including weights, input, output and in-memory cell activations are taken into consideration. In addition, hardware resource cost, power consumption and throughput scalability are explored as a function of precision for FPGA-based implementations of BiLSTM, and multiple approaches of parallelizing the hardware. We provide the first open source HLS library extension of FINN for parameterizable hardware architectures of LSTM layers on FPGAs which offers full precision flexibility and allows for parameterizable performance scaling offering different levels of parallelism within the architecture. Based on this library, we present an FPGA-based accelerator for BiLSTM neural network designed for optical character recognition, along with numerous other experimental proof points for a Zynq UltraScale+ XCZU7EV MPSoC within the given design space. Vladimir Rybalkin, Alessandro Pappalardo, Muhammad Mohsin Ghaffar, Giulio Gambardella, Norbert Wehn, Michaela Blott |
FPL | 5 |
| 2018 | The Role of Memories in Transprecision ComputingabstractComputing paradigms largely evolved over the last decades mainly driven by continuously increasing performance requirements, energy efficiency and power density challenges. Heterogeneous highly parallel architectures enhanced with dedicated accelerators tuned to specific applications, near-threshold computing, and recently approximate computing are examples of these new approaches. In this context the memory part was relatively untouched. However, memories play a central role in any computing system, are a major source of power consumption, and limit in many applications the overall compute performance. In this paper, we focus mainly on Dynamic Random Access Memories (DRAMs), which are today's most prominent external memories. In transprecision computing we address the DRAM memory challenge by several new approaches that are strongly related to the new techniques known on the compute side. In particular, these are the concept of approximate DRAM, advanced power-down modes and the integration of application knowledge into the memory system. Christian Weis, Matthias Jung 0001, Éder Zulian, Chirag Sudarshan, Deepak M. Mathew, Norbert Wehn |
ISCAS | 6 |
| 2018 | A Framework for Non-intrusive Trace-driven Simulation of Manycore Architectures with Dynamic Tracing Configuration
Jasmin Jahic, Matthias Jung 0001, Thomas Kuhn 0001, Claus Kestel, Norbert Wehn |
RV | 5 |
| 2017 | A Heterogeneous SDR MPSoC in 28 nm CMOS for Low-Latency Wireless ApplicationsabstractCurrent and future applications impose high demands on software-defined radio (SDR) platforms in terms of latency, reliability, and flexibility. This paper presents a heterogeneous SDR MPSoC with a hexagonal network-on-chip to address these issues. It features four data processing modules and a baseband processing engine for iterative multiple-input multiple-output (MIMO) receiving. Integrated memory controllers enable dynamic data flow mapping and application isolation. In a 4 x 4 MIMO application scenario, the MPSoC achieves a throughput of 232 Mbit/s with a latency of 20 μs while consuming 414 mW. It outperforms state-of-the-art platforms in terms of throughput by a factor of 4. Sebastian Haas, Tobias Seifert, Benedikt Noethen, Stefan Scholze, Sebastian Höppner, Andreas Dixius, Esther P. Adeva, Thomas R. Augustin, Friedrich Pauls, Sadia Moriam, Mattis Hasler, Erik Fischer, Yong Chen 0014, Emil Matús, Georg Ellguth, Stephan Hartmann 0002, Stefan Schiefer, Love Cederstroem, Dennis Walter, Stephan Henker, Stefan Hänzsche, Johannes Uhlig, Holger Eisenreich, Stefan Weithoffer, Norbert Wehn, René Schüffny, Christian Mayr 0001, Gerhard P. Fettweis |
DAC | 25 |
| 2017 | Hardware architecture of Bidirectional Long Short-Term Memory Neural Network for Optical Character RecognitionabstractOptical Character Recognition is conversion of printed or handwritten text images into machine-encoded text. It is a building block of many processes such as machine translation, text-to-speech conversion and text mining. Bidirectional Long Short-Term Memory Neural Networks have shown a superior performance in character recognition with respect to other types of neural networks. In this paper, to the best of our knowledge, we propose the first hardware architecture of Bidirectional Long Short-Term Memory Neural Network with Connectionist Temporal Classification for Optical Character Recognition. Based on the new architecture, we present an FPGA hardware accelerator that achieves 459 times higher throughput than state-of-the-art. Visual recognition is a typical task on mobile platforms that usually use two scenarios either the task runs locally on embedded processor or offloaded to a cloud to be run on high performance machine. We show that computationally intensive visual recognition task benefits from being migrated to our dedicated hardware accelerator and outperforms high-performance CPU in terms of runtime, while consuming less energy than low power systems with negligible loss of recognition accuracy. Vladimir Rybalkin, Norbert Wehn, Mohammad Reza Yousefi, Didier Stricker |
DATE | 2 |
| 2017 | An advanced embedded architecture for connected component analysis in industrial applicationsabstractIn recent years, connected component analysis (CCA) has become one of the vital image/video processing algorithms due to its wide-range applicability in the field of computer vision. Numerous applications such as pattern recognition, object detection and image segmentation involve connected component analysis. In the context of camera-based inspection systems, CCA plays an important role for quality assurance. State-of-the-art hardware architectures offer high performance implementations of CCA using field programmable gate arrays (FPGAs). However, due to their high memory-demand, most of these implementations inhibit a large resource utilization. In this paper, we propose a hybrid software-hardware architecture of CCA for an industrial application using Xilinx Zynq-7000 All Programmable System on Chip (SoC). By offloading the most resource consuming part of the algorithm to the embedded CPU, we achieved high performance, while reducing the required resources on the FPGA. Our proposed architecture saves more than 30% of on-chip memory (Block RAMs) compared to state-of-the-art hardware architectures without affecting the throughput. Furthermore, due to the embedded CPU, our system provides a versatile and highly flexible feature extraction at run-time without the necessity to reconfigure the FPGA. Menbere Tekleyohannes, MohammadSadegh Sadri, Christian Weis, Norbert Wehn, Martin Klein 0005, Michael Siegrist |
DATE | 4 |
| 2016 | Efficient reliability management in SoCs - an approximate DRAM perspectiveabstractIn today's computing systems Dynamic Random Access Memories (DRAMs) have a large influence on performance and contribute significantly to the total power consumption. Thus, recent research activities bring the idea of approximate DRAM into focus to save power and improve performance by lowering the refresh rate or disabling refresh completely. Hence, fast and accurate models are required for a thoroughly exploration of approximate DRAM for error resilient applications. In this paper we present a holistic simulation environment for investigations on approximate DRAM and show the impact on error resilient applications. Matthias Jung 0001, Deepak M. Mathew, Christian Weis, Norbert Wehn |
ASP-DAC | 4 |
| 2016 | Invited - Approximate computing with partially unreliable dynamic random access memory - approximate DRAMabstractIn the context of approximate computing, Approximate Dynamic Random Access Memory (ADRAM) enables the tradeoff between energy efficiency, performance and reliability. The inherent error resilience of applications allows sacrificing data storage robustness and stability by lowering the refresh rate or disabling refresh in DRAMs completely. Consequently, it is important to know exactly the statistical DRAM behavior with respect to retention time, process variation and temperature to manage this trade-off and thereby deliberately exploiting the error resilience of different target applications. Matthias Jung 0001, Deepak M. Mathew, Christian Weis, Norbert Wehn |
DAC | 4 |
| 2016 | Error resilience and energy efficiency: An LDPC decoder design study
Philipp Schläfer, Chu-Hsiang Huang, Clayton Schoeny, Christian Weis, Yao Li 0007, Norbert Wehn, Lara Dolecek |
DATE | 6 |
| 2016 | Saturated min-sum decoding: An "afterburner" for LDPC decoder hardware
Stefan Scholl, Philipp Schläfer, Norbert Wehn |
DATE | 3 |
| 2016 | A New Architecture for High Speed, Low Latency NB-LDPC Check Node Processing for GF(256)abstractNon-binary low-density parity-check codes have superior communications performance compared to their binary counterparts. However, to be an option for future standards, efficient hardware architectures are mandatory. State-of-the-art decoding algorithms result in architectures suffering from low throughput and high latency. The check node function accounts for the largest part of the decoders overall complexity. To the best of our knowledge, we propose the first architecture for high speed, low latency Non-Binary Low-Density Parity-Check Check Node processing for GF(256). It has state-of-the-art communications performance while largely reducing the hardware complexity. The presented architecture has a 3.3 times higher area efficiency, increases the energy efficiency by factor 2.5 and reduces the latency by factor of 5.5 compared to the first implementation of Check Node for GF(256) based on the state-of-the-art FWBW scheme that was also implemented in the scope of this work. Vladimir Rybalkin, Philipp Schläfer, Norbert Wehn |
VTC Spring | 3 |
| 2016 | Precision-tuning and hybrid pricer for closed-form solution-based Heston calibrationabstractSummary Calibration methods are the heart of modeling any financial process. While for the Heston model (semi) closed‐form solutions exist for simple products, their evaluation involves complex functions and infinite integrals. So far, these integrals can only be solved with time‐consuming numerical methods. For that reason, calibration consumes a large portion of available compute power in the daily finance business. However, more and more theoretical and practical subtleties have been discovered over the years, and today, a large number of calibration methods are available. Currently, there is no clear indication which numerical method should be used for a specific calibration purpose under given speed and accuracy constraints. With this publication, we aim at closing this gap. We derive a novel methodology for systematically finding the best methods for a well‐defined accuracy target. For a practical setup, we study the available popular closed‐form solutions and integration algorithms. In total, we compare 14 numerical methods, including adaptive quadrature and Fourier methods. For a target accuracy of 10−3, we show that adaptive Gauss–Kronrod methods are best on CPUs for the unrestricted parameter set. Furthermore, we introduce hybrid pricer methods that combine quadrature and fast Fourier transform pricers, what gives us another 2.4× speedup. Copyright © 2015 John Wiley & Sons, Ltd. Christian Brugger, Gongda Liu, Christian de Schryver, Norbert Wehn |
Concurr. Comput. Pract. Exp. | 4 |
| 2015 | Exploiting Phase Transitions for the Efficient Sampling of the Fixed Degree Sequence ModelabstractReal-world network data is often very noisy and contains erroneous or missing edges. These superfluous and missing edges can be identified statistically by assessing the number of common neighbors of the two incident nodes. To evaluate whether this number of common neighbors, the so called co-occurrence, is statistically significant, a comparison with the expected co-occurrence in a suitable random graph model is required. For networks with a skewed degree distribution, including most real-world networks, it is known that the fixed degree sequence model, which maintains the degrees of nodes, is favourable over using simplified graph models that are based on an independence assumption. However, the use of a fixed degree sequence model requires sampling from the space of all graphs with the given degree sequence and measuring the co-occurrence of each pair of nodes in each of the samples, since there is no known closed formula for this statistic. While there exist log-linear approaches such as Markov chain Monte Carlo sampling, the computational complexity still depends on the length of the Markov chain and the number of samples, which is significant in large-scale networks. In this article, we show based on ground truth data that there are various phase transition-like tipping points that enable us to choose a comparatively low number of samples and to reduce the length of the Markov chains without reducing the quality of the significance test. As a result, the computational effort can be reduced by an order of magnitudes. Christian Brugger, André Lucas Chinazzo, Alexandre Flores John, Christian de Schryver, Norbert Wehn, Andreas Spitz, Katharina A. Zweig |
ASONAM | 5 |
| 2015 | Reverse longstaff-schwartz american option pricing on hybrid CPU/FPGA systems
Christian Brugger, Javier Alejandro Varela, Norbert Wehn, Songyin Tang, Ralf Korn |
DATE | 3 |
| 2015 | Retention time measurements and modelling of bit error rates of WIDE I/O DRAM in MPSoCs
Christian Weis, Matthias Jung 0001, Peter Ehses, Cristiano Santos, Pascal Vivet, Sven Goossens, Martijn Koedam, Norbert Wehn |
DATE | 8 |
| 2015 | A new architecture for high throughput, low latency NB-LDPC check node processingabstractNon-binary low-density parity-check codes have superior communications performance compared to their binary counterparts. However, to be an option for future standards, efficient hardware architectures must be developed. State-of-the-art decoding algorithms lead to architectures suffering from low throughput and high latency. The check node function accounts for the largest part of the decoders overall complexity. In this paper a new hardware aware check node algorithm and its architecture is proposed. It has state-of-the-art communications performance while reducing the decoding complexity. The presented architecture has a 14 times higher area efficiency, increases the energy efficiency by factor 2.5 and reduces the latency by factor of 3.5 compared to state-of-the-art architectures. Philipp Schläfer, Vladimir Rybalkin, Norbert Wehn, Matthias Alles, Timo Lehnigk-Emden, Emmanuel Boutillon |
PIMRC | 3 |
| 2015 | Latency reduction for LTE/LTE-A turbo-code decoders by on-the-fly calculation of CRCabstractThe major challenge for the development of Turbo-Code decoders for the 3GPP LTE/LTE-A standard is the support of very high code rates while achieving a high throughput with a low decoding latency. For high code rates the decoding process can oscillate, i.e. although the decoder converges to a valid code word in the nth half-iteration, the code word after the n+1th half iteration can be invalid again. Thus, the CRC must be evaluated after every half-iteration to avoid a loss in communications performance. We propose a new CRC calculation scheme that takes into account requirements for parallel Turbo-Code decoding for LTE/LTE-A by calculating the CRC On-the-fly not only during the non-interleaved but also during the interleaved half-iterations. Further, we present a flexible hardware architecture that allows a reduction of the latency for the CRC calculation of up to 34× and is two times more area efficient with respect to latency when compared to State-of-the-Art. Stefan Weithoffer, Norbert Wehn |
PIMRC | 2 |
| 2014 | Mixed precision multilevel Monte Carlo on hybrid computing systemsabstractNowadays, high-speed computations are mandatory for financial and insurance institutes to survive in competition and to fulfill the regulatory reporting requirements that have just toughened over the last years. A majority of these computations are carried out on huge computing clusters, which are an ever increasing cost burden for the financial industry. There, state-of-the-art CPU and GPU architectures execute arithmetic operations with pre-defined precisions only, that may not meet the actual requirements for a specific application. Reconfigurable architectures like field programmable gate arrays (FPGAs) have a huge potential to accelerate financial simulations while consuming only very low energy by exploiting dedicated precisions in optimal ways. In this work we present a novel methodology to speed up multilevel Monte Carlo (MLMC) simulations on reconfigurable architectures. The idea is to aggressively lower the precisions for different parts of the algorithm without loosing any accuracy at the end. For this, we have developed a novel heuristic for selecting an appropriate precision at each stage of the simulation that can be executed with low costs at runtime. Further, we introduce a cost model for reconfigurable architectures and minimize the cost of our algorithm without changing the overall error. We consider the showcase of pricing Asian options in the Heston model. For this setup we improve one of the most advanced simulation methods by a factor of 3-9x on the same platform. Christian Brugger, Christian de Schryver, Norbert Wehn, Steffen Omland, Mario Hefter, Klaus Ritter 0001, Anton Kostiuk, Ralf Korn |
CIFEr | 3 |
| 2014 | Exploiting expendable process-margins in DRAMs for run-time performance optimizationabstractManufacturing-time process (P) variations and runtime voltage (V) and temperature (T) variations can affect a DRAM's performance severely. To counter these effects, DRAM vendors provide substantial design-time PVT timing margins to guarantee correct DRAM functionality under worst-case operating conditions. Unfortunately, with technology scaling these timing margins have become large and very pessimistic for a majority of the manufactured DRAMs. While run-time variations are specific to operating conditions and as a result, their margins difficult to optimize, process variations are manufacturing-time effects and excessive process-margins can be reduced at run-time, on a per-device basis, if properly identified. In this paper, we propose a generic post-manufacturing performance characterization methodology for DRAMs that identifies this excess in process-margins for any given DRAM device at runtime, while retaining the requisite margins for voltage (noise) and temperature variations. By doing so, the methodology ascertains the actual impact of process-variations on the particular DRAM device and optimizes its access latencies (timings), thereby improving its overall performance. We evaluate this methodology on 48 DDR3 devices (from 12 DIMMs) and verify the derived timings under worst-case operating conditions, showing up to 33.3% and 25.9% reduction in DRAM read and write latencies, respectively. Karthik Chandrasekar 0001, Sven Goossens, Christian Weis, Martijn Koedam, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 6 |
| 2014 | Technology transfer towards Horizon 2020abstractEuropean research projects produce many excellent results, and the quality of research papers at DATE and other major European conferences is often outstanding. But how many academic research results in computing technologies and EDA actually make it into industrial practice? In the context of the transition into the Horizon 2020 framework program, the European research community is currently investigating novel ways of stimulating additional academia-industry technology transfer. This special session contributes by discussing concrete transfer experiences and new concepts. Furthermore it will exemplify several success stories from both academic and industrial perspectives. We believe that two major issues currently prevent a wider industrial adoption of research results at European scale: • While most FP7 research projects do provide ambitious exploitation plans, these are rarely implemented to a full extent, because the effort for productization is underestimated and insufficient resources and incentives are available when projects fade out. • There is a lot of emphasis on start-up companies as a primary vehicle for technology transfer. However, the effort of start-up foundation is very high and might not be justified in many cases due to limited market volume. Instead, more focus should be on industrial take-up of specific new technologies or IP generated by research, which does not require large amounts of venture capital. As a consequence, European research needs better mechanisms to provide incentives for technology transfer at small to medium scale. The speakers of this session are experienced actors in this domain. They will point out in a pragmatic way, and using concrete examples, how technology transfer can be initiated and implemented in practice and what are the associated pitfalls and innovation opportunities. The mix of presentations ensures that both academic and industrial viewpoints and concerns are properly addressed. Thus, the session will be of interest to a large audience. Amongst others, it is intended to stimulate more players to engage in international technology transfer. For this purpose, the session will be initiated by a brief presentation of a specific new pilot project (TETRACOM) focused on structured small to medium scale technology transfer. TETRACOM is open to the entire European computing research community and provides both funding and services for bilateral academia-industry collaborations. Rainer Leupers, Norbert Wehn, Marco Roodzant, Johannes Stahl 0004, Luca Fanucci, Albert Cohen 0001, Bernd Janson |
DATE | 2 |
| 2014 | Energy optimization in 3D MPSoCs with Wide-I/O DRAM using temperature variation aware bank-wise refreshabstractHeterogeneous 3D integrated systems with Wide-I/O DRAMs are a promising solution to squeeze more functionality and storage bits into an ever decreasing volume. Unfortunately, with 3D stacking, the challenges of high power densities and thermal dissipation are exacerbated. We improve DRAM refresh power by considering the lateral and vertical temperature variations in the 3D structure and adapting the per-DRAM-bank refresh period accordingly. In order to provide proof of our concepts we develop an advanced virtual platform which models the performance, power, and thermal behavior of a 3D-integrated MPSoC with Wide-I/O DRAMs in detail. On this platform we run the Android OS with real-world benchmarks to quantify the advantages of our ideas. We show improvements of 16% in DRAM refresh power due to temperature variation aware bank-wise refresh. Furthermore, two solutions are investigated to speedup system simulations: (1) Adaptive tuning of sampling intervals based on the estimated chip thermal profile, which results in speedups of 2X. (2) Hardware acceleration of thermal simulations using the Maxeler engine, which shows possible speedups of 12X. MohammadSadegh Sadri, Matthias Jung 0001, Christian Weis, Norbert Wehn, Luca Benini |
DATE | 4 |
| 2014 | Connecting different worlds - Technology abstraction for reliability-aware design and TestabstractThe rapid shrinking of device geometries in the nanometer regime requires new technology-aware design methodologies. These must be able to evaluate the resilience of the circuit throughout all System on Chip (SoC) abstraction levels. To successfully guide design decisions at the system level, reliability models, which abstract technology information, are required to identify those parts of the system where additional protection in the form of hardware or software coun-termeasures is most effective. Interfaces such as the presented Resilience Articulation Point (RAP) or the Reliability Interchange Information Format (RIIF) are required to enable EDA-assisted analysis and propagation of reliability information. The models are discussed from different perspectives, such as design and test. Ulf Schlichtmann, Veit Kleeberger, Jacob A. Abraham, Adrian Evans, Christina Gimmler-Dumont, Michael Glaß, Andreas Herkersdorf, Sani R. Nassif, Norbert Wehn |
DATE | 9 |
| 2014 | Hardware implementation of a Reed-Solomon soft decoder based on information set decodingabstractSoft decision decoding of Reed-Solomon codes can largely improve frame errors rates over currently used hard decision decoding. In this paper, we present a new hardware implementation for soft decoding of Reed-Solomon codes based on information set decoding. To our best knowledge this is the first hardware implementation of information set decoding for long Reed-Solomon codes. We propose a reduced complexity version of the decoding algorithm, that is optimized for efficient hardware implementation and enables high throughput. The decoder was implemented on a Virtex 7 FPGA, achieving a gain of 0.75 dB compared to conventional hard decision decoding and a throughput of up to 1.19 GBit/s for the widely used RS(255,239). This gain in FER is achieved with less complexity and more than 15x larger throughput than other state-of-the-art architectures. Stefan Scholl, Norbert Wehn |
DATE | 2 |
| 2014 | A new architecture for minimum mean square error sorted QR decomposition for MIMO wireless communication systemsabstractMultiantenna telecommunication systems represent channels with multiple inputs and multiple outputs (MIMO) by matrices. QR decomposition (QRD) of the channel matrix is a crucial part of MIMO detection algorithms, such as successive interference cancellation or sphere detection. Modern standards like Long Term Evolution (LTE) require the processing of millions of matrices per second, in order to compensate channel changes that occur due to the mobility of the detector and Doppler spread. We introduce a new architecture for minimum mean square error (MMSE) sorted QR decomposition based on Givens rotations. The architecture is derived from classical systolic array approach but includes modifications to allow sorting and MMSE preprocessing. It balances throughput against area and fulfills the real-time requirements of 1.763 μs and 0.881 μs derived from the LTE MIMO standard when synthesized on ALTERA Stratix III and Stratix V family FPGAs. Moreover, it can trade speed for area and is suitable for tighter time constraints. Victor Tomashevich, Christina Gimmler-Dumont, Christian Fesl, Norbert Wehn, Ilia Polian |
DDECS | 4 |
| 2014 | HyPER: A runtime reconfigurable architecture for monte carlo option pricing in the Heston modelabstractHigh-speed and energy-efficient computations are mandatory in the financial and insurance industry to survive in competition and meet the federal reporting requirements. On a hybrid CPU/FPGA system we propose a modular pricing engine and derive a novel algorithmic extension able to exploit online dynamic reconfiguration. The result is a high-performance and energy-efficient pricing system suitable for exotic option pricing in the state-of-the-art Heston market model. With the online reconfiguration extension our hybrid pricing system is nearly two orders of magnitude faster than high-end Intel CPUs, while consuming the same power. Christian Brugger, Christian de Schryver, Norbert Wehn |
FPL | 3 |
| 2014 | Monitoring household activities and user location with a cheap, unobtrusive thermal sensor arrayabstractWe demonstrate that a cheap (30USD) small, low power 8x8 thermal sensor array can by itself provide a broad range of information relevant for human activity monitoring in home and office environments. In particular the sensor can track people with an accuracy in the range of 1m (which is sufficient to recognize activity relevant regions), detect the operation mode of various appliances such as toaster, water cooker or egg cooker and actions such as opening a refrigerator, the oven or taking a shower. While there are sensing modalities for each of the above types of information (e.g. current sensors for appliances) the fact that they can all be detected by such a simple sensor is highly relevant for practical activity recognition systems. Compared to vision (or thermal imaging systems) the system has the advantage is being less privacy invasive allowing it for example to monitor bathroom activities (as shown in one of our evaluation scenarios). The paper describes the sensor, the methods used for activity detection and the evaluation. Peter Hevesi, Sebastian Wille, Gerald Pirkl, Norbert Wehn, Paul Lukowicz |
UbiComp | 4 |
| 2014 | A simplex algorithm for LP decoding hardwareabstractAn efficient LP decoder is the key building block for a maximum likelihood decoder based on integer programming. In this paper we propose to employ a variant of the simplex algorithm for LP decoding, called the dual simplex algorithm. This algorithm has two advantages: It inherently uses the received LLRs to generate a close to optimum starting solution and it allows to reuse former LP solutions if an adaptive LP decoding scheme is used. It is shown, that the dual simplex algorithm outperforms the standard (primal) simplex by a factor of 15-20 in runtime. This allows for efficient future hardware implementations. Furthermore the use of fixed-point instead of floating-point numbers is investigated to further reduce hardware complexity. Florian Gensheimer, Stefan Ruzika, Stefan Scholl, Norbert Wehn |
PIMRC | 4 |
| 2014 | Optimized active and power-down mode refresh control in 3D-DRAMsabstract3D stacked systems with Wide-I/O DRAMs are the future density optimized mobile computing platforms. Unfortunately, with 3D integration, the power densities and thermal dissipation are increased dramatically. In this paper, we investigate the effectiveness of power-down mode policies (using precharge power down, active power-down and self-refresh) and bank-wise refresh in active mode. We run real-life benchmarks to quantify the impact of each power-down mode setting. We derive a power-down mode policy which shows up to 10% energy reduction in high activity periods and up to 13% in idle phases. Further, we improve DRAM refresh power by considering the lateral and vertical temperature variations in the 3D structure and adapting the per-DRAM-bank refresh period accordingly. To achieve this, a per DRAM array hotspot detector, designed with DRAM cells and circuits, is used to acquire temperature and refresh information directly from the DRAM array. We show 16% improvements in DRAM refresh power due to hotspot detectors inside the DRAM enabling temperature variation aware bank-wise refresh. For all the above mentioned investigations a detailed DRAM controller model with accurate functionality, timing, and power estimation in SystemC TLM-2.0 (Transaction Level Modeling) and a highly sophisticated virtual hardware platform are mandatory to achieve a through analysis. Matthias Jung 0001, Christian Weis, Norbert Wehn, MohammadSadegh Sadri, Luca Benini |
VLSI-SoC | 3 |
| 2014 | A Cross-Layer Reliability Design Methodology for Efficient, Dependable Wireless ReceiversabstractContinued progressive downscaling of CMOS technologies threatens the reliability of chips for future embedded systems. We developed a novel design methodology for dependable wireless communication systems which exploits the mutual trade-offs of system performance, hardware reliability, and implementation complexity. Our cross-layer approach combines resilience techniques on hardware level with algorithmic techniques exploiting the available flexibility in the receiver. The overhead is minimized by recovering only from those hardware errors that have a strong impact on the system behavior. We apply our new methodology on a double-iterative MIMO-BICM receiver which belongs to the most complex systems in current communication standards. Christina Gimmler-Dumont, Norbert Wehn |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Code-aided synchronization with QPSK, 8-PSK and 16-QAM modulationsabstractThe wireless transmission results in phase and frequency offset which degrades the communication performance of a system significantly. Estimation of the unknown parameters of phase and frequency offset is called synchronization. The synchronization is typically performed once prior to the channel decoding by utilizing pilot symbols. Nowadays, turbo decoder can operate at low signal-to-noise ratio (SNR). In low SNR region, the parameters estimation suffers from increased noise. Therefore, the conventional synchronizers can not synchronize receivers perfectly. The synchronization quality can be improved by combining the parameter estimation step with the iterations of turbo decoder which is known as code-aided turbo synchronization. However, if the synchronization is performed after each iteration of turbo decoder then it increases its latency which is not affordable due to tight latency budget. Moreover, as typical wireless communication standards support multiple modulations soit is desirable to design a single synchronization algorithm flexible for them. Due to the tight latency constraints, efficient hardware implementations are mandatory. Hardware implementations require fixed-point representation resulting in quantization losses. We present in this paper an efficient fixed-point realization of a code-aided synchronization algorithm well suited for hardware implementation. The synchronization can run in parallel to turbo decoder to keep its latency unaffected. Moreover, it provides flexibility for following multiple modulations: QPSK, 8-PSK and 16-QAM. We show simulation results to demonstrate the communication performance of the algorithm. Uwe Wasenmüller, Norbert Wehn |
APCC | 3 |
| 2013 | Towards variation-aware system-level power estimation of DRAMs: an empirical approachabstractDRAM vendors provide pessimistic current measures in memory datasheets to account for worst-case impact of process variations and to improve their production yield, leading to unrealistic power consumption estimates. In this paper, we first demonstrate the possible effects of process variations on DRAM performance and power consumption by performing Monte-Carlo simulations on a detailed DRAM cross-section. We then propose a methodology to empirically determine the actual impact for any given DRAM memory by assessing its performance characteristics during the DRAM calibration phase at system boot-time, thereby enabling its optimal use at run-time. We further employ our analysis on Micron's 2Gb DDR3-1600-x16 memory and show considerable over-estimation in the datasheet measures and the energy estimates (up to 28%), by using realistic current measures for a set of MediaBench applications. Karthik Chandrasekar 0001, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DAC | 4 |
| 2013 | Reliable on-chip systems in the nano-era: lessons learnt and future trendsabstractReliability concerns due to technology scaling have been a major focus of researchers and designers for several technology nodes. Therefore, many new techniques for enhancing and optimizing reliability have emerged particularly within the last five to ten years. This perspective paper introduces the most prominent reliability concerns from today's points of view and roughly recapitulates the progress in the community so far. The focus of this paper is on perspective trends from the industrial as well as academic points of view that suggest a way for coping with reliability challenges in upcoming technology nodes. Jörg Henkel, Lars Bauer, Nikil Dutt, Puneet Gupta 0001, Sani R. Nassif, Muhammad Shafique 0001, Mehdi Baradaran Tahoori, Norbert Wehn |
DAC | 8 |
| 2013 | System and circuit level power modeling of energy-efficient 3D-stacked wide I/O DRAMsabstractJEDEC recently introduced its new standard for 3D-stacked Wide I/O DRAM memories, which defines their architecture, design, features and timing behavior. With improved performance/power trade-offs over previous generation DRAMs, Wide I/O DRAMs provide an extremely energy-efficient green memory solution required for next-generation embedded and high-performance computing systems. With both industry and academia pushing to evaluate and employ these highly anticipated memories, there is an urgent need for an accurate power model targeting Wide I/O DRAMs that enables their efficient integration and energy management in DRAM stacked SoC architectures. In this paper, we present the first system-level power model of 3D-stacked Wide I/O DRAM memories that is almost as accurate as detailed circuit-level power models of 3D-DRAMs. To verify its accuracy, we experimentally compare its power and energy estimates for different memory workloads and operations against those of a circuit-level 3D-DRAM power model and show less than 2% difference between the two sets of estimates. Karthik Chandrasekar 0001, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 4 |
| 2013 | A multi-level Monte Carlo FPGA accelerator for option pricing in the Heston modelabstractThe increasing demand for fast and accurate product pricing and risk computation together with high energy costs currently make finance and insurance institutes to rethink their IT infrastructure. Heterogeneous systems including specialized accelerator devices are a promising alternative to current CPU and GPU-clusters towards hardware accelerated computing. It has already been shown in previous work that complex state-of-the-art computations that have to be performed very frequently can be sped up by FPGA accelerators in a highly efficient way in this domain. A very common task is the pricing of credit derivatives, in particular options, under realistic market models. Monte Carlo methods are typically employed for complex or path dependent products. It has been shown that the multi-level Monte Carlo can provide a much better convergence behavior than standard single-level methods. In this work we present the first hardware architecture for pricing European barrier options in the Heston model based on the advanced multi-level Monte Carlo method. The presented architecture uses industry-standard AXI4-Stream flow control, is constructed in a modular way and can be extended to more products easily. We show that it computes around 100 millions of steps in a second with a total power consumption of 3.58 W on a Xilinx Virtex-6 FPGA. Christian de Schryver, Pedro Torruella, Norbert Wehn |
DATE | 3 |
| 2013 | Exploration and Optimization of 3-D Integrated DRAM SubsystemsabstractEnergy efficiency is the major optimization criterion for systems-on-chip (SoCs) for mobile devices (smartphones and tablets). Through silicon via (TSV) technology enables 3-D integration of dies and the heterogeneous stacking of multiple memory or logic layers, allowing increased bandwidth and lower energy consumption of the memory interface compared to traditional approaches. In this paper, we explore the 3-D-DRAM architecture design space. The result is an optimized 2 Gb 3-D-DRAM, which shows a 83% lower energy/bit than a 2 Gb device. Furthermore, we propose a highly energy-efficient DRAM subsystem for next-generation 3-D-integrated SoCs, consisting of a SDR/DDR 3-D-DRAM controller and an attached 3-D-DRAM cube with fine-grained access and a flexible (WIDE-IO) interface. We assess the energy efficiency using a synthesizable model of the SDR/DDR 3-D-DRAM channel controller (CC) as well as functional models of the 3-D-stacked DRAM, including an accurate power estimation engine. We also investigate different DRAM families (WIDE IO SDR/DDR, LPDDR, and LPDDR2) and densities from 256 Mb to 4 Gb per channel. The implementation results of the proposed 3-D-DRAM subsystem show that energy optimized accesses to the 3-D-DRAM enable up to 50% energy savings compared to standard accesses. To the best of our knowledge this is the first design space exploration for 3-D-stacked DRAM considering different technologies based on real-world physical data and the first design of a 3-D-DRAM CC and 3-D-DRAM model featuring co-optimization of memory and controller architecture. Christian Weis, Igor Loi, Luca Benini, Norbert Wehn |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | DRAM selection and configuration for real-time mobile systemsabstractThe performance and power consumption of mobile DRAMs (LPDDRs) depend on the configuration of system-level parameters, such as operating frequency, interface width, request size, and memory map. In mobile systems running both real-time and non-real-time applications, the memory configuration must satisfy bandwidth requirements of real-time applications, meet the power consumption budget, and offer the best average-case execution time to the non-real-time applications. There is currently no well-defined methodology for selecting a suitable memory configuration for real-time mobile systems. The worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across generations have furthermore not been investigated. This paper has two main contributions. 1) We analyze the worst-case bandwidth, average-case execution time, and power consumption of mobile DRAMs across three generations: LPDDR, LPDDR2 and Wide-IO-based 3D-stacked DRAM. 2) Based on our analysis, we propose a methodology for selecting memory configurations in real-time mobile systems.We show that LPDDR (32-bit IO), LPDDR2 (32-bit IO) and 3D-DRAM (128-bit IO) provide worst-case bandwidth up to 0.75 GB/s, 1.6 GB/s and 3.1 GB/s, respectively. We furthermore show for an H.263 decoder that LPDDR2 and 3D-DRAM reduce power consumption with up to 25% and 67%, respectively, compared to LPDDR, and reduce the execution time with up to 18% and 25%. Manil Dev Gomony, Christian Weis, Benny Akesson, Norbert Wehn, Kees Goossens |
DATE | 4 |
| 2012 | An energy efficient DRAM subsystem for 3D integrated SoCsabstractEnergy efficiency is the key driver for the design optimization of System-on-Chips for mobile terminals (smartphones and tablets). 3D integration of heterogeneous dies based on TSV (through silicon via) technology enables stacking of multiple memory or logic layers and has the advantage of higher bandwidth at lower energy consumption for the memory interface. In this work we propose a highly energy efficient DRAM subsystem for next-generation 3D integrated SoCs, which will consist of a SDR/DDR 3D-DRAM controller and an attached 3D-DRAM cube with a fine-grained access and a very flexible (WIDE-IO) interface. We implemented a synthesizable model of the SDR/DDR 3D-DRAM channel controller and a functional model of the 3D-stacked DRAM which embeds an accurate power estimation engine. We investigated different DRAM families (WIDE IO DDR/SDR, LPDDR and LPDDR2) and densities that range from 256Mb to 4Gb per channel. The implementation results of the proposed 3D-DRAM subsystem show that energy optimized accesses to the 3D-DRAM enable an overall average of 37% power savings as compared to standard accesses. To the best of our knowledge this is the first design of a 3D-DRAM channel controller and 3D-DRAM model featuring co-optimization of memory and controller architecture. Christian Weis, Igor Loi, Luca Benini, Norbert Wehn |
DATE | 4 |
| 2012 | A Parallel Adaptive Range Coding Compressor: Algorithm, FPGA Prototype, EvaluationabstractLoss less compression algorithms are employed in a wide variety of communication- and storage-related systems. Many embedded applications, such as real-time communication log compression used in automotive systems, impose strict throughput constraints on the compression unit, creating a demand for hardware-accelerated designs. In this paper we present a modification of the Adaptive Range Coding algorithm used by 7-Zip compressor implemented in an Field Programmable Gate Array (FPGA). We have improved the algorithm to support massive parallelization that allows making use of the distributed FPGA logic and achieving compression throughput of more than 50MB/s when implemented on a Virtex5 FPGA in conjunction with a hardware LZSS coder. Compared to a fixed-table Huffman encoder, our implementation provides the same high throughput and a 20% better compression ratio. Furthermore we explore several variations of algorithm parameters and show various trade-offs between compression efficiency, FPGA utilization and throughput. Ivan Shcherbakov, Norbert Wehn |
DCC | 2 |
| 2012 | Dependable embedded systems: The German research foundation DFG priority program SPP 1500abstractWhen migrating to future technology nodes, dependability becomes a major design problem as variability, aging and susceptibility to soft errors increase. The purpose of this program is to research cross-layer solutions that address the physical problems at system-level i.e. at hardware-level, operating system level, application level etc. The goals and an overview of the DFG SPP 1500 research program are presented. Jörg Henkel, Oliver Bringmann 0001, Andreas Herkersdorf, Wolfgang Rosenstiel, Norbert Wehn |
ETS | 5 |
| 2012 | Combining robotic frameworks with a smart environment framework: MCA2/SimVis3D and TinySEPabstractThis work describes the combination of three software frameworks from two different domains: robotics and smart environments. The two robotic frameworks MCA2 and SimVis-3D that have been in use for several years on a multitude of different robotic systems and TinySEP, a modular framework for smart environments were combined to create a win-win-situation for both roboticists and ubiquitous computing researchers. The possibilities and advantages this combination can offer are discussed, especially in situations where mobile robots and smart environments coexist next to each other. This work is concluded by an experiment that shows the feasibility and the strengths of the proposed approach. Michael Arndt, Karsten Berns, Sebastian Wille, Norbert Wehn, Luiza de Souza |
UbiComp | 4 |
| 2012 | Reliability study on system memories of an iterative MIMO-BICM system
Christina Gimmler-Dumont, Christian Brehm, Norbert Wehn |
VLSI-SoC | 3 |
| 2011 | Reliability: A Cross-Disciplinary and Cross-Layer ApproachabstractReliability is the next big challenge if CMOS scaling will continue. To solve this challenge, cross-disciplinary approaches become mandatory in which experts from different disciplines like hardware, software, OS and methodology have to cooperate deeply. All levels with appropriate abstractions have to be explored and the scope of the investigations has to be expanded to the application and quality-of-service level. Norbert Wehn |
Asian Test Symposium | 1 |
| 2011 | Design space exploration for 3D-stacked DRAMsabstract3D integration based on TSV (through silicon via) technology enables stacking of multiple memory layers and has the advantage of higher bandwidth at lower energy consumption for the memory interface. As in mobile applications energy efficiency is key, 3D integration is especially here a strategic technology. In this paper we focus on the design space exploration of 3D-stacked DRAMs with respect to performance, energy and area efficiency for densities from 256Mbit to 4Gbit per 3D-DRAM channel. We investigate four different technology nodes from 75nm down to 45nm and show the optimal design point for the currently most common commodity DRAM density of 1Gbit. Multiple channels can be combined for main memory sizes of up to 32GB. We present a functional SystemC model for the 3D-stacked DRAM which is coupled with a SDR/DDR 3D-DRAM channel controller. Parameters for this model were derived from detailed circuit level simulations. The exploration demonstrates that an optimized 1Gbit 3D-DRAM stack is 15× more energy efficient compared to a commodity Low-Power DDR SDRAM part without IO drivers and pads. To the best of our knowledge this is the first design space exploration for 3D-stacked DRAM considering different technologies and real world physical commodity DRAM data. Christian Weis, Norbert Wehn, Igor Loi, Luca Benini |
DATE | 2 |
| 2011 | Bringing C++ productivity to VHDL world: From language definition to a case study
Ivan Shcherbakov, Christian Weis, Norbert Wehn |
FDL | 3 |
| 2011 | Energy Efficient Acceleration and Evaluation of Financial Computations towards Real-Time Pricing
Christian de Schryver, Matthias Jung 0001, Norbert Wehn, Henning Marxen, Anton Kostiuk, Ralf Korn |
KES (4) | 3 |
| 2011 | On Complexity, Energy- and Implementation-Efficiency of Channel DecodersabstractFuture wireless communication systems require efficient and flexible baseband receivers. Meaningful efficiency metrics are key for design space exploration to quantify the algorithmic and the implementation complexity of a receiver. Most of the current established efficiency metrics are based on counting operations, thus neglecting important issues like data and storage complexity. In this paper we introduce suitable energy and area efficiency metrics which resolve the afore-mentioned disadvantages. These are decoded information bit per energy and throughput per area unit. Efficiency metrics are assessed by various implementations of turbo decoders, LDPC decoders and convolutional decoders. An exploration approach is presented, which permit an appropriate benchmarking of implementation efficiency, communications performance, and flexibility trade-offs. Two case studies demonstrate this approach and show that design space exploration should result in various efficiency evaluations rather than a single snapshot metric as done often in state-of-the-art approaches. Frank Kienle, Norbert Wehn, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 2010 | A 150Mbit/s 3GPP LTE Turbo code decoderabstract3 GPP long term evolution (LTE) enhances the wireless communication standards UMTS and HSDPA towards higher throughput. A throughput of 150 Mbit/s is specified for LTE using 2×2 MIMO. For this, highly punctured Turbo codes with rates up to 0.95 are used for channel coding, which is a big challenge for decoder design. This paper investigates efficient decoder architectures for highly punctured LTE Turbo codes. We present a 150 Mbit/s 3GPP LTE Turbo code decoder, which is part of an industrial SDR multi-standard baseband processor chip. Matthias May 0001, Thomas Ilnseher, Norbert Wehn, Wolfgang Raab |
DATE | 3 |
| 2010 | A rapid prototyping system for error-resilient multi-processor systems-on-chipabstractStatic and dynamic variations, which have negative impact on the reliability of microelectronic systems, increase with smaller CMOS technology. Thus, further downscaling is only profitable if the costs in terms of area, energy and delay for reliability keep within limits. Therefore, the traditional worst case design methodology will become infeasible. Future architectures have to be error resilient, i.e., the hardware architecture has to tolerate autonomously transient errors. In this paper, we present an FPGA based rapid prototyping system for multi-processor systems-on-chip composed of autonomous hardware units for error-resilient processing and interconnect. This platform allows the fast architectural exploration of various error protection techniques under different failure rates on the microarchitectural level while keeping track of the system behavior. We demonstrate its applicability on a concrete wireless communication system. Matthias May 0001, Norbert Wehn, Abdelmajid Bouajila, Johannes Zeppenfeld, Walter Stechele, Andreas Herkersdorf, Daniel Ziener, Jürgen Teich |
DATE | 2 |
| 2010 | Complete Verification of Weakly Programmable IPs against Their Operational ISA Model
Sacha Loitz, Markus Wedler, Dominik Stoffel, Christian Brehm, Norbert Wehn, Wolfgang Kunz |
FDL | 5 |
| 2010 | Fully integrated UWB impulse transmitter and 402-to-405MHz super-regenerative receiver for medical implant devicesabstractThis paper proposes a multi-standard transceiver architecture, based on UWB impulse transmitter and 402-to-405MHz super-regenerative receiver for medical implant devices. This architecture eliminates the requirement of frequency synthesizer and power amplifier at transmitter side and LNA, mixers, IF amplifiers and ADCs at receiver side of implant transceiver. The test structure of transceiver has been implemented on 0.18μm CMOS technology. The UWB impulse transmitter consumes 0.3mW from 1.5V supply at the data rate of 500Kbits/s. The 402-to-405MHz super-regenerative receiver consumes 0.5mW at the data rate of 120Kbits/s and sensitivity of -95dBm. Muhammad Anis, Maurits Ortmanns, Norbert Wehn |
ISCAS | 3 |
| 2010 | AmICA - A Flexible, Compact, Easy-to-Program and Low-Power WSN Platform
Sebastian Wille, Norbert Wehn, Ivan Martinovic, Simon Kunz, Peter Göhner |
MobiQuitous | 2 |
| 2010 | Low-complexity iteration control for MIMO-BICM systemsabstractAir bandwidth is a precious resource for wireless communication. Multiple-antenna (MIMO) systems enable an increase in channel capacity without increasing the air bandwidth. An iterative demapping and decoding at the receiver improves the communications performance remarkably. However, MIMO demapping and channel decoding have a high computational complexity. Energy consumption, latency and throughput of a hardware implementation strongly depend on the number of iterations. Iteration control techniques are very efficient to reduce the average number of iterations thus increasing decoder throughput and reducing energy consumption and average decoding latency. To the best of our knowledge, we present the first analysis of iteration control in MIMO bit interleaved coded modulation systems. We introduce a novel stopping metric for iteration control, which outperforms existing stopping metrics for middle and high signal-to-noise ratios. Additionally, we analyze state-of-the-art stopping metrics with respect to their algorithmic complexity in due consideration of a later hardware implementation. Christina Gimmler-Dumont, Timo Lehnigk-Emden, Norbert Wehn |
PIMRC | 3 |
| 2010 | A separation algorithm for improved LP-decoding of linear block codesabstractMaximum likelihood (ML) decoding is the optimal decoding algorithm for arbitrary linear block codes and can be written as an integer programming (IP) problem. Feldman relaxed this IP problem and presented linear programming (LP) based decoding. In this paper, we propose a new separation algorithm to improve the error-correcting performance of LP decoding for binary linear block codes. We use an IP formulation with indicator variables that help in detecting the violated parity checks. We derive Gomory cuts from the IP and use them in our separation algorithm. An efficient method of finding cuts induced by redundant parity checks (RPC) is also proposed. Under certain circumstances we can guarantee that these RPC cuts are valid and cut off the fractional optimal solutions of LP decoding. It is demonstrated on three LDPC codes and two BCH codes that our separation algorithm performs significantly better than LP decoding and belief propagation (BP) decoding. Akin Tanatmis, Stefan Ruzika, Horst W. Hamacher, Mayur Punekar, Frank Kienle, Norbert Wehn |
IEEE Trans. Inf. Theory | 6 |
| 2009 | Error correction in single-hop wireless sensor networks - A case studyabstractEnergy efficient communication is a key issue in wireless sensor networks. Common belief is that a multi-hop configuration is the only viable energy efficient technique. In this paper we show that the use of forward error correction techniques in combination with ARQ is a promising alternative. Exploiting the asymmetry between lightweight sensor nodes and a more powerful base station even advanced techniques known from cellular networks can be efficiently applied to sensor networks. Our investigations are based on realistic power models and real measurements and, thus, consider all side-effects. This is to the best of our knowledge the first investigation of advanced forward error correction techniques in sensor networks which is based on real experiments. Daniel Schmidt 0001, Matthias Berning, Norbert Wehn |
DATE | 3 |
| 2009 | A novel LDPC decoder for DVB-S2 IPabstractIn this paper a programmable Forward Error Correction (FEC) IP for a DVB-S2 receiver is presented. It is composed of a Low-Density Parity Check (LDPC), a Bose-Chaudhuri-Hoquenghem (BCH) decoder, and pre- and postprocessing units. Special emphasis is put on LDPC decoding, since it accounts for the most complexity of the IP core by far. We propose a highly efficient LDPC decoder which applies Gauss-Seidel decoding. In contrast to previous publications, we show in detail how to solve the well known problem of superpositions of permutation matrices. The enhanced convergence speed of Gauss-Seidel decoding is used to reduce area and power consumption. Furthermore, we propose a modified version of the lambda-Min algorithm which allows to further decrease the memory requirements of the decoder by compressing the extrinsic information. Compared to the latest published DVB-S2 LDPC decoders, we could reduce the clock frequency by 40% and the memory consumption by 16%, yielding large energy and area savings while offering the same throughput. Stefan Müller 0004, Manuel Schreger, Marten Kabutz, Matthias Alles, Frank Kienle, Norbert Wehn |
DATE | 6 |
| 2009 | Valid inequalities for binary linear codesabstractWe study an integer programming (IP) based separation approach to find the maximum likelihood (ML) codeword for binary linear codes. An algorithm introduced in Tanatmis et al. is extended and improved with respect to decoding performance without increasing the worst case complexity. This is demonstrated on the LDPC and the BCH code classes. Moreover, we propose an integer programming formulation to calculate the minimum distance of a binary linear code. We exemplarily compute the minimum distance of the (204, 102) LDPC code and the (576, 288) WIMAX code. Using the minimum distance of a code, a new class of valid inequalities is introduced. Stefan Ruzika, Akin Tanatmis, Frank Kienle, Horst W. Hamacher, Norbert Wehn, Mayur Punekar |
ISIT | 5 |
| 2008 | A Case Study in Reliability-Aware Design: A Resilient LDPC Code DecoderabstractChip reliability becomes a great threat to the design of future microelectronic systems with the continuation of the progressive downscaling of CMOS technologies. Hence increasing the robustness of chip implementations in terms of error tolerance becomes an important issue. In this paper we present a case study in reliability-aware design tolerating transient errors. A state-of-the-art WiMAX channel decoder for LDPC codes is investigated on all design levels to increase its reliability for a given system performance with minimum hardware overhead. We show that an efficient exploitation of the algorithmic fault-tolerance yields a fairly small area overhead with nearly no degradation in communications performance even under high error injection rates. Matthias May 0001, Matthias Alles, Norbert Wehn |
DATE | 3 |
| 2008 | A Reconfigurable Application Specific Instruction Set Processor for Convolutional and Turbo Decoding in a SDR EnvironmentabstractFuture mobile and wireless communication networks require flexible modem architectures to support seamless services between different network standards. Hence, a common hardware platform that can support multiple protocols implemented or controlled by software, generally referred to as software defined radio (SDR), is essential. This paper presents a family of dynamically reconfigurable application-specific instruction-set processors (ASIP) for the application domain of channel coding in wireless communication systems. As a weakly programmable IP core, it can implement trellis based channel decoding in a SDR environment. It features binary convolutional decoding, and turbo decoding for binary as well as duobinary turbo codes for all current and upcoming standards. The ASIPs consist of a specialized pipeline with 15 stages and a dedicated communication and memory infrastructure. Logic synthesis revealed a maximum clock frequency of 400 MHz and a total area of 0.42 mm2for a 65 nm technology. Simulation results for Viterbi and turbo decoding demonstrate maximum throughput of 196 and 34 Mbps, respectively, and outperforms existing SDR based approaches for channel decoding. Timo Vogt, Norbert Wehn |
DATE | 2 |
| 2008 | Application-specific reconfigurable processorsabstractApplication-specific reconfigurable processor architectures provide a remarkable potential for systems which achieve concurrently high performance, area efficiency, energy efficiency, run-time adaptivity, and sufficient flexibility. Thus, they represent competitive design alternatives that provide significant improvements in some of these figures of merit in comparison to non-reconfigurable architectures. Research results of three projects on the analysis of architecture concepts, design, and evaluation of application-specific reconfigurable processors in different domains are presented. Heiko Hinkelmann, Peter Zipf, Manfred Glesner, Matthias Alles, Timo Vogt, Norbert Wehn, Götz Kappen, Tobias G. Noll |
FPL | 6 |
| 2008 | 3.1-to-7GHz UWB impulse radio transceiver front-end based on statistical correlation techniqueabstractThis paper presents the UWB impulse radio transceiver front-end based on statistical correlation technique between multiple band-pass filters tuned within UWB spectrum. The band-pass filters can detect energy either from transmitted UWB impulses, interfering channels or noise. Transmitted UWB impulses can be extracted by statistical correlation in between the outputs of band-pass filters. Super-regenerative receivers, tuned within 3.1-to-7 GHz UWB spectrum, act as band-pass filters. Summer and comparator perform statistical correlation. The test structure of transceiver has been implemented on 0.18 mum CMOS technology, active area of 1.44 mm2, sensitivity of -99 dBm and -30 dB SIR with power consumption of 15 mW at 1.5 V. Muhammad Anis, Reinhard Tielert, Norbert Wehn |
ISCAS | 3 |
| 2008 | Macro Interleaver Design for Bit Interleaved Coded Modulation with Low-Density Parity-Check CodesabstractBit interleaved coded modulation (BICM) is a pragmatic approach to achieve power and bandwidth efficient transmission. At this low-density parity-check codes (LDPCC) can be utilized which are then concatenated with a modulator. The concatenation of LDPCC and modulator is done via a bit interleaver which has the task to spread information. While the matched design of an optimal LDPC degree distribution with respect to a given mapping scheme is well observed it is often assumed that the probabilistic values passed from demodulator to the LDPCC decoder are independent due to the mapping interleaver. However for higher order modulations, e.g. 16-QAM or 256-QAM, more and more bits are getting correlated during the demodulation process. The connectivity of these correlated bits have to be considered during the design of the mapping interleaver with respect to the graph of the concatenated LDPC codes. In this paper, we present the design of mapping interleavers of BICM schemes utilizing LDPC codes. Simulation results show the strong influence of correlated inputs to the performance of LDPC decodes. This correlations can cause a high error floor if the mapping interleaver is not carefully designed. Frank Kienle, Norbert Wehn |
VTC Spring | 2 |
| 2008 | Designing efficient irregular networks for heterogeneous systems-on-chip
Christian Neeb, Norbert Wehn |
J. Syst. Archit. | 2 |
| 2008 | A Reconfigurable ASIP for Convolutional and Turbo Decoding in an SDR EnvironmentabstractFuture mobile and wireless communication networks require flexible modem architectures to support seamless services between different network standards. Hence, a common hardware platform that can support multiple protocols implemented or controlled by software, generally referred to as software defined radio (SDR), is essential. This paper presents a family of dynamically reconfigurable application-specific instruction-set processors (ASIPs) for channel coding in wireless communication systems. As a weakly programmable intellectual property (IP) core, it can implement trellis-based channel decoding in a SDR environment. It features binary convolutional decoding, and turbo decoding for binary as well as duobinary turbo codes for all current and upcoming standards. The ASIP consists of a specialized pipeline with 15 stages and a dedicated communication and memory infrastructure. Logic synthesis revealed a maximum clock frequency of 400 MHz and an area of 0.11 mm2for the processor's logic using a low power 65-nm technology. Memories require another 0.31 mm2. Simulation results for Viterbi and turbo decoding demonstrate maximum throughput of 196 and 34 Mb/s, respectively. The ASIP hence outperforms state-of-the-art decoder architectures targeting software defined radio by at least a factor of three while consuming only 60% or less of the logic area. Timo Vogt, Norbert Wehn |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Low complexity LDPC code decoders for next generation standardsabstractThis paper presents the design of low complexity LDPC codes decoders for the upcoming WiFi (IEEE 802.11n), WiMax (IEEE802.16e) and DVB-S2 standards. A complete exploration of the design space spanning from the decoding schedules, the node processing approximations up to the top-level decoder architecture is detailed. According to this search state-of-the-art techniques for a low complexity design have been adopted in order to meet feasible high throughput decoder implementations. An analysis of the standardized codes from the decoder-aware point of view is also given, presenting, for each one, the implementation challenges (multi rates-length codes) and bottlenecks related to the complete coverage of the standards. Synthesis results on a present 65nm CMOS technology are provided on a generic decoder architecture Torben Brack, Matthias Alles, Timo Lehnigk-Emden, Frank Kienle, Norbert Wehn, Nicola E. L'Insalata, Massimo Rovini, Luca Fanucci |
DATE | 5 |
| 2007 | Evaluation of High Throughput Turbo-Decoder ArchitecturesabstractThe outstanding forward error correction of Turbo-codes made them part of many today's communications standards. For high throughput applications, efficient parallel Turbo-decoder architectures are the key. In this paper, two fundamentally different parallel architectural approaches in terms of performance and implementation complexity were compared. Both architectures exploit the well known windowing scheme. The first architecture template processes several windows in parallel. Each window is executed on a serial Log-MAP decoder which produces one value per clock cycle. In contrast, the second architecture sequentially processes the individual windows on a fast monolithic pipelined MAP decoder which produces several values per clock cycle and is memory-optimized. This is, to the best of the author's knowledge, the first comparison of this totally different architectural approach for high throughput Turbo-decoder architectures. The 3GPP conditions for performance comparisons were applied. Matthias May 0001, Christian Neeb, Norbert Wehn |
ISCAS | 3 |
| 2007 | Implementation Issues of Turbo Synchronization with Duo-Binary Turbo DecodingabstractThe transmission over a wireless channel results in timing, frequency and phase offsets. To circumvent the severe losses of communications performance caused by these offsets a sophisticated synchronization is mandatory. Synchronization is typically performed only once prior to the channel decoding. In this paper the authors present an FPGA implementation of a joint iterative decoder and synchronizer, which is also referred to as turbo synchronizer. We investigate the additional costs of turbo synchronization in terms of implementation complexity with a 16-state duo-binary turbo decoder. Furthermore we present the communications performance of the turbo synchronizer taking the implementation losses into account. Matthias Alles, Timo Lehnigk-Emden, Uwe Wasenmüller, Norbert Wehn |
PIMRC | 4 |
| 2007 | A Reliability-Aware LDPC Code Decoding AlgorithmabstractWith the continuing downscaling of microelectronic technology, chip reliability becomes a great threat to the design of future complex microelectronic systems. Hence increasing the robustness of chip implementations in terms of tolerating errors becomes mandatory. In this paper we present reliability-aware extensions of the LDPC decoding algorithm. We exploit application specific fault tolerance of the decoding algorithm combined with modifications on the algorithmic level to increase the reliability of a decoder implementation. These modifications lead to a LDPC decoder implementation which tolerates sporadic errors that occur in critical components. To the best of our knowledge this is the first investigation of the LDPC decoding algorithm in terms of implementation reliability. Matthias Alles, Torben Brack, Norbert Wehn |
VTC Spring | 3 |
| 2007 | A Survey on LDPC Codes and Decoders for OFDM-based UWB SystemsabstractCurrent UWB systems apply convolutional codes as their channel coding scheme. For next generation systems LDPC codes are in discussion due to their outstanding communications performance. LDPC codes are already utilized in the new WiMax and WiFi standards. Thus it is reasonable to investigate these codes as candidate LDPC codes for UWB. In this paper the authors present an implementation complexity and performance comparison of LDPC decoders. We will show that it is of great advantage to design new LDPC codes which are tailored to the special latency and throughput constraints of upcoming UWB systems. This new class of LDPC codes is named ultra-sparse LDPC codes. Synthesis results of WiMax, WiFi, and U-S LDPC decoders are presented based on an enhanced 65 nm CMOS process. We show that the implementation complexity of the new U-S LDPC decoders is 55% smaller, utilizing only 0.2 mm2instead of over 0.4 mm2, while the communications performance of all observed LDPC codes are almost identical under all the considered UWB simulation conditions. Torben Brack, Matthias Alles, Timo Lehnigk-Emden, Frank Kienle, Norbert Wehn, Friedbert Berens, Andreas Ruegg |
VTC Spring | 5 |
| 2006 | Disclosing the LDPC code decoder design spaceabstractThe design of future communication systems with high throughput demands will become a critical task, especially when sophisticated channel coding schemes have to be applied. LDPC codes are one of the most promising candidates because of their outstanding communications performance. One major problem for a decoder hardware realization is the huge design space composed of many interrelated parameters which enforces drastic design trade-offs. Another important issue is the need for flexibility of such systems. In this paper we illuminate this design space with special emphasis on the strong interrelations of theses parameters. Three design studies are presented to highlight the effects on a generic architecture if some parameters are constraint by a given standard, given technology, and given area constraints Torben Brack, Frank Kienle, Norbert Wehn |
DATE | 3 |
| 2006 | Designing Efficient Irregular Networks for Heterogeneous Systems-on-ChipabstractNetworks-on-Chip will serve as the central integration platform in future complex SoC designs, composed of a large number of heterogeneous processing resources. Most researchers advocate the use of traditional regular networks like meshes, tori or trees as architectural templates which gained a high popularity in general-purpose parallel computing. However, most SoC platforms are special-purpose tailored to the domain-specific requirements of their application. They are usually built from a large diversity of heterogeneous components which communicate in a very specific, mostly irregular way. In this work, we propose a methodology for the design of customized irregular networks-on-chip, called INoC. We take advantage of a priori knowledge of the applications communication characteristic to generate an optimized network topology and routing algorithm. We show that customized irregular networks are clearly superior to traditional regular architectures in terms of performance at comparable implementation costs for irregular workloads. Even more, they inherently offer true scalability and expansibility which can normally not be accomplished by traditional approaches Christian Neeb, Norbert Wehn |
DSD | 2 |
| 2006 | A Synthesizable IP Core for WIMAX 802.16E LDPC Code DecodingabstractThe upcoming IEEE WiMax 802.16e standard, also referred to as WirelessMAN (2005), is the next step toward very high throughput wireless backbone architectures, supporting up to 500 Mbps. It features as an advanced channel coding scheme low-density parity-check codes. The decoding of LDPC codes is an iterative process, hence many data have to be exchanged between processing units within each iteration. The variety of the specified codes and the envision of different decoding schedules for different codes pose significant challenges to an LDPC decoder hardware realization. In this paper, we present to the best of our knowledge the first published LDPC decoder architecture capable to process all specified WiMax LDPC codes. Detailed synthesis and communications performance results are shown in addition Torben Brack, Matthias Alles, Frank Kienle, Norbert Wehn |
PIMRC | 4 |
| 2006 | Fast convergence algorithm for LDPC CodesabstractLow-density parity-check (LDPC) codes are one of the most powerful codes known today. They are decoded iteratively by a message passing algorithm. There exist many different update schemes of the exchanged messages. The major difference of all update schemes is the convergence speed, i.e. the achieved communications performance for a limited number of iterations. This paper presents a new decoding algorithm which efficiently utilizes the encoder property of linear encodable LDPC codes. The basic idea is to interpret the LDPC encoder as an encoder with puncturing unit which opens as well the door for hybrid ARQ schemes. The presented new decoding algorithm shows a faster convergence behavior than state of art decoding schemes and it results in a lower error floor Frank Kienle, Timo Lehnigk-Emden, Norbert Wehn |
VTC Spring | 3 |
| 2005 | A Synthesizable IP Core for DVB-S2 LDPC Code DecodingabstractThe new standard for digital video broadcast DVB-S2 features low-density parity-check (LDPC) codes as their channel coding scheme. The codes are defined for various code rates with a block size of 64800 which allows a transmission close to the theoretical limits. The decoding of LDPC is an iterative process. For DVB-S2 about 300000 messages are processed and reordered in each of the 30 iterations. These huge data processing and storage requirements are a real challenge for the decoder hardware realization, which has to fulfill the specified throughput of 255 Mbit/s for base station applications. In this paper we show, to the best of our knowledge, the first published IP LDPC decoder core for the DVB-S2 standard. We present a synthesizable IP block based on ST Microelectronics 0.13 /spl mu/m CMOS technology. Frank Kienle, Torben Brack, Norbert Wehn |
DATE | 3 |
| 2004 | Design methodology for IRA codes
Frank Kienle, Norbert Wehn |
ASP-DAC | 2 |
| 2004 | Channel Decoder Architecture for 3G Mobile Wireless TerminalsabstractChannel coding is a key element of any digital wireless communication system since it minimizes the effects of noise and interference on the transmitted signal. In third-generation (3G) wireless systems channel coding techniques must serve both voice and data users whose requirements considerably vary. Thus the third generation partnership project (3GPP) standard offers two coding techniques, convolutional-coding for voice and turbo-coding for data services. In this paper we present a combined channel decoding architecture for 3G terminal applications. It outperforms a solution based on two separate decoders due to an efficient reuse of computational hardware and memory resources for both decoders. Moreover it supports blind transport format detection. Special emphasis is put on low energy consumption. Friedbert Berens, Gerd Kreiselmaier, Norbert Wehn |
DATE | 3 |
| 2004 | Joint graph-decoder design of IRA codes on scalable architectures [LDPC codes]abstractChannel coding is an important building block in communication systems since it ensures the quality of service. Irregular repeat-accumulate (IRA) codes belong to the class of low-density parity-check (LDPC) codes and even outperform the recently introduced turbo-codes of current communication standards. The advantage of IRA codes over LDPC codes is that they come with a linear-time encoding complexity. IRA codes can be represented by a Tanner graph with arbitrary connections between nodes of given degrees. The implementation complexity of IRA decoders is dominated by the randomness of these connections. In this paper, we present a scalable partly parallel IRA decoder architecture. We present a joint graph-decoder design to parallelize IRA codes which can be efficiently processed by this decoder without any RAM access conflicts. We show design examples of these IRA codes which outperform the UMTS turbo-code by 0.2 dB. Frank Kienle, Norbert Wehn |
ICASSP (4) | 2 |
| 2003 | Communication Centric Architectures for Turbo-Decoding on Embedded Multiprocessors
Frank Gilbert, Michael J. Thul, Norbert Wehn |
DATE | 3 |
| 2003 | VLSI-implementation issues of turbo trellis-coded modulationabstractTurbo trellis-coded modulation (TTCM) is a very promising approach for future communication systems. It combines the advantages of channel coding with multilevel signals and the powerful turbo-codes concept. In this paper we consider VLSI implementation aspects of TTCM. We show that techniques, known from binary turbo-decoders, can be applied to TTCM to reduce the implementation complexity significantly. In detail we explore iteration control, quantization, scaling, and the MAP architecture using a bit-true model of an 8-state TTCM decoder. Frank Kienle, Gerd Kreiselmaier, Norbert Wehn |
ICASSP (2) | 3 |
| 2003 | Concurrent interleaving architectures for high-throughput channel codingabstractInterleavers are widely used for a vast range of communications applications. Traditionally used for burst-error separation in distorted channels, they have gained additional interest since the discovery of turbo codes whose performance essentially depends on the interleavers. With the ever increasing data rates demanded by customers, architectures that provide interleaving at high throughput become mandatory. We present an heuristic approach to the design of interleaving architectures based on random graph generation. They can handle any given interleaver pattern and allow for any parallelization degree, and hence speed-up, of the interleaving operation. Moreover, this enables highly parallel architectures for channel decoders such as turbo- and LDPC-decoders. Michael J. Thul, Frank Gilbert, Norbert Wehn |
ICASSP (2) | 3 |
| 2002 | Hardware/Software Trade-Offs for Advanced 3G Channel CodingabstractThird generation's wireless communications systems comprise advanced signal processing algorithms that increase the computational requirements more than ten-fold over 2G's systems. Numerous existing and emerging standards require flexible implementations ("software radio"). Thus efficient implementations of the performance-critical parts as Turbo decoding on programmable architectures are of great interest. Besides high-performance DSPs, application-customized RISC cores offer the required performance while still maintaining the aspired flexibility. This paper presents for the first time Turbo decoder implementations on customized RISC cores and compares the results with implementations on state-of-the-art VLIW DSPs. The results of our studies show that the Log-MAP performance is about 50% higher that on an ST120, a current VLIW architecture. Heiko Michel, Alexander Worm, Norbert Wehn, Michael Münch |
DATE | 3 |
| 2002 | Evaluation of algorithm optimizations for low-power Turbo-Decoder implementationsabstractEnergy aware and low-power implementation of Turbo-Decoders are a must for 3GPP and other communication system designs. This paper explores a reduced-search maximum-a-posteriori algorithm usually referred to as a low-power strategy for implementation. Synthesis results indeed show a noticeable power saving potential in the absence of iteration control. In the presence of iteration control, however, this paper shows that it even leads to an increased power consumption. Michael J. Thul, Timo Vogt, Frank Gilbert, Norbert Wehn |
ICASSP | 4 |
| 2001 | Low power implementation of a turbo-decoder on programmable architecturesabstractLow Power is an extremely important issue for future mobile radio systems. Channel decoders are essential building blocks of base-band signal processing units in mobile terminal architectures. Thus low power implementations of advanced channel decoding techniques are mandatory. In this paper we present a low power implementation of the most sophisticated channel decoding algorithm (Turbo-decoding) on programmable architectures. Low power optimization is performed on two abstraction levels: on system level by the use of an intelligent cancellation technique, on implementation level by the use of dynamic voltage scaling. With these techniques we can reduce the worst case is also applicable for hardware implementations. To the best of our knowledge, this is the first in-depth study of low power implementations of Turbo-decoders based on voltage scheduling for third generation wireless systems. Frank Gilbert, Alexander Worm, Norbert Wehn |
ASP-DAC | 3 |
| 2001 | Design of low-power high-speed maximum a priori decoder architecturesabstractFuture applications demand high-speed maximum a posteriori (MAP) decoders. In this paper, we present an in-depth study of design alternatives for high-speed MAP architectures with special emphasis on low power consumption. We exploit the inherent parallelism of the MAP algorithm to reduce power consumption on various abstraction levels. A fully parameterizable architecture is introduced which allows us to optimally adapt the architecture to the application requirements and the throughput. Intensive design space exploration has been carried out on a state-of-the-art 0.2 /spl mu/m technology, including efficient parallelism techniques, a data flow transformation for reduced power consumption, and an optimized FIFO implementation. Alexander Worm, Holger Lamm, Norbert Wehn |
DATE | 3 |
| 2000 | Automating RT-Level Operand Isolation to Minimize Power Consumption in DatapathsabstractDesigns which do not fully utilize their arithmetic datapath components typically exhibit a significant overhead in power consumption. Whenever a module performs an operation whose result is not used in the downstream circuit, power is being consumed for an otherwise redundant computation. Operand isolation is a technique to minimize the power overhead incurred by redundant operations by selectively blocking the propagation of switching activity through the circuit. This paper discusses how redundant operations can be identified concurrently to normal circuit operation, and presents a model to estimate the power savings that can be obtained by isolation of selected modules at the register-transfer (RT) level. Based on this model, an algorithm is presented to iteratively isolate modules, while minimizing the cost incurred by RTL operand isolation. Experimental results with power reductions of up to 30% demonstrate the effectiveness of the approach. Michael Münch, Norbert Wehn, Bernd Wurth, Renu Mehra, Jim Sproch |
DATE | 2 |
| 1998 | Embedded DRAM Architectural Trade-OffsabstractIn this paper we discuss system-related aspects in embedded DRAM/logic designs. We focus on large embedded memories which have to be implemented as DRAMs. Norbert Wehn, Søren Hein |
DATE | 1 |
| 1997 | An efficient ILP-based scheduling algorithm for control-dominated VHDL descriptionsabstractTo adopt behavioral synthesis techniques in existing design flows, the synthesis methodology must provide the designer with a mechanism to specify a component's interface timing. This will permit pre- and postsynthesis validation through cosimulation with other subsystems or even through formal verification. In control-flow dominated designs, additional timing constraints will result in a complex specification/constraint system for which the scheduling problem has been shown to be NP-complete. In this article, we present a mathematical framework for solving a special instance of the scheduling problem in control-flow dominated behavioral VHDL descriptions given that the timing of I/O signals has been completely or partially specified. It is based on a code-transformation approach that fully preserves the VHDL semantics. The scheduling problem is mapped onto an integer linear program (ILP) solvable in polynomial time assuming a restricted partial order on selected statements. It captures both control-flow and timing constraints in a single model and also exploits dataflow information to optimize the statement sequence across basic block boundaries. Michael Münch, Norbert Wehn, Manfred Glesner |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 1994 | The Hyeti Defect Tolerant Microprocessor: A Practical Experiment and its Cost-Effectiveness AnalysisabstractThis paper summarizes a practical experiment in designing a defect tolerant microprocessor and presents the underlying principles. Unlike memory integrated circuits, microprocessors have an irregular structure which complicates both the task of incorporating redundancy for defect tolerance in the design and the task of analyzing the resulting yield increase. The main goal of this paper is to present the detailed yield analysis of a defect tolerant microprocessor with an irregular structure which has been successfully fabricated. The approaches employed for achieving the goal of yield enhancement in the data path and the control part of the microprocessor are described first. Then, the yield enhancement due to the incorporated redundancy is analyzed. Finally, some practical and theoretical conclusions are drawn.> Régis Leveugle, Zahava Koren, Israel Koren, Gabriele Saucier, Norbert Wehn |
IEEE Trans. Computers | 5 |
| 1993 | The Siemens high-level synthesis system CALLASabstractIn this paper we present the Siemens high-level synthesis system CALLAS and describe its design methodology and synthesis strategy. It supports the synthesis of control-dominated applications and uses a VHDL subset for the algorithmic specification. Its main feature can be characterized as "What you simulate is what you synthesize." This principle permits a validation of the synthesis results by simulation or even formal verification. CALLAS has been successfully applied on real designs which were implemented in silicon. These examples demonstrate that CALLAS fulfils the constraints and objectives of a hardware designer. The circuits are comparable in quality to results achieved by synthesis starting at the register-transfer-level.> Jörg Biesenack, Michael Koster, Anton Langmaier, Stephane Ledeux, Sabine März, Michael Payer, Michael Pilsl, Steffen Rumler, Holger Soukup, Norbert Wehn, Peter Duzy |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 1991 | HADES-high-level architecture development and exploration systemabstractThe authors propose a new approach to high level behavioural synthesis starting from an algorithmic description in Hardware C. The algorithm is compiled into a corresponding data/control flow graph including several optimizations. The behavioural synthesis part of the system performs transformations like loop unrolling, parallelization, etc., whereby the user is supported through a feedback loop. For final structural synthesis an advanced tool based on genetic algorithms is provided. Layout synthesis is assumed to be performed by available tools like the GENESIL silicon compiler.> Peter Pöchmüller, Michael Held, Norbert Wehn, Manfred Glesner |
Great Lakes Symposium on VLSI | 3 |
| 1991 | A new approach to timing driven partitioning of combinational logicabstractThe authors present a new approach to timing driven partitioning of combinational logic. Instead of accessing a predefined library, complex gates based on the line-of-diffusion layout style are automatically synthesized. A new timing model for complex gates is presented which permits a fast pattern independent timing analysis with a deviation of less than 10% and two to three orders of magnitude faster than the exact SPICE simulation taking into account all parasitics and signal slopes. To improve the overall timing, a heuristic is presented which is based on iterative partitioning techniques for complex gates. The overall performance is demonstrated on several examples.> Norbert Wehn, Manfred Glesner |
Great Lakes Symposium on VLSI | 1 |
| 1991 | RAMSES-a rapid prototyping environment for embedded control applicationsabstractAn environment is given for fast prototyping of real-time systems. The supported spectrum of realizations for such systems varies from single task implementations on a general purpose microprocessor to heterogeneous multiprocessor systems with application specific processors. The prototyping environment includes a language to describe a real-time system, a compiler that transforms the description into its specified realization format and a hardware environment that is based on a VMEbus system. The VMEbus system includes a Motorola 68020 CPU with multi-tasking operating system and a prototyping board which can be either configured as multi-DSP board or as emulation board for the implementation of an application specific processor architecture.> Hans-Jürgen Herpel, Norbert Wehn, Manfred Glesner |
RSP | 2 |
| 1988 | A Defect-Tolerant and Fully Testable PLA
Norbert Wehn, Manfred Glesner, K. Caesar, P. Mann, A. Roth |
DAC | 1 |
| 1987 | The ALGIC Silicon Compiler System: Implementation, Design Experience and ResultsabstractIn this paper we present the ALGIC silicon compiler system. Starting from a structural and functional description of digital VLSI circuits, the system generates automatically the corresponding layout in a full custom design style. Main components of the ALGIC system are a system monitor module, a parameterisable macrocell generator, an appropriate floorplanner with 100% routing solution including planar VDD/GND trees and a block-oriented timing verifier. One essential feature of the system is the high degree of integration between all program modules using the concept of abstract data types. The flexibility and performance of the whole system is demonstrated by real design examples in the field of DSP applications. Johannes Schuck, Norbert Wehn, Manfred Glesner, G. Kamp |
DAC | 2 |