Milos Krstic

dblp:33/1954 · also Krstic Milos, Milos D. Krstic · DBLP profile ↗
← Back
95ranked-venue papers
4as first author
49since 2021 · last 2026
0000-0003-0267-0203ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 76 · 3 first-author · 36 since 2021Software engineering, systems software and programming languages · 17 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Building AI Hardware Expertise: Edge AI Curriculum Design and Implementation in German Universities
abstract
Edge artificial intelligence (AI) redistributes AI computation from distant cloud to local processors for real-time processing and enhanced privacy. This fundamental shift underscores a critical gap in current university curricula, which predominantly focus on AI fundamentals and algorithms while often neglecting essential AI hardware topics. To address this deficiency, this paper presents Edge AI, a postgraduate curriculum co-designed by two universities in Germany. Guided by the Dagstuhl triangle, the curriculum is designed to comprehensively cover technical, sociocultural, and application perspectives. Courses are developed using the Four-Component Instructional Design model to encourage action-oriented skill development, with a Learning Management System template available to assist in the design of individual courses. Selected practical courses are formulated as self-managed projects and inverted classrooms, enabling students to learn at their own pace with just-in-time guidance. All curriculum materials are accessible online and maintained by the Open Science Framework to enhance collaboration across institutions and promote applicability in diverse domains. Evaluation results from 176 students over two years (2023-2025) demonstrate universal satisfaction across various curriculum components.
Ann-Marie Gursch, Xuanshu Luo, Lilian Hasse, Carsten Trinitis, Ulrike Lucke, Martin Werner 0001, Milos Krstic
AAAI7
2026 Special Session: Optimizing Edge AI - Current Challenges and the Neuromorphic Outlook
abstract
The increasing deployment of AI (artificial intelligence) on edge devices presents major challenges due to strict constraints on computation, memory, energy, and latency. Effective Edge AI systems thus require multi-objective optimization that balances accuracy, hardware efficiency, and reliability. The Horizon Twinning project AIDA4Edge tackles these challenges by developing methods for efficient and reliable AI on resource-constrained platforms. This paper presents key approaches explored within the project, including neural network quantization, hardware-aware neural architecture search, dynamic neural networks, and self-adaptive resilient AI architectures. Finally, these strategies are placed within a broader, biologically inspired paradigm, highlighting neuromorphic computing as a natural continuation of Edge AI efforts toward highly efficient and resilient intelligent systems.
Milan R. Dincic, Zoran H. Peric, Davide Bertozzi, Alice Bizzarri, Rizwan Tariq Syed, Edward G. Jones, Riccardo Zese, Marko S. Andjelkovic, Fabian Vargas 0001, Milos Krstic, Oliver Rhodes, Modhe Almelihi, Tamara Milovanovic, Ivan Popovic, Sofija Peric
DDECS10
2026 Machine Learning Approach for Cross-Technology Prediction of the Generated Single Event Transient
Konstantinos Varakliotis, Nikolaos Zazatis, Marko S. Andjelkovic, Nikolaos Chatzivangelis, Fabian Vargas 0001, Milos Krstic, Christos P. Sotiriou
ETS6
2026 ReFFT: An Energy-Efficient RRAM-Based FFT Accelerator
abstract
The fast Fourier transform (FFT) is a highly efficient algorithm for computing the discrete Fourier transform (DFT). It is widely employed in various applications, including digital communication, image processing, and signal analysis. Recently, in-memory computing architectures based on emerging technologies, such as resistive RAM (RRAM), have demonstrated promising performance with low hardware cost for data-intensive applications. However, directly mapping FFT onto RRAM crossbars is challenging because the algorithm relies on many small, sequential butterfly operations, while cross-bars are optimized for large-scale, highly parallel vector–matrix multiplications (VMMs). In this paper, we introduce ReFFT, a system architecture that reformulates FFT computations for efficient execution on RRAM crossbars. ReFFT combines the reduced computational complexity of FFT with the parallel VMM capability of RRAM. We incorporate measured device data into our framework to analyze the effect of variability and develop an adaptive mapping scheme that improves twiddle-factor programming accuracy, leading to a 9.9 dB peak signal-to-noise ratio (PSNR) improvement for a 256-point FFT. Compared with prior RRAM-based DFT designs, ReFFT achieves up to 4.6× and 19.5× higher energy efficiency for 256- and 2048-point FFTs, respectively. The system is further validated in digital communication and satellite image compression tasks.
Jianan Wen, Andrea Baroni, Max Uhlmann, Christian Wenger, Milos Krstic
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 RRAM-Based Spectral-Domain Convolution Accelerator for Reliable and Energy-Efficient CNN Inference
abstract
The growing computational demands of convolutional neural networks (CNNs) have motivated the use of spectral-domain inference as an alternative to costly spatial-domain convolutions. In this work, we propose a resistive RAM (RRAM)-based spectral-domain convolutional layer that exploits in-memory computing (IMC) for low energy consumption and high parallelism. Both the 2-D Fourier transform and the elementwise multiplications are directly executed on RRAM crossbar arrays, while Hermitian symmetry is leveraged to further enhance the energy efficiency of the transform and subsequent spectral processing. To ensure robustness, the measured RRAM device data are incorporated into system-level simulations to evaluate inference accuracy under the impact of device variability. Furthermore, we introduce a layer-wise mapping framework that adaptively selects between spatial- and spectral-domain execution based on the tradeoff between energy efficiency and accuracy. Simulation results show that the proposed design achieves up to a$2.18\times $improvement in energy efficiency across various convolutional layer configurations compared with the spatial-domain design. For VGG-8 on CIFAR-100, the proposed architecture with the layer-wise mapping scheme reduces the energy-delay product (EDP) by 45% while incurring negligible accuracy loss. This work presents the first complete RRAM-based spectral-domain convolutional layer that accounts for device variability, providing a promising solution for edge CNN inference.
Jianan Wen, Andrea Baroni, Christian Wenger, Milos Krstic, Letícia Maria Veiras Bolzani
IEEE Trans. Very Large Scale Integr. Syst.4
2025 EMBER: A Cycle-based Framework for Early-Stage Reliability Assessment in Parametric RTL Designs
abstract
Modern trends towards higher architectural complexity and smaller technology nodes do not align with the requirement for reliability that many application domains expose. While system robustness remains one of the most critical aspects for missioncritical application domains, the currently available tools allow designers to accurately investigate the reliability only at the late design stages, limiting the effectiveness of their intervention in addressing architectural vulnerabilities.In this context, this paper presents a novel framework for early-stage reliability assessment that enables fast evaluations to guide the RTL design process, without compromising analysis quality or relying on imprecise high-level fault injection models. In the presented analysis, the paper shows how the proposed framework leads to a significant time saving when targeting highly parametric hardware designs, achieving up to 79 x compile time reduction, and up to $37.6 x$ simulation time reduction when compared to other state-of-the-art approaches.
Alessandro Veronesi, Letícia Maria Veiras Bolzani, Michele Favalli, Milos Krstic, Davide Bertozzi
ATS4
2025 Multi-Partner Project: Twinning for Excellence in Reliable Electronics (TWIN-RELECT)
abstract
Reliable electronics plays a major role in shaping our daily lives, being a key enabler for critical applications, such as space missions, avionics, automotive, medicine, banking, automated industry, wireless communication networks, etc. However, design of highly reliable electronic systems remains a challenge with the advances in semiconductor technology and increase in integrated circuit (IC) complexity. In this work, we introduce the Horizon Europe Twinning project TWIN-RELECT, aimed at strengthening the scientific expertise in designing reliable integrated circuits. The paper presents the general project concept and objectives, and main directions of the joint research activities. The primary scientific goal is to contribute to the development of novel, more efficient, European Electronic Design Automation (EDA) tool-chain for design of reliable chips.
Marko S. Andjelkovic, Fabian Vargas 0001, Milos Krstic, Luigi Dilillo, Alain Michez, Frédéric Wrobel, Davide Bertozzi, Mikel Luján, Christos Georgakidis, Nikolaos Chatzivangelis, Katerina Tsilingiri, Nikolaos Zazatis, Georgios Ioannis Paliaroutis, Pelopidas Tsoumanis, Christos P. Sotiriou
DATE3
2025 AIDA4Edge: Twinning for Excellence in Adaptive Edge Artificial Intelligence
abstract
The growing demand for deployment of Artificial Intelligence (AI) on resource-constrained edge devices has motivated extensive research on the design of efficient edge-compatible AI hardware accelerators. One of the most promising solutions are the self-adaptive AI accelerators, capable of optimizing in real time their performance and energy consumption according to application requirements. This work introduces the EU-funded project Twinning for Excellence in Adaptive Edge Artificial Intelligence (AIDA4Edge), aimed to advance the state-of-the-art in the design of adaptive neural network accelerators for edge applications. The main goal is to develop a novel hybrid self-adaptive neural network architecture combining spiking and artificial neural networks, and supporting runtime adaptation of network functionality, precision and reliability. Furthermore, we aim to enhance the neural network training by incorporating hardware and quantization constraints in an automated tuning engine.
Marko S. Andjelkovic, Rizwan Tariq Syed, Alessandro Veronesi, Fabian Vargas 0001, Markus Ulbricht 0002, Letícia Maria Veiras Bolzani, Milos Krstic, Davide Bertozzi, Edward G. Jones, Oliver Rhodes, Riccardo Zese, Michele Favalli, Alice Bizzarri, Evelina Lamma, Marco Gavanelli, Elena Bellodi, Zoran H. Peric, Jelena Nikolic, Milan R. Dincic, Aleksandra Jovanovic 0001, Dejan Ciric, Nikola Vucic, Sofija Peric, Jelena Jovanovic 0006, Milica Stojanovic, Tatjana R. Nikolic, Goran Nikolic, Jelena Nedeljkovic, Danijel Dankovic, Emilija Zivanovic, Milos Marjanovic, Sandra Veljkovic, Nikola Mitrovic, Bratislav Predic, Tamara Milovanovic
DSD7
2025 Self-Aware Silicon: Enhancing Lifecycle Management with Intelligent Testing and Data Insights
Fabian Vargas 0001, Marko S. Andjelkovic, Milos Krstic, Anirban Kar, Swati Deshwal, Yogesh Singh Chauhan, Hussam Amrouch, Daniel Tille, Sebastian Huhn 0001
ETS3
2025 European Test Symposium Teams: an Anniversary Snapshot
abstract
The IEEE European Test Symposium (ETS) has been facilitating progress in electronic systems testing since its launch in 1996. On the occasion of its 30th anniversary, this collaborative paper gathers sections by 21 ETS teams to outline their influential ideas and milestones. Each team’s section highlights historical perspective, current research, frameworks and projects as well as forward-looking research agendas in the area of electronic-based circuits and systems testing, reliability, safety, security and validation. This anniversary summary documents how research of various ETS teams, exemplifying the test community, has been evolving and transitioning from concepts to practical standards and Electronic Design Automation (EDA) tools and flows. This legacy is a strong base to drive the next generation of advances in electronic systems testing.
Maksim Jenihhin, Jaan Raik, Artur Jutman, Natalia Cherezova, Raimund Ubar, Liviu Miclea, Szilárd Enyedi, Iulia Stefan, Ovidiu Stan, Cosmina Corches, Zebo Peng, Petru Eles, Rolf Drechsler, S. Eggersglüß, Görschwin Fey, Andreas Glowatz, Daniel Tille, Georges Gielen, Anthony Coyette, Wim Dobbelaere, Ronny Vanhooren, Po-Yao Chuang, Erik Jan Marinissen, Giorgio Di Natale, M. Barragan, Paolo Maistri, S. Mir, Vatajelu I. Vatajelu, Paolo Bernardi 0002, Stefano Di Carlo, Paolo Prinetto, Matteo Sonza Reorda, Massimo Violante, Haralampos-G. D. Stratigopoulos, M. K. Michael, Stelios Neophytou, Stavros Hadjitheophanous, Kyriakos Christou, M. Skitsas, Alberto Bosio, Bastien Deveautour, Patrick Girard 0001, Marcello Traiola, Arnaud Virazel, Fernando Santos 0001, Angeliki Kritikakou, Gioele Casagranda, Marzio Vallero, Flavio Vella, Paolo Rech, Letícia Maria Veiras Bolzani, Milos Krstic, Marko S. Andjelkovic, Fabian Vargas 0001, Grigor Tshagharyan, Gurgen Harutunyan, Valery A. Vardanian, Samvel K. Shoukourian, Yervant Zorian, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, Mottaqiallah Taouil, Moritz Fieback, Anteneh Gebregiorgis, Rajendra Bishnoi, Said Hamdioui, Abhijit Chatterjee, Anurup Saha, Suhasini Komarraju, K. Ma, Chandramouli N. Amarnath, Mehdi Baradaran Tahoori, Mahta Mayahinia, Maryam Rajabalipanah, Katayoon Basharkhah, N. Nosrati, Zahra Jahanpeima, Zainalabedin Navabi, Hans-Joachim Wunderlich, Sybille Hellebrand
ETS52
2025 Heterogeneous Integration of Advanced CMOS and Emerging Devices: Challenges and Solutions
Letícia Maria Veiras Bolzani, André Lucas Chinazzo, Mahdi Benkhelifa, Anirban Kar, Hussam Amrouch, Milos Krstic
ETS6
2025 ReDiM: An Efficient Strategy for Read Disturb Mitigation in RRAM-Based Accelerators
abstract
Resistive RAM (RRAM) has emerged as a promising non-volatile memory technology for implementing energy-efficient hardware accelerators within the in-memory computing (IMC) paradigm. However, due to the immature fabrication process and inherent material instabilities, frequent read operations during computations can induce read disturb effects, leading to unintended resistance drift and potential data corruption. Existing mitigation approaches primarily focus on detecting read disturb effects and triggering memory refresh operations. In this work, we propose an architecture-level solution that mitigates read disturb in RRAM-based accelerators. Our strategy employs crossbar duplication and decomposes the single high input pulse into two lower-amplitude pulses, effectively minimizing the risk of read disturb. To validate our approach, we develop a simulation framework that incorporates measurement data from characterized RRAM devices under read disturb stress conditions. Experimental results on VGG-8 with CIFAR-10 demonstrate that the proposed method significantly mitigates inference accuracy degradation caused by read disturb in RRAM-based accelerators, while incurring modest area and energy overheads of 12.32% and 2.15%, respectively. This work provides a practical and scalable solution for enhancing the robustness of RRAM-based accelerators in edge and high-performance computing applications.
Jianan Wen, Andrea Baroni, Alberto Mistroni, Cristian Zambelli, Christian Wenger, Milos Krstic, Letícia Maria Veiras Bolzani
IOLTS7
2025 Neuromorphic Edge Computing: Challenges, Opportunities, and Current Solutions
abstract
Neuromorphic computing is emerging as a paradigm for high-performance, energy-efficient edge intelligence. Yet the transition from laboratory prototypes to deployable edge platforms is slowed by four intertwined obstacles: (1) complex near-sensor integration, where spiking inference must co-locate with analogue sensing to minimise latency and data-movement energy; (2) novel event-based optimisation, requiring weight compression and sparsity techniques tailored to event-driven workloads; (3) heterogeneous integration of emerging devices, such as RRAM and other non-volatile memories, into reliable, manufacturable stacks; and (4) novel security risks, including spike-pattern side channels and model-specific attacks that demand to develop neuromorphic security primitives. This paper surveys the state of the art across these four fronts, drawing on recent advances in spiking microcontrollers, mixed-precision compute-in-memory fabrics, sparsity-aware compilation, and hardware-anchored security primitives (physical unclonable functions, true random number generators, computing-in-memory-based cryptography). By distilling lessons from academic research and industrial prototyping, the paper outlines current solutions and future research directions aimed at accelerating the adoption of neuromorphic platforms in real-world edge AI systems.
Federico Corradi, Amir Zjajo, Letícia Maria Veiras Bolzani, Milos Krstic, Orlando Moreira, Zeqi Zhu, Farhad Merchant
ISLPED4
2025 OTFS Modulation on SDR Platform: Experimental Demonstration and Performance Analysis
abstract
Utilization of the Delay-Doppler (DD) domain enables the recently proposed two-dimensional (2D) Orthogonal Time Frequency Space (OTFS) waveform to provide consistent performance under high mobility communication systems. OTFS outperforms existing standard waveforms under such time-frequency selective channels, making it a waveform candidate for future wireless communication systems. In this work, we present an implementation of an OTFS waveform-based system on a Universal Software Radio Peripheral (USRP) X310 Software Defined Radio (SDR). This system is tested in a realistic indoor office environment at a 5 GHz carrier frequency band with 150 MHz bandwidth. The resulting constellation diagrams were observed for BPSK, 4-QAM, and 16-QAM modulated OTFS symbols. Furthermore, for 4-QAM symbols, several communication metrics, such as Bit Error Rate (BER) and Error Vector Magnitude (EVM), were evaluated against different gains at the USRP. The resulting constellation diagrams, BER, and EVM graphs show a successful implementation of our OTFS system.
Lukasz Lopacinski, Nebojsa Maletic, Matthias Scheide, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass
PIMRC6
2025 OTFS Sensing with SDR: Experimental Results and Analysis
abstract
Localization will be an essential requirement for various$6^{\text{th}}$generation (6G) communication system applications. Integrated Sensing and Communication (ISAC) is seen as a key enabling technology that can provide the capability of combined communication and localization. Reliable ISAC in a high-mobility environment can be challenging and existing communication waveforms suffer from severe degradation due to significant Doppler effect. The Delay-Doppler (DD) domain can be commonly seen in the results of Radio Detection and Ranging (RADAR) systems, which is the baseline for Orthogonal Time Frequency Space (OTFS) modulation. Due to the information encoding in DD domain, OTFS shows significant resilience against doubly-selective channels. In this work, we report the sensing functionality performance of our complete sub-6GHz ISAC system based on the OTFS waveform, implemented on a USRP X310 Software Defined Radio (SDR). The sensing capability of the system is experimentally verified in both an anechoic chamber and in an office scenario for single and for multitargets.
Lukasz Lopacinski, Nebojsa Maletic, Matthias Scheide, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass
VTC2025-Spring6
2025 Analysis and Modeling of Single Event Transient Generation in Standard Combinational Cells
Marko S. Andjelkovic, Milos Krstic
J. Electron. Test.2
2025 Dynamic Fault Mitigation for Space Radiation Using Fault Injection and Machine Learning
Junchao Chen 0001, Marko S. Andjelkovic, Fabian Vargas 0001, Milos Krstic
J. Electron. Test.5
2025 Accelerate SEU Simulation-Based Fault Injection With Spatio-Temporal Graph Convolutional Networks
abstract
Evaluating the sensitivity of circuits to Single Event Upset (SEU) faults has become increasingly important and challenging due to the growing complexity of circuits. Simulation-based fault injection is time-intensive, particularly for highly complex circuits. This paper proposes a novel approach using Spatio-temporal Graph Convolutional Networks (STGCN) to predict SEU fault propagation results in circuits. By representing circuits’ structure as graphs and integrating temporal features from the simulation workload, STGCNs can learn from these spatio-temporal graphs to identify SEU fault propagation patterns. To validate this method, we test it on six evaluation circuits, achieving a prediction accuracy of 93-99%. Given this performance, to accelerate SEU simulation-based fault injection, we divide SEU faults into three subsets and use a STGCN fine-tuned on the training and validation dataset to predict SEU fault propagation in the test dataset, eliminating the need for simulation and reducing the required time. To identify an efficient dataset separation method, we compare three sampling methods: spatial sampling (sampling flip-flops for injected faults), temporal sampling (sampling time points for fault injection), and hybrid sampling (incorporating both spatial and temporal sampling). The hybrid sampling approach is the most promising, optimizing the trade-off between efficiency and accuracy. This approach reduces simulation time by 50% while maintaining accuracy above 95% on the six evaluation circuits.
Junchao Chen 0001, Aneesh Balakrishnan, Markus Ulbricht 0002, Milos Krstic
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 ImSTDP: Implicit Timing On-Chip STDP Learning
abstract
Spike-Timing-Dependent Plasticity (STDP) is a biological-plausible learning mechanism widely adopted for building Spiking Neural Networks (SNNs). It determines plasticity polarity and synapse strength change according to the timing difference between pre- and postsynaptic spikes. The learning curves of STDP differ in temporal window size, magnitude and polarity across different synapse types and brain regions and even within a cell, in different dendritic compartments. To accelerate on-chip STDP learning, various implementations have been proposed. However, they either introduce significant latency due to costly counter-based time difference calculation and substantial area cost due to the implementation of weight change LUTs, or lose biologically-plausible timing information due to oversimplification. For low-cost and efficient on-chip learning, a high-throughput Implicit-timing STDP (ImSTDP) with optimized SR depth and a low-cost register-based Implicit-Timing Look-up (ITL) are proposed. ASIC implementation in 22 nm technology demonstrates that ImSTDP can achieve up to$2\times $throughput improvement and$3.61\times $power efficiency improvement at 27% less area cost compared to the cutting-edge counter-LUT on-chip STDP learning solution.
Dedong Zhao, Oliver Schrape, Zoran Stamenkovic, Milos Krstic
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 RISC-V CPU Design Using RRAM-CMOS Standard Cells
abstract
The breakdown of Dennard scaling has been the driver for many innovations such as multicore CPUs and has fueled the research into novel devices such as resistive random access memory (RRAM). These devices might be a means to extend the scalability of integrated circuits since they allow for fast and nonvolatile operation. Unfortunately, large analog circuits need to be designed and integrated in order to benefit from these cells, hindering the implementation of large systems. This work elaborates on a novel solution, namely, creating digital standard cells utilizing RRAM devices. Albeit this approach can be used both for small gates and large macroblocks, we illustrate it for a 2T2R-cell. Since RRAM devices can be vertically stacked with transistors, this enables us to construct anandstandard cell, which merely consumes the area of two transistors. This leads to a 25% area reduction compared to an equivalent CMOSnandgate. We illustrate achievable area savings with a half-adder circuit and integrate this novel cell into a digital standard cell library. A synthesized RISC-V core using RRAM-based cells results in a 10.7% smaller area than the equivalent design using standard CMOS gates.
Markus Fritscher, Max Uhlmann, Philip Ostrovskyy, Daniel Reiser, Junchao Chen 0001, Jianan Wen, Carsten Schulze, Gerhard Kahmen, Dietmar Fey, Marc Reichenbach, Milos Krstic, Christian Wenger
IEEE Trans. Very Large Scale Integr. Syst.11
2024 Towards SEU Fault Propagation Prediction with Spatio-Temporal Graph Convolutional Networks
abstract
Assessing Single Event Upset (SEU) sensitivity in complex circuits is increasingly important but challenging. This paper proposes an efficient approach using Spatio-temporal Graph Convolutional Networks (STGCN) to predict the results of SEU simulation-based fault injection. Representing circuit structures as graphs and integrating temporal data from the workload's waveform into these graphs, STGCN achieves a 94-96% prediction accuracy on four test circuits.
Junchao Chen 0001, Markus Ulbricht 0002, Milos Krstic
DATE4
2024 Towards Reliable and Energy-Efficient RRAM Based Discrete Fourier Transform Accelerator
abstract
The Discrete Fourier Transform (DFT) holds a prominent place in the field of signal processing. The development of DFT accelerators in edge devices requires high energy efficiency due to the limited battery capacity. In this context, emerging devices such as resistive RAM (RRAM) provide a promising solution. They enable the design of high-density crossbar arrays and facilitate massively parallel and in situ computations within memory. However, the reliability and performance of the RRAM-based systems are compromised by the device non-idealities, especially when executing DFT computations that demand high precision. In this paper, we propose a novel adaptive variability-aware crossbar mapping scheme to address the computational errors caused by the device variability. To quantitatively assess the impact of variability in a communication scenario, we implemented an end-to-end simulation framework integrating the modulation and demodulation schemes. When combining the presented mapping scheme with an optimized architecture to compute DFT and inverse DFT(IDFT), compared to the state-of-the-art architecture, our simulation results demonstrate energy and area savings of up to 57 % and 18 %, respectively. Meanwhile, the DFT matrix mapping error is reduced by 83% compared to conventional mapping. In a case study involving 16-quadrature amplitude modulation (QAM), with the optimized architecture prioritizing energy efficiency, we observed a bit error rate (BER) reduction from 1.6e-2 to 7.3e-5. As for the conventional architecture, the BER is optimized from 2.9e-3 to zero.
Jianan Wen, Andrea Baroni, Max Uhlmann, Markus Fritscher, Karthik KrishneGowda, Markus Ulbricht 0002, Christian Wenger, Milos Krstic
DATE9
2024 6G-TakeOff: Holistic 3D Networks for 6G Wireless Communications
abstract
The unified 3D communication networks, integrating standard terrestrial mobile communication networks and non-terrestrial networks (NTNs), are seen as the key enabler for global connectivity in the next generation (6G) wireless communications. To achieve this goal, new technologies and components are needed in order to meet the requirements for the 6G networks in terms of higher data rates, and enhanced reliability, security and network reconfigurability. This work introduces the German project 6G-TakeOff, aimed at the design of solutions for unified 3D networks for 6G wireless communication systems. The project consortium brings together academic and industrial partners from Germany and Europe, covering the entire value chain from design of electro-nics to applications. This work presents the key hardware components required for 3D networks and the concept for demonstration of their functionality.
Marko S. Andjelkovic, Nebojsa Maletic, Nicola Miglioranza, Milos Krstic, Enrico Koeck, Jan Buchholz, Maike Taddiken, Markus Fehrenz, Shaden Baradie, Dirk Wübben, Markus Breitbach
DSD4
2024 Cross-Layer Reliability Analysis of NVDLA Accelerators: Exploring the Configuration Space
abstract
Investigating the effects of Single Event Upset in domain-specific accelerators represents one of the key enablers to deploy Deep Neural Networks (DNNs) in mission-critical edge applications. Currently, reliability analyses related to DNNs mainly focus either on the DNNs model, at application level, or on the hardware accelerator, at architecture level. This paper presents a systematic cross-layer reliability analysis of NVIDIA Deep-Learning Accelerator, a popular family of industry-grade, open and free DNN accelerators. The goals are i) to analyze the propagation of faults from the hardware to the application level, and ii) to compare different architectural configurations. Our investigation delivers new insights into the performance-accuracy-reliability trade-off spanned by the configuration space of Deep Learning accelerators. In particular, the Failure in Time can be reduced up to 4.3x for the same DNN model accuracy and by up to 9.4x for the same performance, while accounting 6.5x inference latency and 1.1% accuracy drop, respectively.
Alessandro Veronesi, Alessandro Nazzari, Dario Passarello, Milos Krstic, Michele Favalli, Luca Cassano, Antonio Miele, Davide Bertozzi, Cristiana Bolchini
ETS4
2024 High-Efficiency Gesture Recognition Using Multiple mmWave FMCW RADARs
abstract
RADAR-based gesture recognition has attracted a lot of attention in recent years. A large number of studies have been conducted on single RADAR-based gesture recognition. However, there is still a lot of space to explore multi-RADAR-based gesture recognition. Compared to a single-RADAR scenario, a multi-RADAR scenario can provide higher stability and recognition performance. In the context of the multi-RADAR scenario, an efficient, high-performance algorithm is proposed in this paper for the purpose of extracting features from RADAR data and classifying gestures using a simple machine-learning model. The experimental results demonstrate a high average recognition rate of 97.78% on the test set with the proposed algorithm and model, and a significant reduction in runtime compared to the reference work. This can be a very promising application in joint communication and sensing (JCAS) system.
Yanhua Zhao, Vladica Sark, Milos Krstic, Eckhard Grass
ISNCC3
2024 Reliability Assessment of Large DNN Models: Trading Off Performance and Accuracy
abstract
The adoption of Deep Neural Networks (DNNs) in several domains allows for increased effectiveness in applications that deal with massive data-intensive and complex data inputs. When employed in safety-critical scenarios, such as automotive, aerospace, healthcare, and autonomous robotics, assessing the DNNs' reliability and functional safety is crucial to ensure their correct in-field operation, even in the presence of hardware faults. However, the system complexity and the massive amounts of data to be processed by DNNs prevent the effective adoption of traditional strategies for reliability characterization and for identifying the most fault-sensitive structures. Accurate fault assessment strategies usually require unacceptable computational power and large evaluation times. On the other hand, faster strategies commonly lack accuracy in correctly representing system faults. Consequently, it is necessary to develop effective strategies that trade-off between performance and accuracy. This work analyses three reliability assessment strategies for deep neural networks and their underlying hardware, highlighting the main solutions and challenges in terms of evaluation performance and fault characterization accuracy. We overview different solutions to evaluate the hardware accelerators implementing DNNs at three abstraction levels:$i$) by physically injecting faults on a GPU running DNNs, ii) by performing microarchitectural characterization of GPUs to develop application-accurate error models, and iii) by using structure-aware cross-layer error modeling on DNN hardware accelerators. Our experimental results indicate that accurate error representation requires structural features from the targeted hardware.
Junchao Chen 0001, Giuseppe Esposito, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Angeliki Kritikakou, Milos Krstic, Robert Limas Sierra, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Marcello Traiola, Alessandro Veronesi
VLSI-SoC6
2024 Toward Critical Flip-Flop Identification for Soft-Error Tolerance With Graph Neural Networks
abstract
Nanometer circuits are becoming increasingly susceptible to soft errors. Selective hardening is a less expensive technique to improve the reliability of circuits because it hardens the critical components instead of hardening an entire circuit. One challenge of selective hardening is efficiently and effectively identifying the critical parts in circuits. Simulation-based fault injection is commonly used but extremely time-consuming, especially for complex circuits. This article proposes an approach based on graph neural networks (GNNs) to identify critical flip-flops in circuits. GNNs can take advantage of the circuit’s structural features and the features of individual flip-flops. To convert the features into abstract data that can be fed into GNNs, we provide a feature extraction method that uses a graph model to represent the relevant features of the circuit. The method converts the target circuit into a graph representing its architecture. The graph also contains features of individual flip-flops extracted from the circuit’s netlist and the value change dump (VCD) waveforms of the test used for fault simulation. Additionally, we extract edge features in the graph to utilize the information on combinational gates on the path between flip-flops. Datasets generated based on two open-source RISC-V cores are used to validate the proposed approach. We compare the performance of different GNNs on them and discuss the contribution of edge features to their performance. Our experiments show that the prediction accuracy increases significantly with edge features. GraphSAGE and SAGE-GCN with edge features perform best among the selected GNNs. The highest accuracy we achieved on Ibex and RI5CY is 97.75% and 98.67%, respectively. We also provide a method to accelerate the process of critical flip-flop identification.
Junchao Chen 0001, Markus Ulbricht 0002, Milos Krstic
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Machine Learning Methodologies to Predict the Results of Simulation-Based Fault Injection
abstract
Simulation-based fault injection is a widely used technique for early-stage circuit reliability analysis. However, it consumes significant time, particularly for complex circuits. This paper introduces two Machine Learning (ML) methodologies to predict simulation-based fault injection outcomes at the gate level. The initial approach employs Neural Networks (NNs), extracting structural features from synthesis reports and simulation-related characteristics from Value Change Dump (VCD) waveforms. Nevertheless, NNs are restricted to learning from individual gate attributes. To exploit the comprehensive structure of entire circuits, we propose a method to convert circuits into graphs. This facilitates the utilization of Graph Neural Networks (GNNs) as advanced models, resulting in improved prediction performance. We select six open-source circuits with diverse complexities and functions to validate these methodologies and explore their adaptability across various circuits. Our experiments demonstrate the superior performance of GNNs compared to NNs in terms of prediction accuracy, efficiency in hyperparameter search, and the ability to address imbalanced datasets. Additionally, we investigate the feasibility of deploying the trained models to predict results in new circuits. Based on the experimental outcomes, we present an approach for leveraging the proposed methodology to accelerate simulation-based fault injection.
Junchao Chen 0001, Markus Ulbricht 0002, Milos Krstic
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Bits, Flips and RISCs
abstract
Electronic systems can be submitted to hostile environments leading to bit-flips or stuck-at faults and, ultimately, a system malfunction or failure. In safety-critical applications, the risks of such events should be managed to prevent injuries or material damage. This paper provides a comprehensive overview of the challenges associated with designing and verifying safe and reliable systems, as well as the potential of the RISC-V architecture in addressing these challenges.We present several state-of-the-art safety and reliability verification techniques in the design phase. These include a highly-automated verification flow, an automated fault injection and analysis tool, and an AI-based fault verification flow. Furthermore, we discuss core hardening and fault mitigation strategies at the design level. We focus on automated SoC hardening using model-driven development and resilient processing based on sensing and prediction for space and avionic applications.By combining these techniques with the inherent flexibility of the RISC-V architecture, designers can develop tailored solutions that balance cost, performance, and fault tolerance to meet the requirements of various safety-critical applications in different safety domains, such as avionics, automotive, and space. The insights and methodologies presented in this paper contribute to the ongoing efforts to improve the dependability of computing systems in safety-critical environments.
Nicolas Gerlin, Endri Kaja, Fabian Vargas 0001, Anselm Breitenreiter, Junchao Chen 0001, Markus Ulbricht 0002, Maribel Gomez, Ares Tahiraga, Sebastian Siegfried Prebeck, Eyck Jentzsch, Milos Krstic, Wolfgang Ecker
DDECS12
2023 Towards a Smart Multi-Sensor Ionizing Radiation Monitoring System
abstract
Detection and measurement of ionizing radiation is required in a wide range of terrestrial applications, as well as in space missions. For this purpose, special instruments composed of radiation sensors and readout electronics are utilized. As ionizing radiation may affect the operation of electronic systems, radiation hardness is one of the main design requirements for radiation monitoring systems. In this work, we present a concept of a smart multi-sensor radiation monitoring system. The proposed design is based on the results achieved within the framework of EU-funded ELICSIR project. Our solution provides a new perspective on smart radiation monitoring by combining the concepts of self-awareness, self-adaptivity and artificial intelligence. This solution is suitable for applications where long-term autonomous radiation monitoring is required, such as environmental monitoring at terrestrial level or radiation monitoring in space missions.
Marko S. Andjelkovic, Junchao Chen 0001, Rizwan Tariq Syed, Fabian Vargas 0001, Markus Ulbricht 0002, Milos Krstic, Stefan D. Ilic, Milos Marjanovic, Sandra Veljkovic, Nikola Mitrovic, Danijel Dankovic, Goran S. Ristic, Russell Duane, Nikola Vasovic, Aleksandar Jaksic, Alberto J. Palma, Antonio M. Lallena, Miguel Ángel Carvajal
DSD6
2023 PULP Fiction No More - Dependable PULP Systems for Space
abstract
Due to their flexibility and openness, the RISC-V ISA and processor architectures have emerged as notable contenders in various application domains. Their advantages over commercial solutions have attracted the interest of academia and industry and even led to their planned adoption in aeronautics and space. However, in these demanding environments, system reliability is of paramount importance. To address this issue, this paper presents an overview of several hardware-centric approaches for developing reliable systems based on the parallel-ultra low power (PULP) open-source RISC-V hardware platform. These approaches range from gate-level optimizations to system-level improvements and highlight the versatility of the PULP architecture and its potential as a viable architecture for developing various aerospace platforms.
Markus Ulbricht 0002, Yvan Tortorella, Michael Rogenmoser, Junchao Chen 0001, Francesco Conti 0001, Milos Krstic, Luca Benini
ETS7
2023 Amplitude- and phase-modulated PSSS for wide bandwidth mixed analog-digital baseband processors in THz communication
abstract
This paper proposes modifications for the parallel sequence spread spectrum (PSSS) modulation scheme, which improve the bit error rate (BER) performance and peak-to-average power ratio (PAPR) at the same time. In our scheme, some data bits are encoded as phase shifts of the spreading sequences. Thus, the number of transmitted sequences can be reduced, and the resulting PAPR is lowered. The improvements proposed here are inspired by code shift keying (CSK) and parallel combinatory spread spectrum (PCSS) systems. The PSSS variant investigated here is based on recently discovered real-value spreading sequences, where all sequence coefficients are defined in the real domain.
Lukasz Lopacinski, Nebojsa Maletic, Rolf Kraemer, Alireza Hasani, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass
VTC2023-Spring6
2023 Gesture Recognition Using Multiple mmWave FMCW Radars
abstract
Human-computer interaction using radar has gained a lot of attention. Gesture recognition with a single Frequency Modulated Continuous Wave (FMCW) radar has been investigated by many researchers, and several promising results have been achieved. However, the stability of such a system still needs to be improved. This paper proposes a signal-processing approach for gesture recognition based on multiple FMCW radars.VGG16 pre-trained model is employed to extract features from the dataset, after which gesture recognition is performed using a Support Vector Machine (SVM) classifier and an XGBoost classifier, respectively. Moreover, the effect of the incomplete dataset in case of failure of one of the radars on the accuracy of gesture recognition is also analysed. The experimental results show that the SVM classifier in the multiple radar scenario obtains up to 96.67% accuracy on the test set, which is 1.78% to 5.95% higher than the single radar scenario. In the multi-radar scenario, when the extracted dataset is incomplete, the SVM classifier achieves up to 95% recognition accuracy. This proves that multi-radar is more stable than the single radar scheme. This system can be applied in smart homes, in-car entertainment systems or smart factories.
Yanhua Zhao, Vladica Sark, Milos Krstic, Eckhard Grass
VTC Fall3
2023 Fully Parallel Fully Unrolled BP Decoding of LDPC and Polar Codes
abstract
Low-Density Parity-Check (LDPC) and polar codes are two classes of channel codes potentially able to achieve the channel capacity, and both are used in current 5G communication systems. Belief Propagation (BP), which is the main decoding algorithm for LDPC codes could also be used for decoding polar codes. A fully parallel fully unrolled BP decoding architecture has been recently proposed for both of these codes. In this paper, we investigate these two architectures further and compare the two codes from the viewpoint of Bit-Error Rate (BER) performance, achieved throughput, latency, chip area, and also energy efficiency. Toward that, both architectures have been synthesized in a 28 nm CMOS technology, and corresponding results are provided.
Alireza Hasani, Lukasz Lopacinski, Milos Krstic, Eckhard Grass
WCNC3
2023 Cost-Effective Path Delay Defect Testing Using Voltage/Temperature Analysis Based on Pattern Permutation
Tai Song, Zhengfeng Huang, Xiaohui Guo, Milos Krstic
J. Electron. Test.4
2022 The Scale4Edge RISC-V Ecosystem
abstract
This paper introduces the project Scale4Edge. The project is focused on enabling an effective RISC-V ecosystem for optimization of edge applications. We describe the basic components of this ecosystem and introduce the envisioned demonstrators, which will be used in their evaluation.
Wolfgang Ecker, Peer Adelt, Wolfgang Müller 0003, Reinhold Heckmann, Milos Krstic, Vladimir Herdt, Rolf Drechsler, Gerhard Angst, Ralf Wimmer 0001, Andreas Mauderer, Rafael Stahl, Karsten Emrich, Daniel Mueller-Gritschneder, Bernd Becker 0001, Philipp M. Scholl, Eyck Jentzsch, Jan Schlamelcher, Kim Grüttner, Paul Palomero Bernardo, Oliver Bringmann 0001, Brindusa Mihaela Damian-Kosterhon, Julian Oppermann, Andreas Koch 0001, Jörg Bormann, Johannes Partzsch, Christian Mayr 0001, Wolfgang Kunz
DATE5
2022 Exploring Software Models for the Resilience Analysis of Deep Learning Accelerators: the NVDLA Case Study
abstract
Deep learning accelerator models described with software imperative languages are frequently used for their large-scale reliability analysis in order to overcome the prohibitive simulation times of logic-level and RTL models. However, they are faced with the challenge of preserving consistency between software-visible variables and faulty microarchitectural states. The goal of this work is to determine a suitable accelerator modelling that enables analysis without overloading the simulation engine. Toward this goal, the paper explores different accelerator modelling strategies featuring increasing levels of hardware visibility. They are compared in their capability to gain insights into the reliability of the multiply-and-accumulate (MAC) pipeline of an industry-standard deep learning accelerator from NVIDIA. Our results show that subtle microarchitectural details that are typically overlooked by competing approaches play a relevant role in determining accelerator reliability.
Alessandro Veronesi, Francesco Dall'Occo, Davide Bertozzi, Michele Favalli, Milos Krstic
DDECS5
2022 1542 Gbps Fully Pipelined Fast-SSC Decoding of Polar Codes
abstract
Polar codes are one of the candidates as a Forward-Error Correction (FEC) scheme for next-generation high-throughput communication systems. Successive Cancellation (SC) is one of the main decoding algorithms for decoding polar codes. In this paper, a fully pipelined fast Simplified SC (SSC) decoding architecture is proposed which proves a throughput of 1542 Gb/s in a 28 nm CMOS ASIC technology. This architecture is fully pipelined in the sense that sequences are decoded successively in consecutive clock cycles. The implementation results after the design Place&Route show improvements in terms of decoding throughput, clock frequency, core area, and energy efficiency compared with the state-of-the-art.
Alireza Hasani, Lukasz Lopacinski, Milos Krstic, Eckhard Grass
PIMRC3
2022 An Improved Stage-Combined Belief Propagation Decoding of Polar Codes
abstract
Belief Propagation (BP) algorithm is an alternative decoding method to Successive Cancellation (SC) decoding of polar codes with an advantage of higher throughput and lesser latency. BP algorithm is performed over the factor graph of a polar code, and its throughput is oppositely proportional to the number of stages of the factor graph. The stage-combining idea can be applied to the factor graph of a polar code to halve the number of stages, and thus to double the decoding throughput. By this idea, every two adjacent stages of a factor graph are combined and processed in a single step. In this paper, we propose a modified stage-combined BP decoding method that is able to improve the Bit-Error Rate (BER) performance of the decoding algorithm. This improvement is almost 0.5 dB at low to intermediate SNRs, and much more considerable at high SNRs, since the undesired phenomenon of error floor is mitigated as a result of the proposed stage-combining mechanism.
Alireza Hasani, Lukasz Lopacinski, Milos Krstic, Eckhard Grass
PIMRC3
2022 A Hardware Optimized High Throughput LDPC Decoder Supporting 3 Tb/s in 28 nm CMOS
abstract
This paper proposes an optimized pipelined decoding architecture with seven processing stages for unrolled LDPC decoders. Pipelined design with seven register layers significantly increases the resulting clock frequency. Moreover, we investigate the optimal layout shape for unrolled decoders. This paper's fastest decoder is based on the IEEE 802.11n LDPC(1944,1620) parity-check matrix and achieves 2937 Gb/s of coded throughput after physical design. By optimizing the pipeline, floorplan, and employing a codeword length of 1944 bits, we increased the throughput by 241% compared to the previous fastest LDPC decoder presented in the literature. To the best of our knowledge, it is the fastest soft-decision decoder published so far. The standard min-sum approach is employed for decoding, and the proposed improvements consider changes only on the hardware level.
Lukasz Lopacinski, Alireza Hasani, Goran Panic, Nebojsa Maletic, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass, Rolf Kraemer
PIMRC6
2022 High-Speed SC Decoder for Polar Codes achieving 1.7 Tb/s in 28 nm CMOS
abstract
This paper compares three hardware variants of successive cancelation (SC) decoders for polar codes. The fastest implementation, based on the basic SC, provides decoding throughput up to 1700 Gb/s, when implemented in a 28 nm CMOS technology at the worst-case timing corner. This is the fastest polar decoder published so far, to the best of our knowledge. We also discuss the difficulties of implementing single-parity-check nodes and repetition nodes in fast simplified SC (Fast-SSC) decoding algorithm. These two node types are the primary sources of clock frequency reduction, and special care needs to be taken when these elements are implemented. The Fast-SSC decoder requires ~3 times fewer hardware resources than the base version of SC, but achieves ~10% lower decoding throughput.
Lukasz Lopacinski, Alireza Hasani, Goran Panic, Nebojsa Maletic, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass
VLSI-SoC6
2022 Ultra high speed 802.11n LDPC decoder with seven-stage pipeline in 28 nm CMOS
abstract
This paper reports our latest implementation results of a fully unrolled LDPC decoder prototyped in 28 nm CMOS technology. The decoder achieves 1218 Gbps coded throughput and consumes a 5.49 mm2chip area. The standard min-sum decoding algorithm with four-bit quantization, five unrolled iterations, (648,540) parity matrix, and a seven-stage pipeline is employed. Such implementation achieves a higher data rate than adaptive degeneration and finite-alphabet decoding algorithms, requires less silicon than the solutions mentioned above, and is fully compliant with the IEEE 802.11n WLAN standard.
Lukasz Lopacinski, Alireza Hasani, Goran Panic, Nebojsa Maletic, Oliver Schrape, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass, Rolf Kraemer
VTC Spring7
2022 Novel Approach for Gesture Recognition Using mmWave FMCW RADAR
abstract
Hand gesture recognition driven by RADAR technology has attracted significant attention in recent years. Among various RADAR types, frequency-modulated continuous-wave (FMCW) RADAR is used in this work due to its very high range and velocity resolution. However, data collected by RADAR are disturbed by static background and static clutter. Therefore, a novel data preprocessing approach is proposed to remove the static background and clutter in the acquired data. A convolutional neural network is used to extract the features of the acquired data set. To the best of our knowledge, this is the first time that range, velocity and angle features are combined in one map, forming the input signal of a convolutional neural network. Classifiers are applied to recognize gestures. Experimental results show that the proposed method using the XGBoost classifier can achieve a high recognition accuracy of 98.93% on the test set. In contrast, the proposed method with the random forest classifier can achieve a recognition rate of 100% on the same test set with six dynamic hand gestures. This approach could be useful in aspects such as in-car entertainment systems and smart homes.
Yanhua Zhao, Vladica Sark, Milos Krstic, Eckhard Grass
VTC Spring3
2021 Design and Implementation Strategy of Adaptive Processor-Based Systems for Error Resilient and Power-Efficient Operation
abstract
The contemporary computing systems are facing two major challenges: excessive power consumption and susceptibility to faults. In order to take advantage of techniques that efficiently address these challenges, the classic ASIC design flow requires some modifications. In this paper, we present a simple and convenient strategy for design and implementation of processor-based systems using highly configurable, cross-layer framework that encompasses techniques such as Adaptive Voltage and Frequency Scaling (AVFS) and Triple Modular Redundancy (TMR). The proposed strategy augments the conventional design flow with two additional steps to integrate the framework's hardware building blocks into the system. Such system is then able to dynamically switch between low power and error resilient operation modes according to the current requirements. By following the proposed strategy, we were able to implement processor-based system that significantly reduces the power consumption / increases the soft error resilience while preserving the performance at negligible area overhead of less than 1%.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
DDECS3
2021 Behavioral Model of Dot-Product Engine Implemented with 1T1R Memristor Crossbar Including Assessment
abstract
Memristor is an emerging electrical device that enables non-volatile storage and in-memory computing. The memristive crossbar with high memory density and low energy consumption has drawn much attention for the implementation of dot-product engines, which can be deployed in power-hungry applications with intensive multiply-accumulate operations. However, simulating the crossbar containing a group of memristors based on the device-level modeling is time consuming. In this paper, we propose a model to simulate the memristive crossbar with high flexibility and automation at the behavioral level to perform the vector-matrix multiplication. This system-level model captures the non-linearity of memristors aiming for fast and accurate simulation. With the significantly reduced simulation time, this model enables simulating the systems containing memristive crossbar with large scale like neural networks in a more practical way. Moreover, this model can be exploited to analyze the effects of variations, which provides a condition and contributes to revealing potential computational errors. A multilayer perceptron detecting breast cancer is simulated based on this model to assess the classification accuracy with the presence of variabilities.
Jianan Wen, Markus Ulbricht 0002, Xin Fan 0003, Milos Krstic
DDECS5
2021 Classification of Space Particle Events using Supervised Machine Learning Algorithms
abstract
Solar Particle Events (SPEs) generate cosmic radiation of different magnitude in a time span of several hours or even days. This contributes to an increased probability of higher magnitude Single-Event Upsets (SEUs) occurrence in space applications. It is critical to establish early detection of SEU rate or Soft Error Rate (SRE) changes to enable timely radiation hardening measures. This research paper focuses on the high-accuracy detection of SPEs using the manually collected space data. Additionally, the prediction of SRE increase or decrease was established with the seven widely used supervised machine learning algorithms. Excellent performance of 97.82%, including a high F1-score, was achieved during the presence of SPE using$k$-Nearest Neighbor algorithms.
Rijad Saric, Junchao Chen 0001, Milos Krstic, Edhem Custovic, Goran Panic, Jasmin Kevric, Dejan Jokic
DSAA3
2021 Towards Error Resilient and Power-Efficient Adaptive Multiprocessor System using Highly Configurable and Flexible Cross-Layer Framework
abstract
A typical multiprocessor system often needs to support a wide spectrum of applications. Today, error resilience and low power consumption are two crucial, but non-complementary requirements and meeting both simultaneously is difficult. Thus, adaptivity is becoming increasingly important feature for modern computing systems. In this regard, we integrate a highly-configurable framework with a set of cross-layer techniques efficient in improving error resilience / power consumption into a multiprocessor system. The framework intelligently interchanges methods such as Adaptive Voltage and Frequency Scaling (AVFS), Triple Modular Redundancy (TMR) and clock-gating while the system is on-line. Additionally, flexibility as an inherent multiprocessor feature enables dynamical adaptation of the system to the current requirements. Putting all together, a balanced level between the two key metrics is achieved. We conduct numerous experiments to show the advantages of the proposed approach. Finally, we use the results to confirm the benefits and the effectiveness of the framework.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
IOLTS3
2021 Plesiochronous Spread Spectrum Clocking With Guaranteed QoS for In-Band Switching Noise Reduction
abstract
Spread spectrum clocking (SSC) conventionally uses frequency modulation (FM) to suppress digital switching noise in the frequency domain. While clock-FM effectively reduces spectral noise peaks, it maintains the synchronous operation per cycle with total noise unchanged. In this paper, we introduce plesiochronous design as a general applicable de-synchronization solution for the spectral switching noise optimization with guaranteed quality-of-service. By modeling on-chip aperiodic supply current as a poly-cyclostationary random process, we theoretically prove that digital plesiochronous design contributes to reducing both, total and peak switching noise, in a harmonic frequency band of interest logarithmically proportional to the number of adopted clock domains over the synchronous baseline. A complete framework is also developed to implement plesiochronous design with the optimal clock domain partitioning and FIFO-based synchronization that features a minimum depth of six by employing Johnson encoding fully compatible with mainstream design flow. Validated on a 130nm pipelined FFT test chip across 25 dies thus taking process variations into account, our plesiochronous SSC achieves on average 5.1dB total power reductions in addition to 12.8dB peak power reductions of substrate noise at the clock fundamental frequency, which match our predictions, with only marginal hardware overhead in terms of cell area and power consumption.
Xin Fan 0002, Milan Babic, Eckhard Grass, Milos Krstic
IEEE Trans. Circuits Syst. I Regul. Pap.5
2021 Design and Evaluation of Radiation-Hardened Standard Cell Flip-Flops
abstract
Use of a standard non-rad-hard digital cell library in the rad-hard design can be a cost-effective solution for space applications. In this paper we demonstrate how a standard non-rad-hard flip-flop, as one of the most vulnerable digital cells, can be converted into a rad-hard flip-flop without modifying its internal structure. We present five variants of a Triple Modular Redundancy (TMR) flip-flop: baseline TMR flip-flop, latch-based TMR flip-flop, True-Single Phase Clock (TSPC) TMR flip-flop, scannable TMR flip-flop and self-correcting TMR flip-flop. For all variants, the multi-bit upsets have been addressed by applying special placement constraints, while the Single Event Transient (SET) mitigation was achieved through the usage of customized SET filters and selection of optimal inverter sizes for the clock and reset trees. The proposed flip-flop variants feature differing performance, thus enabling to choose the optimal solution for every sensitive node in the circuit, according to the predefined design constraints. Several flip-flop designs have been validated on IHP’s 130nm BiCMOS process, by irradiation of custom-designed shift registers. It has been shown that the proposed TMR flip-flops are robust to soft errors with a threshold Linear Energy Transfer (LET) from ($32.4\, \frac { {\mathrm { \text {M} \text {eV} }}\cdot {\mathrm { \text {c} \text {m} }}^{2}}{ {\mathrm { \text {m} \text {g} }}}$) to ($62.5\, \frac { {\mathrm { \text {M} \text {eV} }}\cdot {\mathrm { \text {c} \text {m} }}^{2}}{ {\mathrm { \text {m} \text {g} }}}$), depending on the variant.
Oliver Schrape, Marko S. Andjelkovic, Anselm Breitenreiter, Steffen Zeidler 0001, Alexey Balashov, Milos Krstic
IEEE Trans. Circuits Syst. I Regul. Pap.6
2020 RESCUE: Interdependent Challenges of Reliability, Security and Quality in Nanoelectronic Systems
abstract
The recent trends for nanoelectronic computing systems include machine-to-machine communication in the era of Internet-of-Things (IoT) and autonomous systems, complex safety-critical applications, extreme miniaturization of implementation technologies and intensive interaction with the physical world. These set tough requirements on mutually dependent extra-functional design aspects. The H2020 MSCAITN project RESCUE is focused on key challenges for reliability, security and quality, as well as related electronic design automation tools and methodologies. The objectives include both research advancements and cross-sectoral training of a new generation of interdisciplinary researchers. Notable interdisciplinary collaborative research results for the first halfperiod include novel approaches for test generation, soft-error and transient faults vulnerability analysis, cross-layer fault-tolerance and error-resilience, functional safety validation, reliability assessment and run-time management, HW security enhancement and initial implementation of these into holistic EDA tools.
Maksim Jenihhin, Said Hamdioui, Matteo Sonza Reorda, Milos Krstic, Peter Langendörfer, Christian Sauer 0001, Anton Klotz, Michael Hübner 0001, Jörg Nolte, Heinrich Theodor Vierhaus, Georgios N. Selimis, Dan Alexandrescu, Mottaqiallah Taouil, Geert Jan Schrijen, Jaan Raik, Luca Sterpone, Giovanni Squillero, Zoya Dyka
DATE4
2020 A Glitch-free Clock Multiplexer for Non-Continuously Running Clocks
abstract
Modern system-on-chips often integrate blocks, which need to be triggered by two or more clock sources depending on the circuit state. Glitchfree clock multiplexers are introduced to such systems for selecting the demanded clock. One specific problem of state-of-the-art solutions is that they need running clocks to perform the switching from one to another source. In this paper, a new clock multiplexer is presented, which overcomes this limitation and enables switching the clock, even if the clock stops running before switching has been performed. The applicability of the multiplexer has been proven in silicon by integrating it into a radhard 1.6-2.5 Gbps SERDES for space applications.
Steffen Zeidler 0001, Oliver Schrape, Anselm Breitenreiter, Milos Krstic
DSD4
2020 Design of Radiation Hardened RADFET Readout System for Space Applications
abstract
Measurement of absorbed dose and dose rate is a common task in radiation environments such as space. This is accomplished with the specialized instruments known as radiation dosimeters. Among the most commonly used radiation dosimeters in space missions are those based on the Radiation Sensitive Field Effect Transistors (RADFETs). In this paper, we propose a design concept for a radiation hardened readout system for the real-time measurement of absorbed dose and dose rate with RADFET. The successive switching between the absorbed dose and dose rate readout modes, as well as the subsequent data processing, are performed by the self-adaptive fault-tolerant Multiprocessing System-on-Chip (MPSoC). The integrated framework controller and the real-time monitoring of particle flux with the embedded Static Random Access Memory (SRAM) enable the autonomous selection of operating and fault-tolerant modes, thus achieving the optimal performance under variable radiation conditions.
Marko S. Andjelkovic, Aleksandar Simevski, Junchao Chen 0001, Oliver Schrape, Zoran Stamenkovic, Milos Krstic, Stefan D. Ilic, Luka Spahic, Laza Kostic, Goran S. Ristic, Aleksandar Jaksic, Alberto J. Palma, Antonio M. Lallena, Miguel Ángel Carvajal
DSD6
2020 Design Concept for Radiation-Hardening of Triple Modular Redundancy TSPC Flip-Flops
abstract
A robust design, which is one of the main requirements for space applications is always a tradeoff between power and area budget, speed requirement and the overall reliability. The occurrence of Single Event Effects (SEE) induced by energetic particle hits in the silicon leads to the insertion of additional replica logic at the design phase in order to tolerate Single Event Upsets (SEU) or Single Event Transients (SET). One of the most traditional circuit design technique is Triple Modular Redundancy (TMR). This paper presents a design concept for Radiation-Hardness-by-Design (RHBD) of TMR standard cell gates with local SET filter on the datapath, composed of True Single-Phase Clock (TSPC) flip-flops. The circuit architecture of the novel TSPC-ΔTMR flip-flops is discussed and compared to the baseline standard cell D-flip-flop. Analog simulations under various process, voltage, and temperature (PVT) conditions show an improvement by 50 % of the gate delay of the baseline TSPC flip-flop. Moreover, the proposed TSPC-ΔTMR gate candidate has a delay overhead of only 130ps under worst case condition compared to the classical unhardened D-latch-based reference flip-flop. Test vehicles for electrical measurements and radiation tests are implemented in 0.13 μm BiCMOS technology.
Oliver Schrape, Marko S. Andjelkovic, Anselm Breitenreiter, Alexey Balashov, Milos Krstic
DSD5
2020 Highly Configurable Framework for Adaptive Low Power and Error-Resilient System-On-Chip
abstract
In this paper, a novel, highly configurable framework for low power and error-resilient System-On-Chip is presented. The framework is composed, on the one hand, of the SWIELD configurable flip-flop and on the other hand, of the Chameleon controller. The SWIELD flip-flop is able to operate in three modes. It is driven/configured during runtime via the dedicated controller called Chameleon System Operation Management Unit. The proposed framework is integrated into a complex SoC based on a 32-bit general-purpose processor and the entire system is synthesized using the IHP 130 nm technology library. Numerous simulation experiments have been conducted in order to estimate the system error resilience and power consumption. At expense of negligible area and complexity overhead, the introduced framework shows great potential and excellent results w.r.t. both metrics of interest.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
DSD3
2020 PISA: Power-robust Multiprocessor Design for Space Applications
abstract
Recently the conservative space industry driven by the requirements of novel applications decided to introduce multiprocessor systems. Following the same line of motivation we introduce the PISA multiprocessor chip with improved power robustness for space applications which is successfully produced and tested in IHP 130 nm technology. The paper brings several novelties in respect to the current state-of-the-art. The chip uses the Waterbear framework in which the multiprocessor cores can be dynamically put in one of three different operating modes according to the current application requirements regarding performance, power consumption and fault tolerance. The chip has special power supply architecture with 13 power domains and Adaptive Voltage Scaling (AVS) mechanism based on voltage regulators which imposes a non-standard IC design flow. The measurement results showed that the power supply of the multiprocessor cores can be reduced from the nominal 1,2 V downto 0,82 V without compromising power integrity.
Aleksandar Simevski, Oliver Schrape, Carlos Benito, Milos Krstic, Marko S. Andjelkovic
IOLTS4
2020 Cross-Layer Hardware/Software Assessment of the Open-Source NVDLA Configurable Deep Learning Accelerator
abstract
The Nvidia Deep Learning Accelerator (NVDLA) is a free and open architecture that aims at promoting a standard way of designing deep neural network (DNN) inference engines. The analogy between open-source software and hardware points to FPGAs as ideal implementation platforms for open hardware accelerators. However, the instantiation flexibility enabled by reconfigurable logic should be correlated to the capacity of cost-effective devices. This paper explores the resource utilization-performance trade-offs spanned by the main precompiled NVDLA accelerator configurations on top of the mainstream Zynq UltraScale+ MPSoC. For the sake of comprehensive end-to-end performance characterization, the inference rate of the software stack is matched to that of the accelerator hardware, thus identifying current bottlenecks and promising optimization directions.
Alessandro Veronesi, Milos Krstic, Davide Bertozzi
VLSI-SOC2
2019 Design of SRAM-Based Low-Cost SEU Monitor for Self-Adaptive Multiprocessing Systems
abstract
Cosmic radiation phenomena such as Solar Particle Events cause high radiation flux lasting from hours to days, thus increasing the probability of Single-Event Upsets (SEUs) for several orders of magnitude. In space applications it is necessary, therefore, to monitor the SEU rate in order to ensure timely detection of high radiation levels and efficient protection of radiation-sensitive circuits. This work proposes an approach combining the SEU monitoring and data storage functions in the same on-chip Static Random Access Memory (SRAM) module, with negligible cost and overheads compared to traditional stand-alone SEU monitors. Furthermore, it also enables the detection of permanent faults in SRAM. The proposed monitor is intended to be further integrated into a highly dependable and self-adaptive multiprocessing platform in which it will drive the selection of the multiprocessor operating modes. Thus, a dynamic trade-off between reliability, performance and power consumption in real-time can be achieved.
Junchao Chen 0001, Marko S. Andjelkovic, Aleksandar Simevski, Patryk Skoncej, Milos Krstic
DSD6
2019 Aspects on Timing Modeling of Radiation-Hardness by Design Standard Cell-Based △TMR Flip-Flops
abstract
The paper presents and discusses the timing modeling approach for digital Radiation-Hardness by Design (RHBD) ΔTMR flip-flops. The basic fault-tolerant Triple Modular Redundancy (TMR) flip-flop architecture and the Single Event Transient-tolerant variant for the datapath (ΔTMR) are briefly introduced. The main focus is set on proper timing library modeling and timing check characterization respectively, as these are required in the digital design flow. The analyses are made with usage of a 130nm high performance BiCMOS technology node as a case study for the paper.
Oliver Schrape, Anselm Breitenreiter, Steffen Zeidler 0001, Milos Krstic
DSD4
2019 Characterization and Modeling of SET Generation Effects in CMOS Standard Logic Cells
abstract
Single event transients (SETs) stand out as one of the major causes of soft errors in nanoscale CMOS integrated circuits. To reduce the need for exhaustive circuit simulations in the design of radiation-hard integrated circuits, the cost-effective approaches for characterization and modeling of SET generation effects in standard logic cells are required. In this work, a SPICE-based methodology for characterization of SET generation effects, employing two different SET current models, is presented. Based on the acquired simulation results, the empirical models for the two main SET generation metrics (SET critical charge and SET pulse width) are derived. The SET generation models and the respective model parameters are intended to be used as inputs for the higher-level analysis of SET effects in digital circuits designed with the characterized standard logic cells. By storing the model parameters for each gate in the look-up table, instead of storing the raw data obtained from SPICE simulations, the amount of characterization data can be significantly reduced, allowing to speed up the subsequent SET analysis of a complex circuit.
Marko S. Andjelkovic, Zoran Stamenkovic, Milos Krstic, Rolf Kraemer
IOLTS4
2019 A Radiation Tolerant 10/100 Ethernet Transceiver for Space Applications
abstract
As space systems evolve to become more complex, they need larger computing and communication capabilities. For example, larger data rates must be supported and more flexible technologies that enable several applications to share the network resources while providing predictable and reliable performance are needed. One of the approaches to address those issues is the adoption of Ethernet standards in space. This has the benefit of reusing existing and field proven technology that also provides evolution to larger data rates. Ethernet is currently used in some space systems and it is being part of the implementation roadmap of many others, such as the next generation of Ariane launchers. Integrated circuits used in space systems must be designed to withstand the effects of radiation that causes errors and failures. These devices, known as rad-hard devices, need to be designed and manufactured using specific techniques and processes. In order to adopt the use of Ethernet in space, corresponding radhard Integrated Circuits (ICs) need to be available. The European industry is working on several such ICs, including an Ethernet switch and a physical layer transceiver. In this paper, SEPHY a 10/100 Mb/s rad-hard Ethernet transceiver designed for space applications is presented.
Anselm Breitenreiter, Jesús López, Pedro Reviriego, Milos Krstic, Úrsula Gutierro, Manuel Sanchez-Renedo, Daniel González
IOLTS4
2019 Selective Fault Tolerance by Counting Gates with Controlling Value
abstract
The protection of flip-flops against soft errors in digital circuits incurs significant overheads. To reduce the protection costs, it is common to identify the flip-flops in which errors can produce an effect on the system output or a persistent error in its state. Then, only those critical flip-flops are protected. To identify those flip-flops, one option is to perform fault injection on all the flips flops during functional simulations but this does not scale well for large circuits as the time required to perform an evaluation would not be practical. Another option is to perform the identification based only on the structural properties of the circuit. For example, flip-flops that are in loops are more likely to produce persistent errors and those closer to the system outputs to affect them. In this paper, an enhancement to the structural analysis is proposed to improve its accuracy. The idea is to identify gates that can mask error propagation as they have a controlling value and use only those to measure distances on the circuit to estimate the criticality of flip-flops. This enables us to better predict the effect of errors while keeping the analysis simple and based only on structural properties and gate types. The proposed scheme has been implemented and tested on a realistic circuit to show its effectiveness.
Anselm Breitenreiter, Stefan Weidling, Oliver Schrape, Steffen Zeidler 0001, Pedro Reviriego, Milos Krstic
IOLTS6
2018 Flip-Flop SEUs Mitigation through Partial Hardening of Internal Latch and Adjustment of Clock Duty Cycle
abstract
A radiation-hardness-by-design (RHBD) method for flip-flop single-event upsets (SEUs) mitigation is studied in this paper. This method applies a certain radiation hardened structure, e.g., the dual-interlocked storage cell (DICE), to implement one stage latch of a flip-flop while the SEUs protection for the other stage is realized by adjusting the clock duty cycle to shorten its hold state duration. Since the radiation hardening technique is used for only one stage latch, the overall area and power costs can be lowered. This technique is compatible with the automatic digital design flow and was implemented for an asynchronous first-in-first-out (FIFO) circuit as a case study in this paper.
Anselm Breitenreiter, Marko S. Andjelkovic, Oliver Schrape, Milos Krstic
DDECS5
2018 A Methodology to Verify Digital IP's within Mixed-Signal Systems
abstract
This paper describes a methodology to improve the quality of verification and an approach to dimension the arithmetic of register transfer level (RTL) model of the digital part of the mixed-signal system. This includes the refinement of the high level model of the system and generation of a MATLAB fixed-point model and test-bench for MATLAB-HDL cosimulation. Additionally an approach for the dimensioning of an adaptive equalizer in frequency domain is discussed. The proposed methodology and results of analysis are applied to verify 10BASE-T/100BASE-TX Ethernet PHY IP.
Navaneetha Channiganathota Manjappa, Anselm Breitenreiter, Markus Ulbricht 0002, Milos Krstic
DDECS4
2018 D-SET Mitigation Using Common Clock Tree Insertion Techniques for Triple-Clock TMR Flip-Flop
abstract
The paper presents a strategy for mitigation of SETs on a data path using three-clock input TMR (Triple Modular Redundancy) flip-flop cells (φTMR approach) as an alternative to the widely used TMR approach with local delay filtering (ΔTMR). The proposed flow enables the use of the common clock tree skewing techniques and is fully compatible with the standard digital design flow. The φTMR architecture is presented, and a shift-register test circuit is implemented in 130nm BiCMOS technology to compare the φTMR with the ΔTMR. Obtained results have shown that the φTMR approach results in power savings of 25% compared to ΔTMR. Furthermore, the robustness is increased by reducing the peak currents about 60% for the clock network, and 20% for the total current when φTMR is selected.
Oliver Schrape, Anselm Breitenreiter, Marko S. Andjelkovic, Milos Krstic
DSD4
2018 Interfacing 3D-stacked Electronic and Optical NoCs with Mixed CMOS-ECL Bridges: a Realistic Preliminary Assessment
abstract
The combination of optical networks-on-chip and 3D stacking represents the most promising system integration framework to overcome the communication bottleneck of future many-core processors. From an architecture viewpoint, the availability of an energy-efficient, low-latency bridge connecting the electronic network-on-chip with the optical one is as important as the maturity of the optical interconnect technology. The key design challenge consists of overcoming the inherent serial nature of optical communications, which is typically pursued by increasing either the data rate or the bit-level parallelism, or by a combination thereof. This paper explores an hybrid CMOS-ECL technology platform for bridge implementation by means of a complete logic synthesis effort. By spanning the wider configuration space of the hybrid bridge with respect to fully-CMOS realizations, the paper identifies the most energy-efficient configurations and provides a comparative assessment of achievable quality metrics. Derived results represent a solid and realistic starting point for future optimizations and for the refinement into an actual layout.
Mahdi Tala, Oliver Schrape, Milos Krstic, Davide Bertozzi
ACM Great Lakes Symposium on VLSI3
2018 Power/Area-Optimized Fault Tolerance for Safety Critical Applications
abstract
Increasing the reliability of a system always comes with a high price in performance/power/area overhead. Enabling error detection and correction features can be obtained by employing different kinds of redundancy including hardware, time, information, software or some combination of them. In many cases the imposed overhead is enormous. Fault tolerance is an important requirement for safety critical applications (e.g., automated driving), but significant power/area overhead is not acceptable. This paper summarizes several strategies and methods how to reduce the introduced overhead, while still providing a respectable level of fault tolerance features. Two main methodologies are discussed: static and dynamic. Static methods address the overhead by performing a static trade-off between the achieved level of fault tolerance and the introduced overhead. Dynamic methods on the other hand are based on the actual application requirement, and are dynamically varying the required overhead to fulfill the safety requirements of the application. This paper summarizes practical examples and results in this field.
Milos Krstic, Aleksandar Simevski, Markus Ulbricht 0002, Stefan Weidling
IOLTS1
2017 A Critical Charge Model for Estimating the SET and SEU Sensitivity: A Muller C-Element Case Study
abstract
This paper presents a critical charge model for estimating the SET and SEU robustness. The proposed model has been derived by analytic fitting of SPICE results, using a Muller C-element designed in 65 and 130 nm bulk CMOS technologies as the target device. The critical charge is expressed in terms of the size of C-element, size of load inverter, supply voltage and temperature, for constant timing parameters of the SET/SEU current pulse. The proposed model could be utilized to calculate the critical charge causing a SET, for both analyzed technologies, with the accuracy comparable to SPICE simulations. The critical charge for SEU was higher than for SET, but the dependencies obtained for SET response were qualitatively similar to those for SEU. This implies that the proposed critical charge model may be applicable for optimizing the SET/SEU robustness evaluation of the circuits involving the Muller C-element. Moreover, the model may also serve as a basis for evaluating the SET/SEU robustness of other standard cells and other technologies, and thus also for analysis of the SET/SEU robustness of complex circuits.
Marko S. Andjelkovic, Milos Krstic, Rolf Kraemer, Varadan Savulimedu Veeravalli, Andreas Steininger
ATS2
2017 An analysis of the operation and SET robustness of a CMOS pulse stretching circuit
abstract
The cascaded asymmetrically sized inverters can be employed as pulse stretchers, for the measurement of very short single event transient (SET) pulse widths (<; 200 ps). This paper analyzes, through the circuit simulations, the effects of various design and operating parameters on the normal operation and SET robustness of a two-inverter pulse stretcher designed in 250 nm bulk CMOS technology. It was shown that the SET hardness of the pulse stretcher can be enhanced by upsizing all transistors in the pulse stretcher without changing the sizing ratio. The SET hardness can also be improved by upsizing the load, but this approach is less effective than the pulse stretcher upsizing. Both upsizing approaches have a negligible impact on the normal operation of the stretcher, i.e. output pulse width. In addition, the operation and SET robustness of the pulse stretcher can be influenced by the operating temperature and supply voltage variations, and these effects should be considered in the design process. Based on the acquired simulation results, a general approach for the design of a radiation hardened CMOS pulse stretcher has been proposed.
Marko S. Andjelkovic, Milos Krstic, Rolf Kraemer
DDECS2
2017 Routing approach for digital, differential bipolar designs using virtual fat-wire boundary pins
abstract
This paper presents an alternative fat-wire routing approach for differential bipolar high-speed designs. The proposed solution obtains parallel routing and well balanced capacitive load of the fully differential signaling. In contrast to other approaches, the proposed flow is optimized for complex bipolar CML/ECL standard cell designs and technology options with few available routing layers. It enables the use of advanced placement and routing methods, such as multi-oriented cell placement and in-place optimization, supported by the standard CAD tools. The standard cell requirements and the corresponding modified digital design flow are proposed and discussed. The presented strategy is evaluated on a 12.5 GHz PLL feedback clock divider which has been fully implemented with differential ECL standard cell gates. A discussion regarding the obtained results finalizes this paper.
Oliver Schrape, Manuel Herrmann, Frank Winkler 0001, Milos Krstic
DDECS4
2017 Design of an On-chip System for the SET Pulse Width Measurement
abstract
This paper presents a design of an on-chip single event transient (SET) pulse width measurement system. The proposed system has been designed and implemented in IHP's 250 nm bulk CMOS technology and is intended for evaluation of SET effects in standard digital library cells. It is composed of an inverter-based target circuit, a pulse stretcher and a processing unit for counting the SET pulses and measuring the SET pulse width. The realized system is based on the combination of best practices from various existing designs, and it has a fairly simple architecture capable to provide reliable SET characterization. It supports serial interfacing with the external data acquisition unit which transfers the acquired data to the personal computer. The preliminary evaluation through the circuit-level simulations has demonstrated that the proposed design can detect and measure the SET pulse widths from around 100 ps up to 3.5 ns, with the measurement resolution of approximately 100 ps.
Marko S. Andjelkovic, Vladimir Petrovic, Miljana Nenadovic, Anselm Breitenreiter, Milos Krstic, Rolf Kraemer
DSD5
2017 Assessment of the amplitude-duration criterion for SET/SEU robustness evaluation
abstract
The relation between amplitude and duration of the current pulse induced by a high energy ionizing particle has been proposed as a criterion for evaluating the SET and SEU robustness of integrated circuits. This criterion has advantage over the widely accepted critical charge concept in the sense that it is less dependent on the current pulse shape. However, to the best of our knowledge, there is no known report on the impact of design and operating parameters on the amplitude-duration criterion. The need for extensive circuit or device simulations to derive the amplitude-duration curves under varying design and operating parameters makes this approach very time-consuming. In that regard, this work proposes a method to establish the amplitude-duration criterion, with a limited number of circuit simulations, as a rational function in terms of the sizing factors of target and load gates and supply voltage. Initial evaluation on a simple circuit composed of two inverters, designed in 130 nm bulk CMOS technology, has shown that the proposed method provides the accuracy comparable to SPICE simulations.
Marko S. Andjelkovic, Milos Krstic, Rolf Kraemer
IOLTS2
2016 Implementation of DBFN processor for Synthetic Aperture Radar application
abstract
One of the main reasons why Synthetic Aperture Radar (SAR) is an attractive solution for earth surface screening applications is its reliability independent on the weather conditions. This paper presents the implementation details of digital beamforming baseband core processor for such a SAR system. The processor chip is part of a distributed beamforming network (DBFN) of 16 baseband processors where each one processes data obtained from four 210MSPS ADC cores. Due to requirements related to space environment, the baseband is implemented in radiation-tolerant manner in order to temper single event effects (SEE). Instead of full-chip protection, the baseband processor uses only partially radhard flip-flops. This trade-off saves 25.87 % of silicon area based on gate-level synthesis results. A prototype is produced in a low-cost variant of a 0.25 μm BiCMOS process. First measurement results show an average operating current of 445.49 mA at a clock speed of 210 MHz and a 2.5 V power supply.
Oliver Schrape, Arkadiusz Koczor, Piotr Penkala, Vladimir Petrovic, Milos Krstic
DDECS5
2016 Implementation of a real time unit for satellite applications
abstract
The significance of low cost small satellites used for scientific research and practical applications continuously grows. Current satellite OBC (On-Board Computer) microcontrollers have integrated various digital peripherals and interfaces. However, a common Real Time Unit (RTU) requires interfacing to simple analogue sensors and actuators. Here we present a novel RTU microcontroller which includes a 13-bit Analog-to-Digital Converter (ADC) and two 12-bit Digital-to-Analog Converters (DAC). Furthermore, it includes a 32KB internal SRAM memory and a 32 KB internal flash memory. This enables an easy construction of a software-controlled embedded system which is easily interfaced to existing hardware sensors and actuators. The chip is produced in IHP 250 nm technology using radiation hardening by design. The operating frequency is 80 MHz. A 3-bit clock divider, as well as clock- and power-gating are used for reducing power consumption which is measured to be 0,8 W in operation.
Aleksandar Simevski, Klaus Schleisiek, Vladimir Petrovic, Norbert Beller, Patryk Skoncej, Günter Schoof, Milos Krstic
DDECS7
2015 A Coarse Model for Estimation of Switching Noise Coupling in Lightly Doped Substrates
abstract
The objective of this paper is to propose a coarse model for coupling of switching noise through lightly doped substrates. This could be achieved by assuming a regular placement of substrate contacts in a digital aggressor. Additionally, an approximation of equal ground bounce in an entire digital aggressor is applied. The proposed model is aimed for use as an estimation before placement, i.e. Before knowing the exact layout details. Consequently, this model could be utilized as a guideline for determining the optimal floor planning of digital blocks with regards to substrate noise coupling to sensitive analog modules. Extraction code is written in MATLAB. Evaluation of the model has shown that reasonable accuracy of the estimation could be expected and that the proposed method could be used as a baseline for early exploration of substrate noise characteristics of the design.
Milan Babic, Milos Krstic
DDECS2
2015 Design Flow for Radhard TMR Flip-Flops
abstract
Protection against radiation effects in digital ASICs chip can be achieved using different design approaches. One of the popular approaches for increasing the reliability is the hardware triplication. However, the hardware triplication does not mean that the susceptibility to radiation effects can be automatically overcome. The automatic random placement of standard cells can result in higher power consumption and more occupied silicon area), with marginal improvement of the linear energy transfer threshold (LET) value. Additionally, the TMR approach usually requires changes in the standard ASIC design flow, even requesting significant modifications of the RTL code. In this paper, we will investigate the issues of design flow for TMR flip-flops, addressing both the issues of compliance to the standard ASIC design methodology, enabling the use of non-modified RTL code, as well as the layout generation of radhard TMR flip-flops, based on standard non-hardened flip-flop components.
Vladimir Petrovic, Milos Krstic
DDECS2
2015 A Design Preconditioning Flow for Low-Noise Circuits
abstract
Mitigating switching noise in highly complex integrated circuits (ICs) is one of the challenging issues in current design flows. The common way to optimize the noise characteristics is to apply current shaping techniques, which introduce clock skew to distribute the switching activity of the circuit. However, this is typically done at late backend design stages, i.e., in layout after cell placement, which limits the maximum clock phase insertion between the domains. Therefore, we propose a novel preconditioning flow, which considers noise optimization up front end design, i.e., Design coding stage. By this, RTL-level techniques, such as clock inversion, can be applied to further optimize the noise characteristics, while common current shaping strategies can still be applied in the backend design.
Steffen Zeidler 0001, Xin Fan 0003, Oliver Schrape, Milos Krstic
DDECS4
2014 Improved circuitry for soft error correction in combinational logic in pipelined designs
abstract
This paper proposes the improvement of the method for soft error correction in the combinational circuit part of sequential circuits as described in [1], for which fault-tolerant master-slave elements are used. The errors in the combinational circuit part are detected by an error detection circuit and the error signal blocks the clock signal as long as the error exists. The system remains in its previous correct state and no complex recovery of the system is needed. In this paper we show how the method can be modified in such a way that it can be integrated into the standard industrial design flow and we demonstrate how the method can be favourably applied to a pipeline with different stages. Experimentally it is shown that the number of corrected errors can be increased by a factor of 4.22.
Milos Krstic, Stefan Weidling, Vladimir Petrovic, Michael Gössel
IOLTS1
2012 Functional Pattern Generation for Asynchronous Designs in a Test Processor Environment
abstract
Testing asynchronous circuits has been a challenge for several years. Especially, the nondeterministic timing behavior leads to problems during test, since the occurrence of test responses is not aligned to tester cycles. For this reason a test processor solution for asynchronous circuits has been recently provided, compensating the timing uncertainty. This is achieved by realizing an elastic test via asynchronous handshaking. Based on this approach we present a method for generating functional test patterns for the provided architecture.
Steffen Zeidler 0001, Christoph Wolf, Milos Krstic, Rolf Kraemer
Asian Test Symposium3
2012 Exploring pausible clocking based GALS design for 40-nm system integration
abstract
Globally asynchronous locally synchronous (GALS) design has attracted intensive research attention during the last decade. Among the existing GALS design solutions, the pausible clocking scheme presents an elegant solution to address the cross-clock synchronization issues with low hardware overhead. This work explored the applications of pausible clocking scheme for area/power efficient GALS design. To alleviate the challenge of timing convergence at the system level, area and power balanced system partitioning was applied for GALS design. An optimized GALS design flow based on the pausible clocking scheme was further proposed. As a practical example, a synchronous/GALS OFDM baseband transmitter chip, named Moonrake, was then designed and fabricated using the 40-nm CMOS process. It is shown that, compared to the synchronous baseline design, 5% reduction in area and 6% saving in power can be achieved in the GALS counterpart.
Xin Fan 0003, Milos Krstic, Eckhard Grass, Birgit Sanders, Christoph Heer
DATE2
2012 Asynchronous circuit design: From basics to practical applications
abstract
After motivating asynchronous techniques in general, the main advantages and disadvantages will be discussed. Several asynchronous timing models such as delay insensitive (DI) quasi delay insensitive (QDI) and speed independent (SI) will be presented and compared. Completion-detection methods used in asynchronous design are reviewed.
Eckhard Grass, Milos Krstic, Xin Fan 0003, Steffen Zeidler 0001
DDECS2
2012 Performance and complexity analysis of channel coding schemes for multi-Gbps wireless communications
abstract
In this paper, a trade-off analysis between concatenated codes consisting of a convolutional (CC) followed by a Reed-Solomon (RS) code versus low-density parity-check (LDPC) codes is presented. The analysis is based on a twofold criterion: coding gain for a target bit error rate (BER) of 10-6and required decoder hardware complexity for a target data throughput of 10 Gbps. Furthermore, we have investigated relevant parameters which directly impact an efficient hardware implementation as well as the error correction performance of the LDPC decoder. These parameters include an attenuation factor in the min-sum layered (MSL) decoding algorithm, the finite word length of soft information and an early-termination (ET) strategy. The error correction performances are evaluated for 16-QAM modulation over an independent Rayleigh fading channel. The complexity of the RS-CC and LDPC decoders is estimated based on synthesis results using an Infineon 40 nm CMOS design kit.
Miroslav Marinkovic, Milos Krstic, Eckhard Grass, Maxim Piz
PIMRC2
2011 Design of a Test Processor for Asynchronous Chip Test
abstract
Due to asynchronous timing and arbitration asynchronous designs may behave no deterministically. For the test of such systems, this means that an exact timing, i.e. a tester cycle, of a test response cannot be guaranteed. This behavior makes functional tests of asynchronous designs relatively complex or even impossible. Therefore, this paper presents a concept for performing functional tests of asynchronous designs using a test processor infrastructure. To this end, we propose a low-cost 16-bit microprocessor solution with special support of asynchronous handshake signalling that can either be integrated into the device-under-test (DUT), mounted on the load board of the tester or a combination of both.
Steffen Zeidler 0001, Christoph Wolf, Milos Krstic, Frank Vater, Rolf Kraemer
Asian Test Symposium3
2011 Low-complexity integrated circuit aging monitor
abstract
Integrated circuit aging effects are more and more pronounced with the continuous technological downscaling. These effects degrade circuit operation which is mainly observed as increased input-to-output delay of circuit components. Eventually, the circuit falls out of its specifications. Countermeasures are needed to prevent or reduce such degradation. Aging monitoring can be very beneficial since it can predict circuit failure and/or activate mechanisms to avoid failure. Most of the present aging monitors are based on reporting abnormal input-to-output signal delays on the critical path of the circuit. However, present approaches introduce additional circuit complexity, use complicated analog design, use non-standard cells etc. We propose a low-complexity aging monitor based on standard library cells, offering simplicity and flexibility of its design, integration and use. The designer could instantiate many monitors throughout the integrated circuit. The user can simply read the “aging code” placed in a register in each monitor and determine the “age” of the circuit, predict a circuit failure and/or take an appropriate action. This is especially useful in microprocessors which are designed with dependability in mind.
Aleksandar Simevski, Rolf Kraemer, Milos Krstic
DDECS3
2010 A GALS FFT processor with clock modulation for low-EMI applications
abstract
With the growth in complexity of digital CMOS circuits, the steep current fluctuations introduced by numerous transistors switching with clock signals are proven to be a significant source of electromagnetic interference (EMI). In recent years the reduction in EMI noise from high speed digital ICs has already gained intensive research attention. In this paper the pausible clocking based globally asynchronous locally synchronous (GALS) design with phase and frequency modulation on the locally generated clocks is proposed as a systematic solution to EMI reduction. As a practical example, a 64-point Radix-23pipelined GALS FFT processor was implemented using the IHP 130nm CMOS technology for low-EMI applications. The on-chip measurements demonstrate 13dB attenuation at the clock fundamental frequency and more than 20dB attenuation at higher clock harmonics, in comparison with the synchronous design.
Xin Fan 0003, Milos Krstic, Christoph Wolf, Eckhard Grass
ASAP2
2010 On-line testing of bundled-data asynchronous handshake protocols
abstract
Asynchronous interfaces, being a popular way of dealing with timing closure problems in deep submicron SoCs, pose a serious problem for on-line testing. Their behavior is specified not as a traditional clocked automaton, but as an asynchronous protocol with timing which depends on the clocks of the communicating blocks, wire and gate delays. The problem is exacerbated by the use of non-encoded bundled data buses, whose transitions cannot be reliably detected, as opposed to very expensive self-timed data encodings. This paper presents a further development of our low-cost checkers for asynchronous handshake protocol, which now support bundled data and are interfaced to the scan chains to assist diagnosis and debugging. Further reduction of area requirements for the delay elements, possibilities to define separate delays for each protocol phase, test results collection and checking the timing relationships between the handshake and the data signals are discussed. An application of the data transition detectors to checking of synchronous SDR/DDR protocols is shown.
Steffen Zeidler 0001, Alexandre V. Bystrov, Milos Krstic, Rolf Kraemer
IOLTS3
2009 Analysis and optimization of pausible clocking based GALS design
abstract
Pausible clocking based globally-asynchronous locally-synchronous (GALS) system design has been proven a promising approach to SoCs and NoCs. In this paper, we analyze the throughput reduction and synchronization failures introduced by the widely used pausible clocking scheme, and propose an optimized scheme for higher throughput and more reliable GALS design. The local clock generator is improved to minimize the acknowledge latency, and a novel input port is applied to maximize the safe timing region for the clock tree insertion. Simulation results using the IHP 0.13-µm standard CMOS process demonstrate that up to one-third increase in data throughput and an almost doubled safe timing region for clock tree distribution can be achieved in comparison to the traditional pausible clocking scheme.
Xin Fan 0003, Milos Krstic, Eckhard Grass
ICCD2
2009 Ultra low cost asynchronous handshake checker
abstract
This paper presents a new on-line checking scheme for asynchronous handshake protocols. The proposed scheme requires very small chip area while maintaining high coverage for all considered faults which are briefly exposed. In addition to simple pass-fail information the checker provides off-line diagnosis capabilities in order to further analyze the cause of a fault and the time of its occurrence. In order to verify its functionality the checker was proven by performing analogue simulations. In addition the area overhead and the power consumption was determined and compared with existing implementations.
Steffen Zeidler 0001, Marcus Ehrig, Milos Krstic, Michael Augustin, Christoph Wolf, Rolf Kraemer
IOLTS3
2009 An OFDM Baseband Receiver for Short-range Communication at 60 GHz
abstract
In this paper, we present implementation details of the digital front end of a wideband OFDM baseband receiver. This receiver offers datarates up to 1.08 GBit/s using a channel bandwidth of 400 MHz. The full baseband has been successfully implemented and tested on a FPGA platform running at a clock speed of only 100 MHz. It is part of a complete 60 GHz demonstrator including MAC, PHY and analog front end. In addition, we give performance results under realistic radio link assumptions for the 480 MBit/s mode to demonstrate the efficiency of the algorithms.
Maxim Piz, Milos Krstic, Marcus Ehrig, Eckhard Grass
ISCAS2
2008 60GHz OFDM hardware demonstrators in SiGe BiCMOS: State-of-the-art and future development
abstract
We present 60 GHz OFDM hardware demonstrators developed so far and outline design considerations for future developments. OFDM schemes have been developed to combat multi-path interferences in indoor wireless environments and further optimized for multi-gigabit data transmission in 60 GHz-band. RF and IF analogues front-ends (AFEs) have been developed with high-speed SiGe BiCMOS technologies. Future developments on AFE are mainly devoted to one-chip integration of RF and IF components based on a sliding IF architecture. We have already achieved wireless transmission of about 1 Gbps in an OFDM demonstrator, which will be continuously upgraded and optimized for higher data transmission with better link adaptability.
Chang-Soon Choi, Eckhard Grass, Frank Herzel, Maxim Piz, Klaus Schmalz, Yaoming Sun, Srdjan Glisic, Milos Krstic, Klaus Tittelbach-Helmrich, Marcus Ehrig, Wolfgang Winkler, Rolf Kraemer, Christoph Scheytt
PIMRC8
2007 60 GHz SiGe-BiCMOS Radio for OFDM Transmission
abstract
This paper reports implementation details of a 60 GHz short range communication system for data rates up to 2 Gbit/s. Based on a MATLAB model the main PHY parameters were derived, and the modules of the analog frontend were specified. For system verification, a demonstrator was developed. The analog RF and IF circuits of the demonstrator are implemented in a 0.25 μm SiGe BiCMOS technology. Measurement results of the analog modules and the complete demonstrator are given. Using OFDM modulation, the maximum data rate transmitted via airlink so far, is 960 Mbit/s.
Eckhard Grass, Frank Herzel, Maxim Piz, Klaus Schmalz, Yaoming Sun, Srdjan Glisic, Milos Krstic, Klaus Tittelbach-Helmrich, Marcus Ehrig, Wolfgang Winkler, Christoph Scheytt, Rolf Kraemer
ISCAS7
2007 Efficient Inner Receiver Design for OFDM-Based WLAN Systems: Algorithm and Architecture
abstract
In this article we propose a complete solution for the so-called inner receiver of an OFDM-WLAN system based on the IEEE 802.11a standard. We concentrate our investigations on three key components forming the inner receiver namely, the synchronizer, the channel estimator and the digital timing loop. The main goal is the joint optimization of the signal processing algorithms along with the implementation friendly VLSI architecture required for these three key components in order to reduce power, area and latency, without compromising the performance excessively. We provide both the mathematical details and extensive computer simulations to validate our design
Alfonso Troya, Koushik Maharatna, Milos Krstic, Eckhard Grass, Ulrich Jagdhold, Rolf Kraemer
IEEE Trans. Wirel. Commun.3
2005 BIST Technique for GALS Systems
abstract
In this paper a test technique based on the built-in self-test (BIST) is proposed. Our BIST concept is based on hierarchical testing of the digital systems. The presented test scheme is optimized for globally asynchronous locally synchronous (GALS) systems. The BIST technique, described here, is implemented on a GALS baseband processor compliant to the IEEE 802.11a standard. Some results on the performance of our test solution are given. The GALS processor with embedded BIST was fabricated in IHP's 0.25 /spl mu/m CMOS technology and test results are presented.
Milos Krstic, Eckhard Grass
DSD1
2005 Modified virtually scaling-free adaptive CORDIC rotator algorithm and architecture
abstract
In this paper, we proposed a novel Coordinate Rotation Digital Computer (CORDIC) rotator algorithm that converges to the final target angle by adaptively executing appropriate iteration steps while keeping the scale factor virtually constant and completely predictable. The new feature of our scheme is that, depending on the input angle, the scale factor can assume only two values, viz., 1 and 1//spl radic/2, and it is independent of the number of executed iterations, nature of iterations, and word length. In this algorithm, compared to the conventional CORDIC, a reduction of 50% iteration is achieved on an average without compromising the accuracy. The adaptive selection of the appropriate iteration step is predicted from the binary representation of the target angle, and no further arithmetic computation in the angle approximation datapath is required. The convergence range of the proposed CORDIC rotator is spanned over the entire coordinate space. The new CORDIC rotator requires 22% less adders and 53% less registers compared to that of the conventional CORDIC. The synthesized cell area of the proposed CORDIC rotator core is 0.7 mm/sup 2/ and its power dissipation is 7 mW in IHP in-house 0.25-/spl mu/m BiCMOS technology.
Koushik Maharatna, Swapna Banerjee, Eckhard Grass, Milos Krstic, Alfonso Troya
IEEE Trans. Circuits Syst. Video Technol.4
2004 A 16-bit CORDIC rotator for high-speed wireless LAN
abstract
We propose a novel 16-bit low power CORDIC rotator that is used for high-speed wireless LAN. The algorithm converges to the final target angle by adaptively selecting appropriate iteration steps while keeping the scale factor virtually constant. The VLSI architecture of the proposed design eliminates the entire arithmetic hardware in the angle approximation datapath and reduces the number of iterations by 50% on an average. The cell area of the processor is 0.7 mm and it dissipates 7 mW power at 20 MHz frequency.
Koushik Maharatna, Alfonso Troya, Swapna Banerjee, Eckhard Grass, Milos Krstic
PIMRC5
2003 Optimized low-power synchronizer design for the IEEE 802.11a standard
abstract
The authors propose a low-power synchronizer design for the IEEE 802.11a standard capable to estimate frequency offsets in the range /spl plusmn/468 kHz (80 ppm @ 5.8 GHz) with very simple and effective frame detection and timing synchronization. The core area of the design after layout is 13 mm/sup 2/, including the CORDIC and FFT processors, with a total estimated power consumption of 140 mW.
Milos Krstic, Alfonso Troya, Koushik Maharatna, Eckhard Grass
ICASSP (2)1