Georgios Karakonstantis

dblp:34/2281 · also George Karakonstantis · DBLP profile ↗
← Back
59ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-5693-8503ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 56 · 9 first-author · 18 since 2021Software engineering, systems software and programming languages · 18 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Partner Project: COIN-3D - Collaborative Innovation in 3D VLSI Reliability
abstract
As semiconductor manufacturing advances from the 3-nm process toward the sub-nanometer regime and transitions from FinFETs to gate-all-around field-effect transistors (GAAFETs), the resulting complexity and manufacturing challenges continue to increase. In this context, 3D chiplet-based approaches have emerged as key enablers to address these limitations while exploiting the expanded design space. Specifically, chiplets help address the lower yields typically associated with large monolithic designs. This paradigm enables the modular design of heterogeneous systems consisting of multiple chiplets (e.g., CPUs, GPUs, memory) fabricated using different technology nodes and processes. Consequently, it offers a capable and cost-effective strategy for designing heterogeneous systems.This paper introduces the Horizon Europe Twinning project COIN-3D (Collaborative Innovation in 3D VLSI Reliability), which aims to strengthen research excellence in 2.5D/3D VLSI systems reliability through collaboration between leading European institutions. More specifically, our primary scientific goal is the provision of novel open-source Electronic Design Automation (EDA) tools for reliability assessment of 3D systems, integrating advanced algorithms for physical- and system-level reliability analysis.
George Rafael Gourdoumanis, Fotoini Oikonomou, Maria Pantazi-Kypraiou, Pavlos Stoikos, Olympia Axelou, Athanasios Tziouvaras, Georgios Karakonstantis, Tahani Aladwani, Christos Anagnostopoulos 0001, Yixian Shen, Anuj Pathania, Alberto García Ortiz, George Floros 0002
DATE7
2026 On the Implementation of Low-Cost Neural-Network Models for Cardiac Arrhythmia Detection
Fotios Tsipis, Nikolaos Kostakis, Georgios Karakonstantis
ISCAS3
2026 Late Breaking Results - Compressed Bit-Level Timing-Error Predictors via Binary Neural Networks
Georgios Chatzitsompanis, Nikolaos Kostakis, Georgios Karakonstantis
VTS3
2024 ePredictNet: Low Cost Error Prediction Neural Network
abstract
The pursuit of miniature energy-efficient chips, and push for scaled voltages leads to increased timing errors that threaten the correct system functionality. Conventional error mitigation schemes based on redundancy, require costly and disruptive design changes, while they detect errors after they occur requiring expensive follow up correction schemes. In contrast to existing approaches, this paper introduces ePredictNet, an accurate workload-aware error-predictor based on compressed neural networks that can estimate early error prone instructions and avoid the manifestation of errors even under scaled voltages. Our work shows for the first time that accurate error prediction models are realizable on hardware with very low cost. This is achieved by training a quantized neural network that is converted to a netlist of truth tables using LogicNets, which can efficiently be mapped on circuit. Our results indicate, that ePredictNet once mapped on FPGA can achieve 99.39% error classification accuracy while utilizing only 270 LUTs and comes with only 2.1% area and 5% power overhead once integrated with a RISC core and costs up-to 98% less than the area and power incurred by redundancy based schemes. Apart from a cost effective error estimation scheme, ePredictNet can be used to avoid errors by guiding complimentary error mitigation schemes. For instance, in a power conscious use case,ePredictNet can allow the operation of a Open-RISC core under 12% reduced voltage, allowing to save 17% power with minimal 2.7% throughput reduction and guide the dynamic relaxation of frequency for avoiding errors only once error-prone instructions are predicted.
Georgios Chatzitsompanis, Georgios Karakonstantis
ISLPED2
2024 Adaptive approximate computing in edge AI and IoT applications: A review
abstract
Recent advancements in hardware and software systems have been driven by the deployment of emerging smart health and mobility applications. These developments have modernized the traditional approaches by replacing conventional computing systems with cyber-physical and intelligent systems combining the Internet of Things (IoT) with Edge Artificial Intelligence. Despite the many advantages and opportunities of these systems within various application domains, the scarcity of energy, extensive computing needs, and limited communication must be considered when orchestrating their deployment. Inducing savings in these directions is central to the Approximate Computing (AxC) paradigm, in which the accuracy of some operations is traded off with energy, latency, and/or communication reductions. Unfortunately, the dynamics of the environments in which AxC-equipped IoT systems operate have been paid little attention. We bridge this gap by surveying adaptive AxC techniques applied to three emerging application domains, namely autonomous driving, smart sensing and wearables, and positioning, paying special attention to hardware acceleration. We discuss the challenges of such applications, how adaptive AxC can aid their deployment, and which savings it can bring based on traits of the data and devices involved. Insights arising thereof may serve as inspiration to researchers, engineers, and students active within the considered domains.
Hans Jakob Damsgaard, Antoine Grenier, Dewant Katare, Zain Taufique, Salar Shakibhamedan, Tiago Troccoli, Georgios Chatzitsompanis, Anil Kanduri, Aleksandr Ometov, Aaron Yi Ding, Nima Taherinejad, Georgios Karakonstantis, Roger F. Woods, Jari Nurmi
J. Syst. Archit.12
2024 Enabling Voltage Over-Scaling in Multiplierless DSP Architectures via Algorithm-Hardware Co-Design
abstract
The design of low-power digital signal processing (DSP) architectures have gained a lot of attention due to their use in a variety of smart edge applications and portable devices. Recent efforts have focused on the replacement of power-hungry multipliers with various approximation frameworks such as multiplierless architectures that require only a few bit-shifts, additions and/or multiplexers when the multiplicand coefficients are known a priori. However, most existing multiplierless and approximation-based works have not been combined systematically with voltage over-scaling (VOS), which is considered one of the most effective power saving approaches, while the few that have tried, were applied to specific case studies with custom modifications. In this article, we are proposing a generic optimization framework that not only minimizes the hardware units in any time-multiplexed directed acyclic graph (TM-DAG) multiplier but also allows the reliable completion of most operations and the avoidance of random timing errors under VOS. This is achieved by synthesizing alternative coefficients that approximate well the original ones, while also activating shorter critical paths. As a result when VOS is applied, minor quality degradation occurs due to the coefficient approximations which are deterministic by design, while the gained timing slack of the new multiplicands allow us to reduce the supply voltage and circumvent the random timing errors induced by the increased delay under iso-frequency/throughput. Our experiments have indicated that when our framework is applied on fast Fourier transform (FFT) and discrete cosine transform (DCT) architectures, it results in up to 34.07% power savings, when compared to conventional multiplierless architectures, while it induces minimal signal-to-noise ratio (SNR) degradation, even when voltage is reduced by up to 20%.
Charalampos Eleftheriadis, Georgios Chatzitsompanis, Georgios Karakonstantis
IEEE Trans. Very Large Scale Integr. Syst.3
2023 AI-based Timing Error Modelling: A Case Study on a Pipelined Floating-point Core
abstract
The adoption of aggressively down-scaled voltages along with worsening process variations render nanometer devices prone to timing errors that threaten system functionality [1] , [2] . Recent studies tried to predict timing errors using machine learning (ML), while considering some workload characteristics [3] , [4] , [5] . However, successfully training such models is challenging, since traditionally acquired samples are insufficient, especially in operating regions where timing errors occur rarely.
Styliani Tompazi, Georgios Karakonstantis
ARITH2
2023 ACOR: On the Design of Energy-Efficient Autocorrelation for Emerging Edge Applications
abstract
The identification of patterns and changes in timeseries using the autocorrelation function (ACF) is traditionally used in several applications from communications, multimedia to remote health monitoring. Existing ACF implementations have tried to meet the throughput requirements of specific domains by mainly using time-domain approaches, however such techniques require several costly multiplications, which hinder their use in power-constrained devices, essential in emerging ACF-based edge applications. Frequency-domain (FD-ACF) approaches could reduce the computational complexity of the ACF calculation, but their use is limited in specific domains, leaving room for further power-aware algorithmic and architectural optimizations. This paper presents a framework, named ACOR, for the design of energy-efficient pipelined ACF architectures under various settings, throughput and energy requirements that vary across ACF-based applications. The proposed framework allows the quick exploration of ACF architectures for different sampling window sizes, window overlapping ratios, number of lags, and precision levels, which is impossible with the existing scattered domain-specific works. Our experimental results show that when compared with existing ACF architectures used in bio-signal analysis, linear predictive coding and telecommunications our proposed framework achieves up to 27.18%, and 51.47% reduction in the circuit area and energy consumption, respectively, with a slight throughput reduction of 8%.
Charalampos Eleftheriadis, Georgios Karakonstantis
ICCAD2
2023 A Compressed and Accurate Sparse Deep Learning-based Workload-Aware Timing Error Model
abstract
This paper showcases the novel application of Deep-Learning (DL) in the development of accurate microarchitecture and workload-aware timing error models and investigates methods such as sparsification for reducing their complexity, while maintaining high accuracy. Our study shows that DL can help increase the accuracy and true positive rate (TPR) of workload-aware models for a pipelined floating-point core compared to existing models. In addition, we demonstrate that removing up to 40% of the total neurons has minimal impact on the accuracy and overall predictive performance (up to 2.2%) of our DL-based timing error models, while significantly reducing the computational complexity. In fact, the complexity of the sparse model is approximately 2× smaller than the dense one.
Styliani Tompazi, Georgios Karakonstantis
ICCD2
2023 On the Facilitation of Voltage Over-Scaling and Minimization of Timing Errors in Floating-Point Multipliers
abstract
Voltage over-scaling (VoS) may be one of the most effective power reduction approaches, however, it makes circuits susceptible to timing failures. Various techniques were proposed to facilitate VoS by detecting and correcting errors and moving away from traditional voltage and timing guardbands. However, such approaches require the addition of extra redundant hardware leading to area and power overheads, especially in case of large timing errors. Recent, complementary approaches tried to redesign the target circuits, however, they were not yet applied on complex pipelined architectures like floating point multiplier which is extensively used in neural networks and other popular applications. In this paper, we develop a low-power pipelined floating-point IEEE-754 compatible multiplier that can operate reliably under voltage over-scaling (VoS). This is achieved by applying a path-shaping approach that helps minimize the number of paths susceptible to timing errors under lower voltages. In addition, our micro-architectural modifications isolate the critical paths only to a single pipeline stage, thus minimizing the error-prone stages under VoS. To showcase the efficacy of our approach, we perform post-place dynamic timing analysis using various benchmarks, indicating that our design can lead up to 171% better SNR, 80% less BER and 36% less power under 18% less voltage compared to the baseline multiplier. The applied analysis reveals that our approach can help limit the erroneous outputs of the unit by up to 74% and reduce by 20% the multi-bit error probability, in such a widely used arithmetic unit.
Georgios Chatzitsompanis, Georgios Karakonstantis
IOLTS2
2023 Microarchitecture-Aware Timing Error Prediction via Deep Neural Networks
abstract
Nanometer circuits are becoming increasingly prone to timing errors due to worsening parametric variations and operation close to voltage and frequency limits. Such errors threaten the system functionality and make circuits increasingly vulnerable to fault injection attacks, thus escalating the need to accurately predict and avoid them. Recent studies focus on modelling these errors by exploiting various supervised Machine Learning (ML)-based techniques. However, such efforts have not yet explored Neural Network (NN) methods that could improve accuracy, while being more easily scalable to complex, deep-pipelined architectures. This is the first study to explore the application of NN models on the accurate prediction of timing errors while considering various microarchitecture and workload parameters. To enable this study, we utilized stochastic search-based techniques to generate error-prone microarchitecture-aware samples, even in operating regions where samples are limited, the large number of which is an essential requirement in deep learning modelling. Our novel framework combines post-layout dynamic timing analysis and genetic algorithms, considering the data-dependent path sensitization and instruction execution history. The generated samples are used to train and evaluate various NN models for timing error prediction under multiple operating conditions. To evaluate the high efficacy of the NN models, we tested them on 6 applications with more than 8.5M instruction sequences. Evaluation results show over 99.8% predictive accuracy, combined with up to a 121.35% increase (on average) of the true positive rate in real test data compared to prior studies.
Styliani Tompazi, Georgios Karakonstantis
IOLTS2
2023 Energy-Efficient Short-Time Fourier Transform for Partial Window Overlapping
abstract
This paper presents an energy-efficient short-time Fourier transform (STFT) architecture. The proposed architecture is called frequency decomposition STFT (FD-STFT) and it achieves significant computational complexity reduction by effectively re-utilizing previously computed spectrums between overlapped sampling windows. Such an algorithmic modification not only reduces the required hardware units, but also achieves low accumulative error compared to conventional approaches. In addition, the quality of the resulting spectrogram is improved by integrating an efficient Hanning windowing technique that replaces the multiplication in the time domain with a low-cost filtering in the frequency domain. For an$N=256$-point window with$R=32$overlapping samples, our results indicate that our approach achieves up-to 40.86% and 65.56% area and power savings respectively, compared to recent approaches.
Charalampos Eleftheriadis, Mario Garrido, Georgios Karakonstantis
ISCAS3
2023 ARETE: Accurate Error Assessment via Machine Learning-Guided Dynamic-Timing Analysis
abstract
Nanometer circuits are increasingly prone to timing errors, escalating the need forfault injectionframeworks to accurately evaluate their impact on applications. In this paper, we propose ARETE, a novel cross-layer, fault-injection framework that combines dynamic-binary instrumentation with machine learning-guided dynamic-timing analysis. ARETE enables accurate fault-injection into any application by estimating the location of the injecting errors via dynamic-timing analysis. To accelerate fault-injection, we develop a novel, data-aware, machine learning-based mechanism that dynamically pre-selects the error-prone instructions and limits the application of the costly dynamic-timing analysis only to them. To evaluate ARETE's accuracy, our fully automated toolflow is configured to support fault-injection based on detailed post-layout gate-level simulations as well as via existing workload-agnostic error models. Our results for various workloads, including an autonomous-driving library, show that the location and time of injected errors performed by ARETE, is 89.9% consistent with fault-injection based on full gate-level simulation. On average, ARETE executes 84.6× faster than gate-level simulation and at a cost of 3.4% loss in the program output quality estimation. When compared to the existing statistical fault-injection tools that are based on workload-agnostic error models, ARETE improves the accuracy of fault-injection rate and output quality estimation by 143.9% and 40.4% on average, respectively.
Ioannis Tsiokanos, Styliani Tompazi, Giorgis Georgakoudis, Lev Mukhanov, Georgios Karakonstantis
IEEE Trans. Computers5
2023 Optimal Adder-Multiplexer Co-Optimization for Time-Multiplexed Multiplierless Architectures
abstract
Many digital signal processing (DSP) applications in multimedia, telecommunications, and artificial intelligence require several multiplications, which are considered among the most expensive arithmetic operations. To optimize these operations, several approaches have been proposed, mainly by representing multiplications with additions and bit-shifts. While these approaches may have limited the number of required adders, they have not given much attention to the overhead of multiplexers, which can grow significantly. In this paper, a comprehensive framework is presented that not only reduces the number of adders as in prior works, but also optimizes the number of multiplexers needed in modern DSP architectures. This framework is based on a new accelerated depth first search (A-DFS) algorithm that yields superior results, both for single input single output (SISO) and single input dual output (SIDO) architectures, the latter of which were not covered by existing approaches. At the initial stage of the proposed approach, all possible directed acyclic graph (DAG) multipliers are generated to produce minimum adder graphs of one or two outputs. Then, a systematic strategy to efficiently merge the produced graphs is presented, while preventing the number of multiplexers from growing exponentially as the set of multiplicands increases. Our experimental results show that when applied on several popular fast Fourier transform (FFT) and discrete cosine transform (DCT) coefficient sets, the proposed framework achieves significant savings in terms of the number of required multiplexers, leading to substantial area, power and power-delay-product (PDP) reduction, compared to existing works and a commercial synthesis tool.
Charalampos Eleftheriadis, Georgios Karakonstantis
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 Instruction-aware Learning-based Timing Error Models through Significance-driven Approximations
abstract
The adoption of aggressively down-scaled voltages along with worsening process variations, render nanometer devices prone to timing errors that threaten system functionality. The increased vulnerability of nanometer circuits to these errors attracted recent efforts in the development of timing error prediction models using machine learning (ML) methods. However, the majority of such models may be inaccurate, since they either neglect important microarchitecture properties and workload-dependent parameters, affecting timing error manifestation, or are constrained to limited operating areas. In this paper, we propose microarchitecture- and workload-aware ML models for timing error prediction that jointly consider various instruction types as well as all in-flight instructions in a pipeline. Our proposed models are able to predict the exact time and location (i.e., cycle, instruction and bit position) of timing errors with over 98% accuracy across multiple, critical operating regions. To circumvent the increased model complexity due to the considered features, we apply for the first time significance-driven approximations. Evaluation results for various workloads and voltage reduction levels show that our significance-driven precision scaling improves the models’ inference time up to 4.66×, with less than 4% accuracy loss. Finally, we use the proposed model to accurately and realistically inject timing errors during the evaluation of application resiliency. When compared to prior timing error evaluation frameworks that rely on workload-agnostic models, our framework improves the output quality estimation up to 82.6%.
Styliani Tompazi, Ioannis Tsiokanos, Jesús Martínez del Rincón, Lev Mukhanov, Georgios Karakonstantis
ICCD5
2022 Efficient, Dynamic Multi-Task Execution on FPGA-Based Computing Systems
abstract
With growing Field Programmable Gate Array (FPGA) device sizes and their integration in environments enabling sharing of computing resources such as cloud and edge computing, there is a requirement to share the FPGA area between multiple tasks. The resource sharing typically involves partitioning the FPGA space into fix-sized slots. This results in suboptimal resource utilisation and relatively poor performance, particularly as the number of tasks increase. Using OpenCL's exploration capabilities, we employ clever clustering and custom, task-specific partitioning and mapping to create a novel, area sharing methodology where task resource requirements are more effectively managed. Using models with varying resource/throughput profiles, we select the most appropriate distribution based on the runtime, workload needs to enhance temporal compute density. The approach is enabled in the system stack by a corresponding task-based virtualisation model. Using 11 high performance tasks from graph analysis, linear algebra and media streaming, we demonstrate an average 2.8× higher system throughput at 2.3× better energy efficiency over existing approaches.
Umar Ibrahim Minhas, Roger F. Woods, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IEEE Trans. Parallel Distributed Syst.4
2022 On the Evaluation of the Total-Cost-of-Ownership Trade-Offs in Edge vs Cloud Deployments: A Wireless-Denial-of-Service Case Study
abstract
We are witnessing an explosive growth in the number of Internet-connected devices and the emergence of several new classes of Internet of Things (IoT) applications that require rapid processing of an abundance of data. To overcome the resulting need for more network bandwidth and low network latency, a new paradigm has emerged that promotes the offering of Cloud services at the Edge, closer to users. However, the Edge is a highly constrained environment with limited power budget for servers per Edge installation which, in turn, limits the number of Internet-connected devices, such as sensors, that an installation can service. Consequently, the limited number of sensors leads to a reduction in the area coverage provided by them and puts in question the effectiveness for deploying IoT applications at the Edge. In this paper, we investigate the benefits of running an emerging security focused IoT application, (jamming detection), at the Edge vs. the Cloud by developing a Total Cost of Ownership (TCO) model, which considers the application's requirements as well as the Edge's constraints. For the first time, we build such a model based on realistic performance and energy-efficiency measurements obtained from commodity 64-bit ARM based micro-servers that are excellent candidates for supporting Cloud services at the Edge. Such servers represent the type of devices that can provide the right balance between power and performance, without requiring any complicate cooling and power supply infrastructure, which will not be available at the de-centralized deployments. Aiming at improving the energy efficiency, we exploit the pessimistic design margins adopted conventionally in such devices and investigate their operation under lower than nominal supply voltage and memory refresh-rate. Our results show that the jamming detection application deployed at an Edge environment is superior to a Cloud based solution by up to 2.13 times in terms of TCO. Moreover, when servers operate below nominal conditions, we can achieve up to 9 percent power savings which enables in several situations 100 percent gains in the TCO/area-coverage metric, i.e double area can be served with the same TCO.
Panagiota Nikolaou, Yiannakis Sazeides, Alejandro Lampropulos, Denis Guilhot, Andrea Bartoli, George Papadimitriou 0001, Athanasios Chatzidimitriou, Dimitris Gizopoulos, Konstantinos Tovletoglou, Lev Mukhanov, Georgios Karakonstantis
IEEE Trans. Sustain. Comput.11
2021 DTA-PUF: Dynamic Timing-aware Physical Unclonable Function for Resource-constrained Devices
abstract
In recent years, physical unclonable functions (PUFs) have gained a lot of attention as mechanisms for hardware-rooted device authentication. While the majority of the previously proposed PUFs derive entropy using dedicated circuitry, software PUFs achieve this from existing circuitry in a system. Such software-derived designs are highly desirable for low-power embedded systems as they require no hardware overhead. However, these software PUFs induce considerable processing overheads that hinder their adoption in resource-constrained devices. In this article, we propose DTA-PUF, a novel, software PUF design that exploits the instruction- and data-dependent dynamic timing behaviour of pipelined cores to provide a reliable challenge-response mechanism without requiring any extra hardware. DTA-PUF accepts sequences of instructions as an input challenge and produces an output response based on the manifested timing errors under specific over-clocked settings. To lower the required processing effort, we systematically select instruction sequences that maximise error-rate. The application to a post-layout pipelined floating-point unit, which is implemented in 45 nm process technology, demonstrates the effectiveness and practicability of our PUF design. Finally, DTA-PUF requires up to 50× fewer instructions than existing software processor PUF designs, limiting processing costs and resulting in up to 26% power savings.
Ioannis Tsiokanos, Jack Miskelly, Chongyan Gu, Máire O'Neill, Georgios Karakonstantis
ACM J. Emerg. Technol. Comput. Syst.5
2021 Revealing DRAM Operating GuardBands Through Workload-Aware Error Predictive Modeling
abstract
Improving the energy efficiency of DRAMs becomes very challenging due to the growing demand for storage capacity and failures induced by the manufacturing process. To protect against failures, vendors adopt conservative margins in the refresh period and supply voltage. Previously, it was shown that these margins are too pessimistic and will become impractical due to high-power costs, especially in future DRAM technologies. In this article, we present a new technique for automatic scaling the DRAM refresh period under reduced supply voltage that minimizes the probability of failures. The main idea behind the proposed approach is that DRAM error behavior is workload-dependent and can be predicted based on particular program inherent features. We use a Machine Learning (ML) method to build a workload-aware DRAM error behavior model based on the program features which we extract from real workloads during our DRAM error characterization campaign. With such a model, we identify the marginal value of the DRAM refresh period under relaxed voltage for each DRAM module of a server that enable us to reduce the DRAM power. We implement a temperature-driven OS governor which automatically sets the module-specific marginal DRAM parameters discovered by the ML model. Our governor reduces the DRAM power by 24 percent on average while minimizing the probability of failures. Unlike previous studies, our technique: i) does not require intrusive changes to hardware; ii) is implemented on a real server; iii) uses a mechanism that prevents any abnormal DRAM error behavior; and iv) can be easily deployed in data centers.
Lev Mukhanov, Konstantinos Tovletoglou, Hans Vandierendonck, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IEEE Trans. Computers5
2020 HaRMony: Heterogeneous-Reliability Memory and QoS-Aware Energy Management on Virtualized Servers
abstract
The explosive growth of data increases the storage needs, especially within servers, making DRAM responsible for more than 40% of the total system power. Such a reality has made researchers focus on energy saving schemes that relax the pessimistic DRAM circuit parameters at the cost of potential faults. In an effort to limit the resultant risk of critical data disruption, new methods were introduced that split DRAM into domains with varying reliability and power. The benefits of such schemes may have been showcased on simulators but have neither been implemented on real systems with a complete software stack, nor have been combined with any energy-reliability OS management policies. In this paper, we are the first to implement and evaluate HaRMony, a heterogeneous-reliability memory framework, in conjunction with QoS-aware energy management policies on a server with a complete virtualization stack. HaRMony overcomes the practical restrictions stemming from default hardware specifications, which were neglected in prior works, by introducing a software-based memory interleaving scheme. Furthermore, we expose the capabilities of HaRMony to the QEMU-KVM hypervisor through two unique policies. The first policy enables the hypervisor to seek the most power efficient DRAM circuit parameters based on the server availability requested by the user. The second policy enables users to exploit the inherent application error-resiliency by allowing them to limit the error protection mechanisms and allocate data structures on variably-reliable memory domains. Our evaluation shows that HaRMony reduces the performance overhead incurred due to disabling hardware interleaving from 29.3% down to 1.1% and leads to 17.7% DRAM energy savings and 8.6% total system energy savings on average in case of native execution of 28 benchmarks on an ARMv8-based server. Finally, we demonstrate that our QoS-aware scaling governor integrated with QEMU-KVM can dynamically scale the DRAM parameters, while reducing the system energy by 8.4% and meeting the targeted QoS even under extreme temperatures.
Konstantinos Tovletoglou, Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
ASPLOS4
2020 Increasing the Profit of Cloud Providers through DRAM Operation at Reduced Margins
abstract
Energy reduction is a key objective in cloud computing, and DRAM memories are responsible for an important amount of the energy consumption of data center nodes. Vendors adopt very conservative margins for DRAM operating parameters, such as the refresh rate and supply voltage, to guarantee correct operation even under the worst process variation and operating conditions. In this paper, we investigate the exploitation of DRAM margins to improve the energy efficiency of data center nodes, without triggering penalties due to service level agreement (SLA) violations. We introduce a model that captures the most important aspects of job management and system configuration. We also introduce RM-DRAM, a scheduling and node configuration policy that exploits the extended margins of DRAMs to reduce the operator's cost, considering the tradeoff between the cost of energy consumption and potential SLA violations. RM-DRAM also employs cost-aware (rather than threshold-based) VM consolidation. We extract the parameters used in the simulation (particularly power consumption and error rates) by characterizing a commercial ARM-based server. We perform simulations to evaluate the effectiveness of our approach, showing that significant gains, up to 34.84% and 29.53% in energy and cost, respectively, can be achieved compared with a state-of-the-art policy.
Christos Kalogirou, Christos D. Antonopoulos, Nikolaos Bellas, Spyros Lalis, Lev Mukhanov, Georgios Karakonstantis
CCGRID6
2020 DEFCON: Generating and Detecting Failure-prone Instruction Sequences via Stochastic Search
abstract
The increased variability and adopted low supply voltages render nanometer devices prone to timing failures, which threaten the functionality of digital circuits. Recent schemes focused on developing instruction-aware failure prediction models and adapting voltage/frequency to avoid errors while saving energy. However, such schemes may be inaccurate when applied to pipelined cores since they consider only the currently executed instruction and the preceding one, thereby neglecting the impact of all the concurrently executing instructions on failure occurrence. In this paper, we first demonstrate that the order and type of instructions in sequences with a length equal to the pipeline depth affect significantly the failure rate. To overcome the practically impossible evaluation of the impact of all possible sequences on failures, we present DEFCON, a fully automated framework that stochastically searches for the most failure-prone instruction sequences (ISQs). DEFCON generates such sequences by integrating a properly formulated genetic algorithm with accurate post-layout dynamic timing analysis, considering the data-dependent path sensitization and instruction execution history. The generated micro-architecture aware ISQs are then used by DEFCON to estimate the failure vulnerability of any application. To evaluate the efficacy of the proposed framework, we implement a pipelined floating-point unit and perform dynamic timing analysis based on input data that we extract from a variety of applications consisting of up-to 43.5M ISQs. Our results show that DEFCON reveals quickly ISQs that maximize the output quality loss and correctly detects 99.7% of the actual faulty ISQs in different applications under various levels of variation-induced delay increase. Finally, DEFCON enable us to identify failure-prone ISQs early at the design cycle, and save 26.8% of energy on average when combined with a clock stretching mechanism.
Ioannis Tsiokanos, Lev Mukhanov, Giorgis Georgakoudis, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
DATE5
2020 DStress: Automatic Synthesis of DRAM Reliability Stress Viruses using Genetic Algorithms
abstract
Failures become inevitable in DRAM devices, which is a major obstacle for scaling down the density of cells in future DRAM technologies. These failures can be detected by specific DRAM tests that implement the data and memory access patterns having a strong impact on DRAM reliability. However, the design of such tests is very challenging, especially for testing DRAM devices in operation, due to an extremely large number of possible cell-to-cell interference effects and combinations of patterns inducing these effects.In this paper, we present a new framework for the synthesis of DRAM reliability stress viruses, DStress. This framework automatically searches for the data and memory access patterns that induce the worst-case DRAM error behavior regardless the internal DRAM design. The search engine of our framework is based on Genetic Algorithms (GA) and a programming tool that we use to specify the patterns examined by GA. To evaluate the effect of program viruses on DRAM reliability, we integrate DStress with an experimental server where 72 DRAM chips can operate under various operating parameters and temperatures.We present the results of our 7-month experimental study on the search of DRAM reliability stress viruses. We show that DStress finds the worst-case data pattern virus and the worst-case memory access virus with probabilities of 1 - 4 × 10-7and 0.95, respectively. We demonstrate that the discovered patterns induce by at least 45% more errors than the traditional data pattern micro-benchmarks used in previous studies. We show that DStress enables us to detect the marginal DRAM operating parameters reducing the DRAM power by 17.7 % on average without compromising reliability. Overall, our framework facilitates the exploration of new data patterns and memory access scenarios increasing the probability of DRAM errors, which is essential for improving the state-of-the-art DRAM testing mechanisms.
Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
MICRO3
2019 Low-Power Variation-Aware Cores based on Dynamic Data-Dependent Bitwidth Truncation
abstract
Increasing variability of transistor parameters in nanoscale era renders modern circuits prone to timing failures. To address such failures, designers adopt pessimistic timing/voltage guardbands, which are estimated under rare worst-case conditions, thus leading to power and performance overheads. Recent approximation schemes based on precision reduction may help to limit the incurred overheads, but the precision is reduced statically in all operations. This results in unnecessary quality loss, since these schemes neglect the fact that only few long latency paths (LLPs) may be prone to failures, and such paths may be activated rarely. In this paper, we propose a variation-aware framework that minimizes any quality loss by dynamically truncating the bitwidth only for operands triggering the LLPs. This is achieved by predicting at runtime the excitation of the LLPs based on the processed operands. The applied truncation, which we implement by setting a number of least-significant bits to a constant value of zero, can effectively reduce the delay of the excited LLPs, providing sufficient timing slack to avoid failures without using conservative guardbands. To facilitate the adoption of such a scheme within pipelined cores and limit the incurred overheads, we also shape the path distribution appropriately for isolating the LLPs in a single pipeline stage. Additionally, to evaluate the efficacy of our framework, we perform post-layout dynamic timing analysis based on real operands that we extract from a variety of applications. When applied to the implementation of an IEEE-754 compatible double precision floating-point unit (FPU) in a 45nm technology, our approach eliminates timing failures under 8% delay variations with no performance loss. Our design comes at a cost of up-to 4.48% power and 0.34% area overheads, while the occasional operand truncation incurs minimal quality-loss in terms of relative error, up-to 4.1 · 10-6. Finally, when compared to an FPU with pessimistic margins, our techniqutechniquee can save up-to 44.3% power.
Ioannis Tsiokanos, Lev Mukhanov, Georgios Karakonstantis
DATE3
2018 An energy-efficient and error-resilient server ecosystem exceeding conservative scaling limits
abstract
The explosive growth of Internet-connected devices will soon result in a flood of generated data, which will increase the demand for network bandwidth as well as compute power to process the generated data. Consequently, there is a need for more energy efficient servers to empower traditional centralized Cloud data-centers as well as emerging decentralized data-centers at the Edges of the Cloud. In this paper, we present our approach, which aims at developing a new class of micro-servers - the UniServer - that exceed the conservative energy and performance scaling boundaries by introducing novel mechanisms at all layers of the design stack. The main idea lies on the realization of the intrinsic hardware heterogeneity and the development of mechanisms that will automatically expose the unique varying capabilities of each hardware. Low overhead schemes are employed to monitor and predict the hardware behavior and report it to the system software. The system software including a virtualization and resource management layer is responsible for optimizing the system operation in terms of energy or performance, while guaranteeing non-disruptive operation under the extended operating points. Our characterization results on a 64-bit ARMv8 micro-server in 28nm process reveal large voltage margins in terms of Vmin variation among the 8 cores of the CPU chip, among three different sigma chips, and among different benchmarks with the potential to obtain up-to 38.8% energy savings. Similarly, DRAM characterizations show that refresh rate and voltage can be relaxed by 35x and 5%, respectively, leading to 23.2% power savings on average.
Georgios Karakonstantis, Konstantinos Tovletoglou, Lev Mukhanov, Hans Vandierendonck, Dimitrios S. Nikolopoulos, Peter Lawthers, Panos K. Koutsovasilis, Manolis Maroudas, Christos D. Antonopoulos, Christos Kalogirou, Nikolaos Bellas, Spyros Lalis, Srikumar Venugopal, Arnau Prat-Pérez, Alejandro Lampropulos, Marios Kleanthous, Andreas Diavastos, Zacharias Hadjilambrou, Panagiota Nikolaou, Yiannakis Sazeides, Pedro Trancoso, George Papadimitriou 0001, Manolis Kaliorakis, Athanasios Chatzidimitriou, Dimitris Gizopoulos, Shidhartha Das
DATE1
2018 Facilitating Easier Access to FPGAs in the Heterogeneous Cloud Ecosystems
abstract
With FPGAs being increasingly integrated into existing software-based heterogeneous cloud environments, novel evaluation mechanisms are required to reveal the energy-performance trade-offs of accelerators (FPGAs, GPUs, etc) in high-level heterogeneous programming environments. For FPGAs, this involves also a reconsideration of scheduling policies and reconfiguration methods with an aim of integrating software-based approaches as well as performance optimizations for wider workload sizes. The approaches are evaluated using various reconfiguration methodologies for a number of applications.
Umar Ibrahim Minhas, Roger F. Woods, Georgios Karakonstantis
FPL3
2018 Userspace Hypervisor Data Characterization in Virtualized Environment
abstract
Memory-intensive applications have grown rapidly in order to adapt to the changing needs of businesses. Memory errors become even more concerning, since configuration at extended operating points makes hardware more susceptible to system failures. Some data structures may be more sensitive to errors and may cause system crashes more easily than others. A failure in a critical data structure can cause a complete system crash. In this paper, we propose a sensitivity characterization to the hypervisor structures. First, we implement an error-injection framework using Syscall ptrace. We profile both static and dynamic data structures through our framework. Then we provide the detailed analysis and data characterization to QEMU hypervisor. Finally, we discuss the checkpointing and selective mechanism in case the QEMU hypervisor crash.
Hans Vandierendonck, Georgios Karakonstantis, Dimitrios S. Nikolopoulos
ICPADS3
2018 DRAM Characterization under Relaxed Refresh Period Considering System Level Effects within a Commodity Server
abstract
Today's rapid generation of data and the increased need for higher memory capacity has triggered a lot of studies on aggressive scaling of refresh period, which is currently set according to rare worst case conditions. Such studies analysed in detail the data-dependent circuit level factors and indicated the need for online DRAM characterization due to the variable cell retention time. They have done so by executing few test data patterns on FPGAs under controlled temperatures by using thermal testbeds, which however cannot be available in the field. Moreover, the existing studies were not able to reveal any system level effects, which may be excited under the execution of workloads on real systems and directly or indirectly affect DRAM reliability. In this paper, we develop an experimental framework based on a state-of-the-art 64-bit ARM based server with Linux OS, in which we enabled the DRAM characterization under relaxed refresh period by executing conventional test data patterns as well as popular HPC and Cloud workloads. Our results indicate that common test patterns are ineffective in identifying error-prone locations at low DRAM temperatures. Furthermore, we reveal that there is a strong correlation between the SOC utilization and DRAM reliability. By exploiting such findings, we developed a benchmark, which can indirectly stress the DRAM temperature and thus used for characterization in the field without needing any complicated thermal equipment. Our study shows that the refresh period can be relaxed by 35 times on such a commodity system with all errors being corrected by the available error correcting codes, resulting in 11.5% power savings on average.
Lev Mukhanov, Konstantinos Tovletoglou, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IOLTS4
2018 Minimization of Timing Failures in Pipelined Designs via Path Shaping and Operand Truncation
abstract
The continuous scaling of transistor sizes and the increased static and dynamic parametric variations render nanometer circuits more prone to timing failures. To protect circuits from such failures, typically designers adopt pessimistic timing margins, which are estimated statically under rare worst-case conditions. In this paper, we aim at minimizing the timing failures, while avoiding such pessimistic margins by proposing an approach that initially minimizes the number of long latency paths within each processor pipeline stage and con- straints them in as few stages as possible. Such an approach, not only reduces the timing failures, but also limits the potential error prone locations to only few pipeline registers/stages. To further reduce these failures, we exploit the path excitation dependence on data patterns and we truncate the bit-width of the operands in the few remaining long latency paths by setting a number of least significant bits to a constant value zero. Such truncation may incur quality loss, but this can be controlled by carefully selecting the number of truncated bits and will be in any case less than the catastrophic loss that may be incurred under random timing failures. Additionally, our framework performs post- place and route dynamic timing analysis based on real operands that are extracted from a variety of applications, helping to estimate the dynamic timing failures, while considering the data dependent path excitation. When applied to an IEEE-754 compatible double precision Floating Point Unit (FPU), the proposed approach reduces the timing failures by 104.5× on average compared to a reference FPU design under an assumed 8.1% variation-induced worst-case path delay increase in a 45 nm process. Finally, results show that path shaping alone introduces an insignificant 0.25% area and 5.7% power overhead with no performance cost. The combination of path shaping with aggressive operand bit- width truncation leads to up-to 44.7% on average power savings due to the substantially reduced switching activity at a minimal quality loss.
Ioannis Tsiokanos, Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IOLTS4
2018 Variation-Aware Pipelined Cores through Path Shaping and Dynamic Cycle Adjustment: Case Study on a Floating-Point Unit
abstract
In this paper, we propose a framework for minimizing variation-induced timing failures in pipelined designs, while limiting any overhead incurred by conventional guardband based schemes. Our approach initially limits the long latency paths (LLPs) and isolates them in as few pipeline stages as possible by shaping the path distribution. Such a strategy, facilitates the adoption of a special unit that predicts the excitation of the isolated LLPs and dynamically allows an extra cycle for the completion of only these error-prone paths. Moreover, our framework performs post-layout dynamic timing analysis based on real operands that we extract from a variety of applications. This allows us to estimate the bit error rates under potential delay variations, while considering the dynamic data dependent path excitation. When applied to the implementation of an IEEE-754 compatible double precision floating-point unit (FPU) in a 45nm process technology, the path shaping helps to reduce the bit error rates on average by 2.71 x compared to the reference design under 8% delay variations. The integrated LLPs prediction unit and the dynamic cycle adjustment avoid such failures and any quality loss at a cost of up-to 0.61% throughput and 0.3% area overheads, while saving 37.95% power on average compared to an FPU with pessimistic margins.
Ioannis Tsiokanos, Lev Mukhanov, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
ISLPED4
2017 Relaxing DRAM refresh rate through access pattern scheduling: A case study on stencil-based algorithms
abstract
The main memory in today's systems is based on DRAMs, which may offer low cost and high density storage for large amounts of data but it comes with a main drawback; DRAM cells need to be refreshed frequently for retaining the stored data. The refresh rate in modern DRAMs is set based on the worst-case retention time without considering access statistics, thereby resulting in very frequent refresh operations. Such high refresh rate leads eventually to large power and performance overheads, which are increasing with higher DRAM densities. However, such high refresh rates may not even required due to extremely low probability of the actual occurrence of the assumed worst-case scenarios, or due to the implicit refresh operation that occur during every memory access, a feature that has not been yet been studied in depth. In this paper, we enhance the state-of-the-art by systematically exploiting the implicit refresh of memory access for relaxing the refresh rate, while minimizing the resulting memory errors. This is achieved by modifying the algorithmic parameters that influence the access patterns such that all stored data are being touched within a target time interval that is necessary for meeting a target error rate. The proposed method is applied to stencil-based algorithms which represent a wide class of algorithms used in numerical analysis, image processing and cellular automata applications. The efficacy of the proposed method is demonstrated on an off-the-shelf server running a fully fledged Linux OS and results show that it is even possible to completely disable DRAM refresh with minor quality loss.
Konstantinos Tovletoglou, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IOLTS3
2016 Statistical fault injection for impact-evaluation of timing errors on application performance
abstract
This paper proposes a novel approach to modeling of gate level timing errors during high-level instruction set simulation. In contrast to conventional, purely random fault injection, our physically motivated approach directly relates to the underlying circuit structure, hence allowing for a significantly more detailed characterization of application performance under scaled frequency / voltage (including supply noise). The model uses gate level timing statistics extracted by dynamic timing analysis from the post place & route netlist of a general-purpose processor to perform instruction-aware fault injections. We employ a 28 nm OpenRISC core as a case study, to demonstrate how statistical fault injection provides a more accurate and realistic analysis of power vs. error performance.
Jeremy Constantin, Andreas Peter Burg, Zheng Wang 0020, Anupam Chattopadhyay, Georgios Karakonstantis
DAC5
2016 A low overhead error confinement method based on application statistical characteristics
Zheng Wang 0020, Georgios Karakonstantis, Anupam Chattopadhyay
DATE2
2016 Brief Announcement: Energy Optimization of Memory Intensive Parallel Workloads
abstract
Energy consumption is an important concern in modern multicore processors. The energy consumed during the execution of an application can be minimized by tuning the hardware state utilizing knobs such as frequency, voltage etc. The existing theoretical work on energy minimization using Global DVFS (Dynamic Voltage and Frequency Scaling), despite being thorough, ignores the energy consumed by the CPU on memory accesses and the dynamic energy consumed by the idle cores. This article presents an analytical energy-performance model for parallel workloads that accounts for the energy consumed by the CPU chip on memory accesses in addition to the energy consumed on CPU instructions. In addition, the model we present also accounts for the dynamic energy consumed by the idle cores. We present an analytical framework around our energy-performance model to predict the operating frequencies for global DVFS that minimize the overall CPU energy consumption. We show how the optimal frequencies in our model differ from the optimal frequencies in a model that does not account for memory accesses.
Chhaya Trehan, Hans Vandierendonck, Georgios Karakonstantis, Dimitrios S. Nikolopoulos
SPAA3
2015 Mitigating the impact of faults in unreliable memories for error-resilient applications
abstract
Inherently error-resilient applications in areas such as signal processing, machine learning and data analytics provide opportunities for relaxing reliability requirements, and thereby reducing the overhead incurred by conventional error correction schemes. In this paper, we exploit the tolerable imprecision of such applications by designing an energy-efficient fault-mitigation scheme for unreliable data memories to meet target yield. The proposed approach uses a bit-shuffling mechanism to isolate faults into bit locations with lower significance. This skews the bit-error distribution towards the low order bits, substantially limiting the output error magnitude. By controlling the granularity of the shuffling, the proposed technique enables trading-off quality for power, area, and timing overhead. Compared to error-correction codes, this can reduce the overhead by as much as 83% in read power, 77% in read access time, and 89% in area, when applied to various data mining applications in 28nm process technology.
Shrikanth Ganapathy, Georgios Karakonstantis, Adam Teman, Andreas Peter Burg
DAC2
2015 On the statistical memory architecture exploration and optimization
Charalampos Antoniadis, Georgios Karakonstantis, Nestoras E. Evmorfopoulos, Andreas Peter Burg, Georgios I. Stamoulis
DATE2
2015 Exploiting dynamic timing margins in microprocessors for frequency-over-scaling with instruction-based clock adjustment
Jeremy Constantin, Lai Wang, Georgios Karakonstantis, Anupam Chattopadhyay, Andreas Peter Burg
DATE3
2015 Energy versus data integrity trade-offs in embedded high-density logic compatible dynamic memories
Adam Teman, Georgios Karakonstantis, Robert Giterman, Pascal Andreas Meinerzhagen, Andreas Peter Burg
DATE2
2015 Refresh-free dynamic standard-cell based memories: Application to a QC-LDPC decoder
abstract
The area and power consumption of low-density parity check (LDPC) decoders are typically dominated by embedded memories. To alleviate such high memory costs, this paper exploits the fact that all internal memories of a LDPC decoder are frequently updated with new data. These unique memory access statistics are taken advantage of by replacing all static standard-cell based memories (SCMs) of a prior-art LDPC decoder implementation by dynamic SCMs (D-SCMs), which are designed to retain data just long enough to guarantee reliable operation. The use of D-SCMs leads to a 44% reduction in silicon area of the LDPC decoder compared to the use of static SCMs. The low-power LDPC decoder architecture with refresh-free D-SCMs was implemented in a 90nm CMOS process, and silicon measurements show full functionality and an information bit throughput of up to 600 Mbps (as required by the IEEE 802.11n standard).
Pascal Andreas Meinerzhagen, Andrea Bonetti, Georgios Karakonstantis, Christoph Roth, Frank Giirkaynak, Andreas Peter Burg
ISCAS3
2015 Enhancing Design Space Exploration by Extending CPU/GPU Specifications onto FPGAs
abstract
The design cycle for complex special-purpose computing systems is extremely costly and time-consuming. It involves a multiparametric design space exploration for optimization, followed by design verification. Designers of special purpose VLSI implementations often need to explore parameters, such as optimal bitwidth and data representation, through time-consuming Monte Carlo simulations. A prominent example of this simulation-based exploration process is the design of decoders for error correcting systems, such as the Low-Density Parity-Check (LDPC) codes adopted by modern communication standards, which involves thousands of Monte Carlo runs for each design point. Currently, high-performance computing offers a wide set of acceleration options that range from multicore CPUs to Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). The exploitation of diverse target architectures is typically associated with developing multiple code versions, often using distinct programming paradigms. In this context, we evaluate the concept of retargeting a single OpenCL program to multiple platforms, thereby significantly reducing design time. A single OpenCL-based parallel kernel is used without modifications or code tuning on multicore CPUs, GPUs, and FPGAs. We use SOpenCL (Silicon to OpenCL), a tool that automatically converts OpenCL kernels to RTL in order to introduce FPGAs as a potential platform to efficiently execute simulations coded in OpenCL. We use LDPC decoding simulations as a case study. Experimental results were obtained by testing a variety of regular and irregular LDPC codes that range from short/medium (e.g., 8,000 bit) to long length (e.g., 64,800 bit) DVB-S2 codes. We observe that, depending on the design parameters to be simulated, on the dimension and phase of the design, the GPU or FPGA may suit different purposes more conveniently, thus providing different acceleration factors over conventional multicore CPUs.
Muhsen Owaida, Gabriel Falcão Paiva Fernandes, João Andrade, Christos D. Antonopoulos, Nikolaos Bellas, Madhura Purnaprajna, David Novo, Georgios Karakonstantis, Andreas Peter Burg, Paolo Ienne
ACM Trans. Embed. Comput. Syst.8
2014 A quality-scalable and energy-efficient approach for spectral analysis of heart rate variability
abstract
Today there is a growing interest in the integration of health monitoring applications in portable devices necessitating the development of methods that improve the energy efficiency of such systems. In this paper, we present a systematic approach that enables energy-quality trade-offs in spectral analysis systems for bio-signals, which are useful in monitoring various health conditions as those associated with the heart-rate. To enable such trade-offs, the processed signals are expressed initially in a basis in which significant components that carry most of the relevant information can be easily distinguished from the parts that influence the output to a lesser extent. Such a classification allows the pruning of operations associated with the less significant signal components leading to power savings with minor quality loss since only less useful parts are pruned under the given requirements. To exploit the attributes of the modified spectral analysis system, thresholding rules are determined and adopted at design- and run-time, allowing the static or dynamic pruning of less-useful operations based on the accuracy and energy requirements. The proposed algorithm is implemented on a typical sensor node simulator and results show up-to 82% energy savings when static pruning is combined with voltage and frequency scaling, compared to the conventional algorithm in which such trade-offs were not available. In addition, experiments with numerous cardiac samples of various patients show that such energy savings come with a 4.9% average accuracy loss, which does not affect the system detection capability of sinus-arrhythmia which was used as a test case.
Georgios Karakonstantis, Aviinaash Sankaranarayanan, Mohamed M. Sabry, David Atienza 0001, Andreas Peter Burg
DATE1
2014 Enabling complexity-performance trade-offs for successive cancellation decoding of polar codes
abstract
Polar codes are one of the most recent advancements in coding theory and they have attracted significant interest. While they are provably capacity achieving over various channels, they have seen limited practical applications. Unfortunately, the successive nature of successive cancellation based decoders hinders fine-grained adaptation of the decoding complexity to design constraints and operating conditions. In this paper, we propose a systematic method for enabling complexity-performance tradeoffs by constructing polar codes based on an optimization problem which minimizes the complexity under a suitably defined mutual information based performance constraint. Moreover, a low-complexity greedy algorithm is proposed in order to solve the optimization problem efficiently for very large code lengths.
Alexios Balatsoukas-Stimming, Georgios Karakonstantis, Andreas Peter Burg
ISIT2
2013 Exploiting application resiliency for energy-efficient and adequately-reliable operation
abstract
Summary form only given. Currently, manufacturers go to great lengths for mitigating the effects of parametric variations and scaled supply voltages by adopting conservative layout rules and by introducing several mechanisms providing redundancy on various layers of design abstraction. Such measures may have accomplished to hide any inaccurate behavior of nanometer circuits from the application layers and maintain acceptable yield levels, but unfortunately the large energy, performance, and area overheads that they incur limit their viability, especially as we move beyond the 45nm node. Such a reality has urged us to rethink the current design flows and question if such considerable overhead is really required given that many modern signal-processing workloads such as in multimedia, communications, or biomedical systems are inherently complexity/energy-scalable and can even tolerate a degree of imprecision in their computations and stored data. By taking advantage of this inherent resilience of many applications and trading off output precision and quality of service we could reduce energy usage and reliability costs, since by allowing some computations to be approximate we can alleviate the burden of correctness overhead imposed by the traditional design paradigm. In this paper, we discuss methods that could reveal and exploit the resilience and certain characteristics of biomedical and communication applications for achieving adequately-reliable and energy efficient operation, while limiting or even avoiding the penalties required by traditional approaches.
Georgios Karakonstantis, David Atienza 0001, Andy Burg
IOLTS1
2012 On the exploitation of the inherent error resilience of wireless systems under unreliable silicon
abstract
In this paper, we investigate the impact of circuit misbehavior due to parametric variations and voltage scaling on the performance of wireless communication systems. Our study reveals the inherent error resilience of such systems and argues that sufficiently reliable operation can be maintained even in the presence of unreliable circuits and manufacturing defects. We further show how selective application of more robust circuit design techniques is sufficient to deal with high defect rates at low overhead and improve energy efficiency with negligible system performance degradation.
Georgios Karakonstantis, Christoph Roth, Christian Benkeser, Andreas Peter Burg
DAC1
2012 Shortening Design Time through Multiplatform Simulations with a Portable OpenCL Golden-model: The LDPC Decoder Case
abstract
Hardware designers and engineers typically need to explore a multi-parametric design space in order to find the best configuration for their designs using simulations that can take weeks to months to complete. For example, designers of special purpose chips need to explore parameters such as the optimal bit width and data representation. This is the case for the development of complex algorithms such as Low-Density Parity-Check (LDPC) decoders used in modern communication systems. Currently, high-performance computing offers a wide set of acceleration options, that range from multicore CPUs to graphics processing units (GPUs) and FPGAs. Depending on the simulation requirements, the ideal architecture to use can vary. In this paper we propose a new design flow based on Open CL, a unified multiplatform programming model, which accelerates LDPC decoding simulations, thereby significantly reducing architectural exploration and design time. Open CL-based parallel kernels are used without modifications or code tuning on multicore CPUs, GPUs and FPGAs. We use SOpen CL (Silicon to Open CL), a tool that automatically converts Open CL kernels to RTL for mapping the simulations into FPGAs. To the best of our knowledge, this is the first time that a single, unmodified Open CL code is used to target those three different platforms. We show that, depending on the design parameters to be explored in the simulation, on the dimension and phase of the design, the GPU or the FPGA may suit different purposes more conveniently, providing different acceleration factors. For example, although simulations can typically execute more than 3× faster on FPGAs than on GPUs, the overhead of circuit synthesis often outweighs the benefits of FPGA-accelerated execution.
Gabriel Falcão Paiva Fernandes, Muhsen Owaida, David Novo, Madhura Purnaprajna, Nikolaos Bellas, Christos D. Antonopoulos, Georgios Karakonstantis, Andreas Peter Burg, Paolo Ienne
FCCM7
2011 Significance driven computation on next-generation unreliable platforms
abstract
In this paper, we propose a design paradigm for energy efficient and variation-aware operation of next-generation multicore heterogeneous platforms. The main idea behind the proposed approach lies on the observation that not all operations are equally important in shaping the output quality of various applications and of the overall system. Based on such an observation, we suggest that all levels of the software design stack, including the programming model, compiler, operating system (OS) and runtime system should identify the critical tasks and ensure correct operation of such tasks by assigning them to dynamically adjusted reliable cores/units. Specifically, based on error rates and operating conditions identified by a sense-and-adapt (SeA) unit, the OS selects and sets the right mode of operation of the overall system. The run-time system identifies the critical/less-critical tasks based on special directives and schedules them to the appropriate units that are dynamically adjusted for highly-accurate/approximate operation by tuning their voltage/frequency. Units that execute less significant operations can operate at voltages less than what is required for correct operation and consume less power, if required, since such tasks do not need to be always exact as opposed to the critical ones. Such scheme can lead to energy efficient and reliable operation, while reducing the design cost and overheads of conventional circuit/micro-architecture level techniques.
Georgios Karakonstantis, Nikolaos Bellas, Christos D. Antonopoulos, Georgios Tziantzioulis, Vaibhav Gupta, Kaushik Roy 0001
DAC1
2011 Multi-level wordline driver for low power SRAMs in nano-scale CMOS technology
abstract
In this paper, a multi-level wordline driver scheme is presented to improve SRAM read and write stability while lowering power consumption during hold operation. The proposed circuit applies a shaped wordline voltage pulse during read mode and a boosted wordline pulse during write mode. During read, the applied shaped pulse is tuned at nominal voltage for short period of time, whereas for the remaining access time, the wordline voltage is reduced to a lower level. This pulse results in improved read noise margin without any degradation in access time which is explained by examining the dynamic and nonlinear behavior of the SRAM cell. Furthermore, during hold mode, the wordline voltage starts from a negative value and reaches zero voltage, resulting in a lower leakage current compared to conventional SRAM. Our simulations using TSMC 65nm process show that the proposed wordline driver results in 2X improvement in static read noise margin while the write margin is improved by 3X. In addition, the total leakage of the proposed SRAM is reduced by 10% while the total power is improved by 12% in the worst case scenario of a single SRAM cell. The total area penalty is 10% for a 128Kb standard SRAM array.
Farshad Moradi, Georgios Panagopoulos, Georgios Karakonstantis, Dag T. Wisland, Hamid Mahmoodi, Jens Kargaard Madsen, Kaushik Roy 0001
ICCD3
2010 VEDA: Variation-aware energy-efficient Discrete Wavelet Transform architecture
abstract
In this paper, we present a unified approach to an energy-efficient variation-tolerant design of Discrete Wavelet Transform (DWT) in the context of image processing applications. It is to be noted that it is not necessary to produce exactly correct numerical outputs in most image processing applications. We exploit this important feature and propose a design methodology for DWT which shows energy quality tradeoffs at each level of design hierarchy starting from the algorithm level down to the architecture and circuit levels by taking advantage of the limited perceptual ability of the Human Visual System. A unique feature of this design methodology is that it guarantees robustness under process variability and facilitates aggressive voltage over-scaling. Simulation results show significant energy savings (74%-83%) with minor degradations in output image quality and avert catastrophic failures under process variations compared to a conventional design.
Vaibhav Gupta, Georgios Karakonstantis, Debabrata Mohapatra, Kaushik Roy 0001
ICCD2
2010 A self-consistent model to estimate NBTI degradation and a comprehensive on-line system lifetime enhancement technique
abstract
Lifetime reliability and the resultant temporal performance degradation due to Negative Bias Temperature Instability (NBTI) has emerged as a critical challenge in design and test of integrated circuits in nanometer technology nodes. In this work, we have developed a model that self-consistently estimates the NBTI degradation by considering the impact on the circuit lifetime of inter-dependent parameters such as Vddand temperature simultaneously. Using the proposed model, we observed that a circuit with lower Vddcan provide better lifetime performance than with higher Vdd. This interesting observation can be attributed to the reduction of electric field in the transistor along with the circuit power/temperature reduction that leads to lesser NBTI degradation. Based on this observation we have developed a on-line detection and mitigation scheme that allows Vddscaling to enhance system lifetime. The proposed scheme was applied to various arithmetic units and results in 45nm IBM process technology show 18% lifetime improvement with 57% reduction in power compared to conventional mitigation techniques. We also show that by using existing NBTI estimation models, the error in delay estimation can be as large as 7.6%.
Georgios Karakonstantis, Charles Augustine, Kaushik Roy 0001
IOLTS1
2010 HERQULES: system level cross-layer design exploration for efficient energy-quality trade-offs
abstract
In this paper, we present a unique cross-layer design framework that allows systematic exploration of the energy-delay-quality trade-offs at the algorithm, architecture and circuit level of design abstraction for each block of a system. In addition, taking into consideration the interactions between different sub-blocks of a system, it identifies the design solutions that can ensure the least energy at the "right amount of quality" for each sub-block/system under user quality/delay constraints. This is achieved by deriving sensitivity based design criteria, the balancing of which form the quantitative relations that can be used early in the system design process to evaluate the energy efficiency of various design options. The proposed framework when applied to the exploration of energy-quality design space of the main blocks of a digital camera and a wireless receiver, achieves 58% and 33% energy savings under 41% and 20% error increase, respectively.
Georgios Karakonstantis, Georgios Panagopoulos, Kaushik Roy 0001
ISLPED1
2010 Low-power DWT-based quasi-averaging algorithm and architecture for epileptic seizure detection
abstract
In this paper, we have developed a low-complexity algorithm for epileptic seizure detection with a high degree of accuracy. The algorithm has been designed to be feasibly implementable as battery-powered low-power implantable epileptic seizure detection system or epilepsy prosthesis. This is achieved by utilizing design optimization techniques at different levels of abstraction. Particularly, user-specific critical parameters are identified at the algorithmic level and are explicitly used along with multiplier-less implementations at the architecture level. The system has been tested on neural data obtained from in-vivo animal recordings and has been implemented in 90nm bulk-Si technology. The results show up to 90 % savings in power as compared to prevalent wavelet based seizure detection technique while achieving 97% average detection rate.
Himanshu Markandeya, Georgios Karakonstantis, Raghunathan Shriram, Pedro P. Irazoqui, Kaushik Roy 0001
ISLPED2
2010 Voltage Scalable High-Speed Robust Hybrid Arithmetic Units Using Adaptive Clocking
abstract
In this paper, we explore various arithmetic units for possible use in high-speed, high-yield ALUs operated at scaled supply voltage with adaptive clock stretching. We demonstrate that careful logic optimization of the existing arithmetic units (to create hybrid units) indeed make them further amenable to supply voltage scaling. Such hybrid units result from mixing right amount of fast arithmetic into the slower ones. Simulations on differenthybridadder and multipliers in BPTM 70 nm technology show 18%-50% improvements in power compared to standard adders with only 2%-8% increase in die-area at iso-yield. These optimized datapath units can be used to construct voltage scalable robust ALUs that can operate at high clock frequency with minimal performance degradation due to occasional clock stretching.
Swaroop Ghosh, Debabrata Mohapatra, Georgios Karakonstantis, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Process-Variation Resilient and Voltage-Scalable DCT Architecture for Robust Low-Power Computing
abstract
In this paper, we present a novel discrete cosine transform (DCT) architecture that allows aggressive voltage scaling for low-power dissipation, even under process parameter variations with minimal overhead as opposed to existing techniques. Under a scaled supply voltage and/or variations in process parameters, any possible delay errors appear only from the long paths that are designed to be less contributive to output quality. The proposed architecture allows a graceful degradation in the peak SNR (PSNR) under aggressive voltage scaling as well as extreme process variations. Results show that even under large process variations (±3σ around mean threshold voltage) and aggressive supply voltage scaling (at 0.88 V, while the nominal voltage is 1.2 V for a 90-nm technology), there is a gradual degradation of image quality with considerable power savings (71% at PSNR of 23.4 dB) for the proposed architecture, when compared to existing implementations in a 90-nm process technology.
Georgios Karakonstantis, Nilanjan Banerjee, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2009 Significance driven computation: a voltage-scalable, variation-aware, quality-tuning motion estimator
abstract
In this paper we present a design methodology for algorithm/architecture co-design of a voltage-scalable, process variation aware motion estimator based on significance driven computation. The fundamental premise of our approach lies in the fact that all computations are not equally significant in shaping the output response of video systems. We use a statistical technique to intelligently identify these significant/not-so-significant computations at the algorithmic level and subsequently change the underlying architecture such that the significant computations are computed in an error free manner under voltage over-scaling. Furthermore, our design includes an adaptive quality compensation (AQC) block which tunes the algorithm and architecture depending on the magnitude of voltage over-scaling and severity of process variations. Simulation results show average power savings of ~ 33% for the proposed architecture when compared to conventional implementation in the 90 nm CMOS technology. The maximum output quality loss in terms of Peak Signal to Noise Ratio (PSNR) was ~ 1 dB without incurring any throughput penalty.
Debabrata Mohapatra, Georgios Karakonstantis, Kaushik Roy 0001
ISLPED2
2009 Design Methodology for Low Power and Parametric Robustness Through Output-Quality Modulation: Application to Color-Interpolation Filtering
abstract
Power dissipation and robustness to process variation have conflicting design requirements. Scaling of voltage is associated with larger variations, while Vdd upscaling or transistor up-sizing for parametric-delay variation tolerance can be detrimental for power dissipation. However, for a class of signal-processing systems, effective tradeoff can be achieved between Vdd scaling, variation tolerance, and ldquooutput quality.rdquo In this paper, we develop a novel low-power variation-tolerant algorithm/architecture for color interpolation that allows a graceful degradation in the peak-signal-to-noise ratio (PSNR) under aggressive voltage scaling as well as extreme process variations. This feature is achieved by exploiting the fact that all computations used in interpolating the pixel values do not equally contribute to PSNR improvement. In the presence of Vdd scaling and process variations, the architecture ensures that only the ldquoless important computationsrdquo are affected by delay failures. We also propose a different sliding-window size than the conventional one to improve interpolation performance by a factor of two with negligible overhead. Simulation results show that, even at a scaled voltage of 77% of nominal value, our design provides reasonable image PSNR with 40% power savings.
Nilanjan Banerjee, Georgios Karakonstantis, Jung Hwan Choi, Chaitali Chakrabarti, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Process variation tolerant low power DCT architecture
abstract
2D discrete cosine transform (DCT) is widely used as the core of digital image and video compression. In this paper, the authors present a novel DCT architecture that allows aggressive voltage scaling by exploiting the fact that not all intermediate computations are equally important in a DCT system to obtain "good" image quality with peak signal to noise ratio (PSNR) > 30 dB. This observation has led us to propose a DCT architecture where the signal paths that are less contributive to PSNR improvement are designed to be longer than the paths that are more contributive to PSNR improvement It should also be noted that robustness with respect to parameter variations and low power operation typically impose contradictory requirements in terms of architecture design. However, the proposed architecture lends itself to aggressive voltage scaling for low-power dissipation even under process parameter variations. Under a scaled supply voltage and/or variations in process parameters, any possible delay errors would only appear from the long paths that are less contributive towards PSNR improvement, providing large improvement in power dissipation with small PSNR degradation. Results show that even under large process variation and supply voltage scaling (0.8V), there is a gradual degradation of image quality with considerable power savings (62.8%) for the proposed architecture when compared to existing implementations in 70 nm process technology
Nilanjan Banerjee, Georgios Karakonstantis, Kaushik Roy 0001
DATE2
2007 An Optimal Algorithm for Low Power Multiplierless FIR Filter Design using Chebychev Criterion
abstract
In this paper, we propose a novel finite impulse response (FIR) filter design methodology that reduces the number of operations with a motivation to reduce power consumption and enhance performance. The novelty of our approach lies in the generation of filter coefficients such that they conform to a given low-power architecture, while meeting the given filter specifications. The proposed algorithm is formulated as a mixed integer linear programming problem that minimizes Chebychev error and synthesizes coefficients which consist of pre-specified alphabets. The new modified coefficients can be used for low-power VLSI implementation of vector scaling operations such as FIR filtering using computation sharing multiplier (CSHM). Simulations in 0.25 μm technology show that CSHM FIR filter architecture can result in 55% power and 34% speed improvement compared to carry save multiplier (CSAM) based filters.
Georgios Karakonstantis, Kaushik Roy 0001
ICASSP (2)1
2007 Design methodology to trade off power, output quality and error resiliency: application to color interpolation filtering
abstract
Power dissipation and tolerance to process variations pose conflicting design requirements. Scaling of voltage is associated with larger variations, while Vdd upscaling or transistor up-sizing for process tolerance can be detrimental for power dissipation. However, for certain signal processing systems such as those used in color image processing, we noted that effective trade-offs can be achieved between Vdd scaling, process tolerance and “output quality”. In this paper we demonstrate how these tradeoffs can be effectively utilized in the development of novel low-power variation tolerant architectures for color interpolation. The proposed architecture supports a graceful degradation in the PSNR (Peak Signal to Noise Ratio) under aggressive voltage scaling as well as extreme process variations in sub-70nm technologies. This is achieved by exploiting the fact that some computations are more important and contribute more to the PSNR improvement compared to the others. The computations are mapped to the hardware in such a way that only the less important computations are affected by Vdd-scaling and process variations. Simulation results show that even at a scaled voltage of 60% of nominal Vdd value, our design provides reasonable image PSNR with 69% power savings
Georgios Karakonstantis, Nilanjan Banerjee, Kaushik Roy 0001, Chaitali Chakrabarti
ICCAD1
2007 Low-power process-variation tolerant arithmetic units using input-based elastic clocking
abstract
In this paper we propose a design methodology for low-power high-performance, process-variation tolerant architecture for arithmetic units. The novelty of our approach lies in the fact that possible delay failures due to process variations and/or voltage scaling are predicted in advance and addressed by employing an elastic clocking technique. The prediction mechanism exploits the dependence of delay of arithmetic units upon input data patterns and identifies specific inputs that activate the critical path. Under iso-yield conditions, the proposed design operates at a lower scaled down Vdd without any performance degradation, while it ensures a superlative yield under a design style employing nominal supply and transistor threshold voltage. Simulation results show power savings of upto 29%, energy per computation savings of upto 25.5% and yield enhancement of upto 11.1% compared to the conventional adders and multipliers implemented in the 70nm BPTM technology. We incorporated the proposed modules in the execution unit of a five stage DLX pipeline to measure performance using SPEC2000 benchmarks [9]. Maximum area and throughput penalty obtained were 10% and 3% respectively.
Debabrata Mohapatra, Georgios Karakonstantis, Kaushik Roy 0001
ISLPED2