Georgios Zervakis 0001

dblp:135/0008 · DBLP profile ↗
← Back
53ranked-venue papers
9as first author
40since 2021 · last 2026
0000-0001-8110-7122ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 53 · 9 first-author · 40 since 2021Software engineering, systems software and programming languages · 14 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Design and Optimization of Mixed-Kernel Mixed-Signal SVMs for Flexible Electronics
abstract
Flexible Electronics (FE) have emerged as a promising alternative to silicon-based technologies, offering on-demand low-cost fabrication, conformality, and sustainability. However, their large feature sizes severely limit integration density, imposing strict area and power constraints, thus prohibiting the realization of Machine Learning (ML) circuits, which can significantly enhance the capabilities of relevant near-sensor applications. Support Vector Machines (SVMs) offer high accuracy in such applications at relatively low computational complexity, satisfying FE technologies’ constraints. Existing SVM designs rely solely on linear or Radial Basis Function (RBF) kernels, forcing a trade-off between hardware costs and accuracy. Linear kernels, implemented digitally, minimize overhead but sacrifice performance, while the more accurate RBF kernels are prohibitively large in digital, and their analog realization contains inherent functional approximation. In this work, we propose the first mixed-kernel and mixed-signal SVM design in FE, which unifies the advantages of both implementations and balances the cost/accuracy trade-off. To that end, we introduce a co-optimization approach that trains our mixed-kernel SVMs and maps binary SVM classifiers to the appropriate kernel (linear/RBF) and domain (digital/analog), aiming to maximize accuracy whilst reducing the number of costly RBF classifiers. Our designs deliver 7.7% higher accuracy than state-of-the-art single-kernel linear SVMs, and reduce area and power by 108× and 17× on average compared to digital RBF implementations.
Florentia Afentaki, Maha Shatta, Konstantinos Balaskas, Georgios Panagopoulos, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
DATE5
2026 Bespoke Co-processor for Energy-Efficient Health Monitoring on RISC-V-based Flexible Wearables
abstract
Flexible electronics offer unique advantages for conformable, lightweight, and disposable healthcare wearables. However, their limited gate count, large feature sizes, and high static power consumption make on-body machine learning classification highly challenging. While existing bendable RISC-V systems provide compact solutions, they lack the energy efficiency required. We present a mechanically flexible RISC-V that integrates a bespoke multiply-accumulate co-processor with fixed coefficients to maximize energy efficiency and minimize latency. Our approach formulates a constrained programming problem to jointly determine co-processor constants and optimally map Multi-Layer Perceptron (MLP) inference operations, enabling compact, model-specific hardware by leveraging the low fabrication and non-recurring engineering costs of flexible technologies. Post-layout results demonstrate near-real-time performance across several healthcare datasets, with our circuits operating within the power budget of existing flexible batteries and occupying only 2.42mm2, offering a promising path toward accessible, sustainable, and conformable healthcare wearables. Our microprocessors achieve an average 2.35x speedup and 2.15x lower energy consumption compared to the state of the art.
Theofanis Vergos, Polykarpos Vergos, Mehdi Baradaran Tahoori, Georgios Zervakis 0001
DATE4
2026 Lightweight Fault Resilient Flexible Flash-ADCs
Florentia Afentaki, Paula L. Duarte, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
IOLTS3
2025 Design and In-training Optimization of Binary Search ADC for Flexible Classifiers
abstract
Flexible Electronics (FE) offer distinct advantages, including mechanical flexibility and low process temperatures, enabling extremely low-cost production. To address the demands of applications such as smart sensors and wearables, flexible devices must be small and operate at low supply voltages. Additionally, target applications often require classifiers to operate directly on analog sensory input, necessitating the use of Analog to Digital Converters (ADCs) to process the sensory data. However, ADCs present serious challenges, particularly in terms of high area and power consumption, especially when considering stringent area and energy budget. In this work, we target common classifiers in this domain such as MLPs and SVMs and present a holistic approach to mitigate the elevated overhead of analog to digital interfacing in FE. First, we propose a novel design for Binary Search ADC that reduces area overhead 2× compared with the state-of-the-art Binary design and up to 5.4× compared with Flash ADC. Next, we present an in-training ADC optimization in which we keep the bare-minimum representations required and simplifying ADCs by removing unnecessary components. Our in-training optimization further reduces on average the area in terms of transistor count of the required ADCs by 5× for less than 1% accuracy loss.
Paula L. Duarte, Florentia Afentaki, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
ASP-DAC3
2025 Sequential Printed Multilayer Perceptron Circuits for Super-TinyML Multi-Sensory Applications
abstract
Super-TinyML aims to optimize machine learning models for deployment on ultra-low-power application domains such as wearable technologies and implants. Such domains also require conformality, flexibility, and non-toxicity which traditional silicon-based systems cannot fulfill. Printed Electronics (PE) offers not only these characteristics, but also cost-effective and on-demand fabrication. However, Neural Networks (NN) with hundreds of features ---often necessary for target applications--- have not been feasible in PE because of its restrictions such as limited device count due to its large feature sizes. In contrast to the state of the art using fully parallel architectures and limited to smaller classifiers, in this work we implement a super-TinyML architecture for bespoke (application-specific) NNs that surpasses the previous limits of state of the art and enables NNs with large number of parameters. With the introduction of super-TinyML into PE technology, we address the area and power limitations through resource sharing with multi-cycle operation and neuron approximation. This enables, for the first time, the implementation of NNs with up to 35.9× more features and 65.4× more coefficients than the state of the art solutions.
Gurol Saglam, Florentia Afentaki, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
ASP-DAC3
2025 Late Breaking Results: Energy-Efficient Printed Machine Learning Classifiers with Sequential SVMs
abstract
Printed Electronics (PE) provide a mechanically flexible and cost-effective solution for machine learning (ML) circuits, compared to silicon-based technologies. However, due to large feature sizes, printed classifiers are limited by high power, area, and energy overheads, which restricts the realization of battery-powered systems. In this work, we design sequential printed bespoke Support Vector Machine (SVM) circuits that adhere to the power constraints of existing printed batteries while minimizing energy consumption, thereby boosting battery life. Our results show 6.5x energy savings while maintaining higher accuracy compared to the state of the art.
Spyridon Besias, Ilias Sertaridis, Florentia Afentaki, Konstantinos Balaskas, Georgios Zervakis 0001
DATE5
2025 Late Breaking Results: Leveraging Approximate Computing for Carbon-Aware DNN Accelerators
abstract
The rapid growth of Machine Learning (ML) has increased demand for DNN hardware accelerators, but their embodied carbon footprint poses significant environmental challenges. This paper leverages approximate computing to design sustainable accelerators by minimizing the Carbon Delay Product (CDP). Using gate-level pruning and precision scaling, we generate area-aware approximate multipliers and optimize the accelerator design with a genetic algorithm. Results demonstrate reduced embodied carbon while meeting performance and accuracy requirements.
Aikaterini Maria Panteleaki, Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Iraklis Anagnostopoulos
DATE3
2025 Computing with Printed and Flexible Electronics
Mehdi Baradaran Tahoori, Georgios Zervakis 0001, Konstantinos Balaskas, Priyanjana Pal
ETS3
2025 R2T-Tiny: Runtime-Reconfigurable Throughput-Optimized TinyML for Hybrid Inference Acceleration on FPGA SoCs
abstract
The emergence of tiny machine learning (TinyML) has represented a paradigm shift toward energy efficient and on-device inference, with TinyML research primarily focusing on low-cost and energy-efficiency microcontroller units (MCUs). However, the low computational capabilities of MCUs greatly limit the performance of TinyML applications, especially in the case of throughput driven tasks. Field programmable gate arrays (FPGAs), are therefore a promising alternative to MCUs, offering high parallelism and low energy requirements. FPGA-based accelerators typically follow one of two directions: high-throughput resource-intensive pipelined designs, or low-throughput sequential systolic arrays. Incorporating both approaches into a hybrid streaming/sequential accelerator strikes a balance between resource efficiency and throughput, but incurs resource contention within a tiny FPGA. In this work, we address the aforementioned limitations and propose a throughput-driven hybrid acceleration methodology for TinyML. We introduce R2T-Tiny, an adaptive framework that brings layer-wise customizability into throughput-driven inference. By leveraging runtime partial reconfiguration, R2T-Tiny dynamically adjusts the accelerator type and applies tailored approximations per layer, achieving high throughput while adhering to the tight resource constraints of tiny embedded FPGAs. Our comprehensive evaluation on popular TinyML benchmarks showcases the capabilities of our framework in achieving high throughput inference on the PYNQ-Z2 FPGA board, increasing throughput by an average of 1.6x across 3 popular deep neural networks (DNNs) in the TinyML domain, when compared to DNNDK, a systolic array based accelerator from Xilinx, while incurring less than 1% accuracy loss.
Georgios Mentzos, Valentin Alexander Frey, Konstantinos Balaskas, Georgios Zervakis 0001, Jörg Henkel
ICCAD4
2025 Invited Paper: Feature-to-Classifier Co-Design for Mixed-Signal Smart Flexible Wearables for Healthcare at the Extreme Edge
abstract
Flexible Electronics (FE) offer a promising alternative to rigid silicon-based hardware for wearable healthcare devices, enabling lightweight, conformable, and low-cost systems. However, their limited integration density and large feature sizes impose strict area and power constraints, making ML-based healthcare systems–integrating analog frontend, feature extraction and classifier–particularly challenging. Existing FE solutions often neglect potential system-wide solutions and focus on the classifier, overlooking the substantial hardware cost of feature extraction and Analog-to-Digital Converters (ADCs)–both major contributors to area and power consumption. In this work, we present a holistic mixed-signal feature-to-classifier co-design framework for flexible smart wearable systems. To the best of our knowledge, we design the first analog feature extractors in FE, significantly reducing feature extraction cost. We further propose an hardware-aware NAS-inspired feature selection strategy within ML training, enabling efficient, application-specific designs. Our evaluation on healthcare benchmarks shows our approach delivers highly accurate, ultra-area-efficient flexible systems–ideal for disposable, low-power wearable monitoring.
Maha Shatta, Konstantinos Balaskas, Paula L. Duarte, Georgios Panagopoulos, Mehdi Baradaran Tahoori, Georgios Zervakis 0001
ICCAD6
2025 Leveraging Image Difficulty for Run-Time Adaptive DNN Inference on Embedded Devices
abstract
Deep Neural Networks (DNNs) impose great challenges on resource-constraint embedded devices since they employ billions of computational operations. To satisfy these computational demands such devices utilize hardware accelerators, which can lead to increased power consumption whatsoever. To that end, the compression of DNNs to lower precision has been proposed, in order to achieve savings in energy consumption at the cost of some accuracy loss during inference. However, DNNs do not behave similarly under lower-precision execution and the accuracy degradation can be severe. In this work, we utilize the notion of image difficulty and explore how we can change DNN precision during inference to achieve gains in energy consumption without big drops in accuracy. We evaluate our work on the ImageNet dataset and show how the proposed framework achieves energy savings at run-time.
Vasileios Pentsos, Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos
ISCAS3
2025 Compact Yet Highly Accurate Printed Classifiers Using Sequential Support Vector Machine Circuits
abstract
Printed Electronics (PE) technology has emerged as a promising alternative to silicon-based computing. It offers attractive properties such as on-demand ultra-low-cost fabrication, mechanical flexibility, and conformality. However, PE are governed by large feature sizes, prohibiting the realization of complex printed Machine Learning (ML) classifiers. Leveraging PE’s ultra-low non-recurring engineering and fabrication costs, designers can fully customize hardware to a specific ML model and dataset, significantly reducing circuit complexity. Despite significant advancements, state-of-the-art solutions achieve area efficiency at the expense of considerable accuracy loss. Our work mitigates this by designing area- and power-efficient printed ML classifiers with little to no accuracy degradation. Specifically, we introduce the first sequential Support Vector Machine (SVM) classifiers, exploiting the hardware efficiency of bespoke control and storage units and a single Multiply-Accumulate compute engine. Our SVMs yield on average 6x lower area and 4.6% higher accuracy compared to the printed state of the art.
Ilias Sertaridis, Spyridon Besias, Florentia Afentaki, Konstantinos Balaskas, Georgios Zervakis 0001
ISCAS5
2025 Approximate Multiplier Mapping for Unfairness Mitigation in Energy-Efficient DNNs
abstract
Embedded devices struggle with the heavy computational demands of extensive neural network models, a problem partially addressed by integrating accelerators with numerous multiply-accumulate units. However, this solution increases energy consumption. While using approximate circuits in accelerators can lower energy usage, it compromises accuracy and raises concerns about maintaining fair inference across diverse populations, particularly in the medical field. This work leverages approximate multipliers in deep neural networks to retain fairness and reduce energy consumption while keeping the inference accuracy within strict thresholds.
Ourania Spantidi, Georgios Zervakis 0001, Jörg Henkel, Iraklis Anagnostopoulos
ISCAS2
2025 Exploration of Low-Power Flexible Stress Monitoring Classifiers for Conformal Wearables
abstract
Conventional stress monitoring relies on episodic, symptom-focused interventions, missing the need for continuous, accessible, and cost-efficient solutions. State-of-the-art approaches use rigid, silicon-based wearables, which, though capable of multitasking, are not optimized for lightweight, flexible wear, limiting their practicality for continuous monitoring. In contrast, flexible electronics (FE) offer flexibility and low manufacturing costs, enabling real-time stress monitoring circuits. However, implementing complex circuits like machine learning (ML) classifiers in FE is challenging due to integration and power constraints. Previous research has explored flexible biosensors and ADCs, but classifier design for stress detection remains underexplored. This work presents the first comprehensive design space exploration of low-power, flexible stress classifiers. We cover various ML classifiers, feature selection, and neural simplification algorithms, with over 1200 flexible classifiers. To optimize hardware efficiency, fully customized circuits with low-precision arithmetic are designed in each case. Our exploration provides insights into designing real-time stress classifiers that offer higher accuracy than current methods, while being low-cost, conformable, and ensuring low power and compact size.
Florentia Afentaki, Sri Sai Rakesh Nakkilla, Konstantinos Balaskas, Paula L. Duarte, Shiyi Jiang, Georgios Zervakis 0001, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori
ISLPED6
2025 Enabling Printed Multilayer Perceptrons Realization via Area-Aware Neural Minimization
abstract
Printed Electronics (PE) set up a new path for the realization of ultra low-cost circuits that can be deployed in every-day consumer goods and disposables. In addition, PE satisfy requirements such as porosity, flexibility, and conformity. However, the large feature sizes in PE and limited device counts incur high restrictions and increased area and power overheads, prohibiting the realization of complex circuits. As a result, although printed Machine Learning (ML) circuits could open new horizons and bring “intelligence” in such domains, the implementation of complex classifiers, as required in target applications, is hardly feasible. In this paper, we aim to address this and focus on the design of battery-powered printed Multilayer Perceptrons (MLPs). To that end, we exploit fully-customized circuit (bespoke) implementations, enabled in PE, and propose a hardware-aware neural minimization framework dedicated for such customized MLP circuits. Our evaluation demonstrates that, for up to 3% accuracy loss, our co-design methodology enables, for the first time, battery-powered operation of complex printed MLPs.
Argyris Kokkinis, Georgios Zervakis 0001, Kostas Siozios, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Computers2
2024 Late Breaking Results: Language-level QoR modeling for High-Level Synthesis
abstract
This paper proposes a language-level modeling approach for HighLevel Synthesis based on the state-of-the-art Transformer architecture. Our approach estimates the performance and required resources of HLS applications directly from the source code when different synthesis directives, in terms of HLS #pragmas, are applied. Results show that the proposed architecture achieves 96.02% accuracy for predicting the feasibility class of applications and an average of 0.95 and 0.91 R2 scores for predicting the actual performance and required resources, respectively.
Dimosthenis Masouros, Aggelos Ferikoglou, Georgios Zervakis 0001, Sotirios Xydis, Dimitrios Soudris
DAC3
2024 Synthesis of Resource-Efficient Superconducting Circuits with Clock-Free Alternating Logic
abstract
Gate-level clocking, typical in traditional approaches to Single Flux Quantum (SFQ) technology, makes the effective synthesis of superconducting circuits a significant engineering hurdle. This paper addresses this challenge by employing the recently introduced alternating SFQ (xSFQ) logic family. xSFQ leverages dual-rail alternating encoding to eliminate the clock dependency from the superconducting gate semantics. This obviates the need for ad hoc modifications to existing synthesis tools and avoids unnecessary circuit resource overheads, marking a significant advancement in superconducting circuit design automation. Our implementation results demonstrate an average reduction of over 80% in the Josephson junction count for circuits from the ISCAS85, EPFL, and ISCAS89 benchmark suites.
Jennifer Volk, Panagiotis Papanikolaou, Georgios Zervakis 0001, Georgios Tzimpragos
DAC3
2024 Embedding Hardware Approximations in Discrete Genetic-Based Training for Printed MLPs
abstract
Printed Electronics (PE) stands out as a promising technology for widespread computing due to its distinct attributes, such as low costs and flexible manufacturing. Unlike traditional silicon-based technologies, PE enables stretchable, conformal, and non-toxic hardware. However, PE are constrained by larger feature sizes, making it challenging to implement complex circuits such as machine learning (ML) classifiers. Approximate computing has been proven to reduce the hardware cost of ML circuits such as Multilayer Perceptrons (MLPs). In this paper, we maximize the benefits of approximate computing by integrating hardware approximation into the MLP training process. Due to the discrete nature of hardware approximation, we propose and implement a genetic-based, approximate, hardware-aware training approach specifically designed for printed MLPs. For a 5% accuracy loss, our MLPs achieve over 5 × area and power reduction compared to the baseline while outperforming state-of-the-art approximate and stochastic printed MLPs.
Florentia Afentaki, Michael Hefenbrock, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
DATE3
2024 On-Sensor Printed Machine Learning Classification via Bespoke ADC and Decision Tree Co-Design
abstract
Printed electronics (PE) technology provides cost-effective hardware with unmet customization, due to their low non-recurring engineering and fabrication costs. PE exhibit features such as flexibility, stretchability, porosity, and conformality, which make them a prominent candidate for enabling ubiquitous computing. Still, the large feature sizes in PE limit the realization of complex printed circuits, such as machine learning classifiers, especially when processing sensor inputs is necessary, mainly due to the costly analog-to-digital converters (ADCs). To this end, we propose the design of fully customized ADCs and present, for the first time, a co-design framework for generating bespoke Decision Tree classifiers. Our comprehensive evaluation shows that our co-design enables self-powered operation of on-sensor printed classifiers in all benchmark cases.
Giorgos Armeniakos, Paula L. Duarte, Priyanjana Pal, Georgios Zervakis 0001, Mehdi Baradaran Tahoori, Dimitrios Soudris
DATE4
2024 Fault Sensitivity Analysis of Printed Bespoke Multilayer Perceptron Classifiers
abstract
Printed Electronics (PE) is an emerging technology with flexible substrates and ultra-low-cost manufacturing, providing an appealing alternative to traditional wafer-scale silicon fabrication. With the increasing integration of various printed neural network (NN) architectures in diverse applications, the reliability of printed circuits has become a critical concern. This work provides a comprehensive analysis of the fault sensitivity on a variety of classification tasks for various digital and analog realizations of printed multilayer perceptrons (MLPs). We further evaluate different digital architectures, i.e., generic, bespoke, and approximate, to provide a comprehensive fault analysis on different benchmark datasets.
Priyanjana Pal, Florentia Afentaki, Haibin Zhao, Gurol Saglam, Michael Hefenbrock, Georgios Zervakis 0001, Michael Beigl, Mehdi Baradaran Tahoori
ETS6
2024 Evolutionary Approximation of Ternary Neurons for On-sensor Printed Neural Networks
abstract
Printed electronics offer ultra-low manufacturing costs and the potential for on-demand fabrication of flexible hardware. However, significant intrinsic constraints stemming from their large feature sizes and low integration density pose design challenges that hinder their practicality. In this work, we conduct a holistic exploration of printed neural network accelerators, starting from the analog-to-digital interface---a major area and power sink for sensor processing applications---and extending to networks of ternary neurons and their implementation. We propose bespoke ternary neural networks using approximate popcount and popcount-compare units, developed through a multi-phase evolutionary optimization approach and interfaced with sensors via customizable analog-to-binary converters. Our evaluation results show that the presented designs outperform the state of the art, achieving at least 6× improvement in area and 19× in power. To our knowledge, they represent the first open-source digital printed neural network classifiers capable of operating with existing printed energy harvesters.
Vojtech Mrazek, Argyris Kokkinis, Panagiotis Papanikolaou, Zdenek Vasícek, Kostas Siozios, Georgios Tzimpragos, Mehdi Baradaran Tahoori, Georgios Zervakis 0001
ICCAD8
2023 Hardware-Aware Automated Neural Minimization for Printed Multilayer Perceptrons
abstract
The demand of many application domains for flexibility, stretchability, and porosity cannot be typically met by the silicon VLSI technologies. Printed Electronics (PE) has been introduced as a candidate solution that can satisfy those requirements and enable the integration of smart devices on consumer goods at ultra low-cost enabling also in situ and on-demand fabrication. However, the large features sizes in PE constraint those efforts and prohibit the design of complex ML circuits due to area and power limitations. Though, classification is mainly the core task in printed applications. In this work, we examine, for the first time, the impact of neural minimization techniques, in conjunction with bespoke circuit implementations, on the area-efficiency of printed Multilayer Perceptron classifiers. Results show that for up to 5 % accuracy loss up to 8× area reduction can be achieved.
Argyris Kokkinis, Georgios Zervakis 0001, Kostas Siozios, Mehdi Baradaran Tahoori, Jörg Henkel
DATE2
2023 Bespoke Approximation of Multiplication-Accumulation and Activation Targeting Printed Multilayer Perceptrons
abstract
Printed Electronics (PE) feature distinct and remarkable characteristics that make them a prominent technology for achieving true ubiquitous computing. This is particularly relevant in application domains that require conformal and ultra-low cost solutions, which have experienced limited penetration of computing until now. Unlike silicon-based technologies, PE offer unparalleled features such as non-recurring engineering costs, ultra-low manufacturing cost, and on-demand fabrication of conformal, flexible, non-toxic, and stretchable hardware. However, PE face certain limitations due to their large feature sizes, that impede the realization of complex circuits, such as machine learning classifiers. In this work, we address these limitations by leveraging the principles of Approximate Computing and Bespoke (fully-customized) design. We propose an automated framework for designing ultra-low power Multilayer Perceptron (MLP) classifiers which employs, for the first time, a holistic approach to approximate all functions of the MLP's neurons: multiplication, accumulation, and activation. Through comprehensive evaluation across various MLPs of varying size, our framework demonstrates the ability to enable battery-powered operation of even the most intricate MLP architecture examined, significantly surpassing the current state of the art.
Florentia Afentaki, Gurol Saglam, Argyris Kokkinis, Kostas Siozios, Georgios Zervakis 0001, Mehdi Baradaran Tahoori
ICCAD5
2023 Co-Design of Approximate Multilayer Perceptron for Ultra-Resource Constrained Printed Circuits
abstract
Printed Electronics (PE) exhibits on-demand, extremely low-cost hardware due to its additive manufacturing process, enabling machine learning (ML) applications for domains that feature ultra-low cost, conformity, and non-toxicity requirements that silicon-based systems cannot deliver. Nevertheless, large feature sizes in PE prohibit the realization of complex printed ML circuits. In this work, we present, for the first time, an automated printed-aware software/hardware co-design framework that exploits approximate computing principles to enable ultra-resource constrained printed multilayer perceptrons (MLPs). Our evaluation demonstrates that, compared to the state-of-the-art baseline, our circuits feature on average 6x (5.7x) lower area (power) and less than 1% accuracy loss.
Giorgos Armeniakos, Georgios Zervakis 0001, Dimitrios Soudris, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Computers2
2023 Model-to-Circuit Cross-Approximation For Printed Machine Learning Classifiers
abstract
Printed electronics (PEs) promises on-demand fabrication, low nonrecurring engineering costs, and subcent fabrication costs. It also allows for high customization that would be infeasible in silicon, and bespoke architectures prevail to improve the efficiency of emerging PE machine learning (ML) applications. Nevertheless, large feature sizes in PE prohibit the realization of complex ML models in PE, even with bespoke architectures. In this work, we present an automated, cross-layer approximation framework tailored to bespoke architectures that enable complex ML models, such as multilayer perceptrons (MLPs) and support vector machines (SVMs), in PE. Our framework adopts cooperatively a hardware-driven coefficient approximation of the ML model at algorithmic level, a netlist pruning at logic level, and a voltage overscaling at the circuit level. Extensive experimental evaluation on 12 MLPs and 12 SVMs and more than 6000 approximate and exact designs demonstrates that our model-to-circuit cross-approximation delivers power and area optimal designs that, compared to the state-of-the-art exact designs, feature on average 51% and 66% area and power reduction, respectively, for less than 5% accuracy loss. Finally, we demonstrate that our framework enables 80% of the examined classifiers to be battery-powered with almost identical accuracy with the exact designs, paving thus the way toward smart complex printed applications.
Giorgos Armeniakos, Georgios Zervakis 0001, Dimitrios Soudris, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 AdaPT: Fast Emulation of Approximate DNN Accelerators in PyTorch
abstract
Current state-of-the-art employs approximate multipliers to address the highly increased power demands of deep neural network (DNN) accelerators. However, evaluating the accuracy of approximate DNNs is cumbersome due to the lack of adequate support for approximate arithmetic in DNN frameworks. We address this inefficiency by presenting AdaPT, a fast emulation framework that extends PyTorch to support approximate inference as well as approximation-aware retraining. AdaPT can be seamlessly deployed and is compatible with the most DNNs. We evaluate the framework on several DNN models and application fields, including CNNs, LSTMs, and GANs for a number of approximate multipliers with distinct bitwidth values. The results show substantial error recovery from approximate retraining and reduced inference time up to$53.9 \times $with respect to the baseline approximate implementation.
Dimitrios Danopoulos, Georgios Zervakis 0001, Kostas Siozios, Dimitrios Soudris, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Cross-Layer Approximation For Printed Machine Learning Circuits
abstract
Printed electronics (PE) feature low non-recurring engineering costs and low per unit-area fabrication costs, enabling thus extremely low-cost and on-demand hardware. Such low-cost fabrication allows for high customization that would be infeasible in silicon, and bespoke architectures prevail to improve the efficiency of emerging PE machine learning (ML) applications. However, even with bespoke architectures, the large feature sizes in PE constraint the complexity of the ML models that can be implemented. In this work, we bring together, for the first time, approximate computing and PE design targeting to enable complex ML models, such as Multi-Layer Perceptrons (MLPs) and Support Vector Machines (SVMs), in PE. To this end, we propose and implement a cross-layer approximation, tailored for bespoke ML architectures. At the algorithmic level we apply a hardware-driven coefficient approximation of the ML model and at the circuit level we apply a netlist pruning through a full search exploration. In our extensive experimental evaluation we consider 14 MLPs and SVMs and evaluate more than 4300 approximate and exact designs. Our results demonstrate that our cross approximation delivers Pareto optimal designs that, compared to the state-of-the-art exact designs, feature 47% and 44% average area and power reduction, respectively, and less than 1% accuracy loss.
Giorgos Armeniakos, Georgios Zervakis 0001, Dimitrios Soudris, Mehdi Baradaran Tahoori, Jörg Henkel
DATE2
2022 Approximate Computing and the Efficient Machine Learning Expedition
abstract
Approximate computing (AxC) has been long accepted as a design alternative for efficient system implementation at the cost of relaxed accuracy requirements. Despite the AxC research activities in various application domains, AxC thrived the past decade when it was applied in Machine Learning (ML). The by definition approximate notion of ML models but also the increased computational overheads associated with ML applications-that were effectively mitigated by corresponding approximations-led to a perfect matching and a fruitful synergy. AxC for AI/ML has transcended beyond academic prototypes. In this work, we enlighten the synergistic nature of AxC and ML and elucidate the impact of AxC in designing efficient ML systems. To that end, we present an overview and taxonomy of AxC for ML and use two descriptive application scenarios to demonstrate how AxC boosts the efficiency of ML systems.
Jörg Henkel, Hai Li 0001, Anand Raghunathan, Mehdi Baradaran Tahoori, Swagath Venkataramani, Xiaoxuan Yang 0001, Georgios Zervakis 0001
ICCAD7
2022 Impact of NCFET Technology on Eliminating the Cooling Cost and Boosting the Efficiency of Google TPU
abstract
Recent breakthroughs in Neural Networks (NNs) led to significant accuracy improvements of several machine learning applications such as image classification and voice recognition. However, this accuracy improvement comes at the cost of an immense increase in computation demands. NNs became one of the most common and computationally intensive workloads in today's datacenters. To address these computational demands, Google announced in 2016 the Tensor Processing Unit (TPU), an advanced custom ASIC accelerator for NN inference. Two new TPU versions (v2 and v3) followed in 2017 and 2018 that support also training. Google TPUv3 packs an immense processing power ($\mathrm{90TFLOPS}$per chip) in a tiny and condensed area, leading to very high on-chip power densities and thus excessive temperature. In this article, superlattice thermoelectric cooling, which is one of the emerging on-chip cooling, is considered as an advanced cooling example for Google TPU and we investigate the impact of Negative Capacitance FET (NCFET), which is one of the recent emerging technologies, on the cooling and efficiency of TPU. Through full-chip design, of the computational core of the TPU, based on$14\mathrm{nm}$Intel FinFET technology and multiphysics temperature simulations, we demonstrate that NCFET can significantly minimize the required cooling-cost. More than 4000 NCFET configurations are evaluated in order to traverse the entire design space defined by the thickness of the ferroelectric layer of NCFET, the operating voltage, cooling, and the operating frequency, in addition to all possible FinFET's configurations. Moreover, our experimental evaluation shows that by eliminating the cooling cost, NCFET delivers 2.8x higher efficiency compared to the conventional FinFET baseline.
Sami Salamin, Georgios Zervakis 0001, Florian Klemme, Hammam Kattan, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch
IEEE Trans. Computers2
2022 Thermal-Aware Design for Approximate DNN Accelerators
abstract
Recent breakthroughs in Neural Networks (NNs) have made DNN accelerators ubiquitous and led to an ever-increasing quest on adopting them from Cloud to edge computing. However, state-of-the-art DNN accelerators pack immense computational power in a relatively confined area, inducing significant on-chip power densities that lead to intolerable thermal bottlenecks. Existing state of the art focuses on using approximate multipliers only to trade-off efficiency with inference accuracy. In this work, we present a thermal-aware approximate DNN accelerator design in which we additionally trade-off approximation with temperature effects towards designing DNN accelerators that satisfy tight temperature constraints. Using commercial multi-physics tool flows for heat simulations, we demonstrate how our thermal-aware approximate design reduces the temperature from 139$^{\circ }$C, in an accurate circuit, down to 79$^{\circ }$C. This enables DNN accelerators to fulfill tight thermal constraints, while still maximizing the performance and reducing the energy by around 75% with a negligible accuracy loss of merely 0.44% on average for a wide range of NN models. Furthermore, using physics-based transistor aging models, we demonstrate how reductions in voltage and temperature obtained by our approximate design considerably improve the circuit’s reliability. Our approximate design exhibits around 40% less aging-induced degradation compared to the baseline design.
Georgios Zervakis 0001, Iraklis Anagnostopoulos, Sami Salamin, Ourania Spantidi, Isai Roman-Ballesteros, Jörg Henkel, Hussam Amrouch
IEEE Trans. Computers1
2022 Energy-Efficient DNN Inference on Approximate Accelerators Through Formal Property Exploration
abstract
Deep neural networks (DNNs) are being heavily utilized in modern applications, putting energy-constraint devices to the test. To bypass high energy consumption issues, approximate computing has been employed in DNN accelerators to balance out the accuracy-energy reduction trade-off. However, the approximation-induced accuracy loss can be very high and drastically degrade the performance of the DNN. Therefore, there is a need for a fine-grain mechanism that would assign specific DNN operations to approximation to maintain acceptable DNN accuracy, while achieving low energy consumption. We present an automated framework for weight-to-approximation mapping through formal property exploration for approximate DNN accelerators. At the MAC unit level, our experimental evaluation surpassed already energy-efficient mappings by more than$\times 2$in terms of energy gains, while supporting a fine-grain control over the introduced approximation.
Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Variability-Aware Approximate Circuit Synthesis via Genetic Optimization
abstract
One of the major barriers that CMOS devices face at nanometer scale is increasing parameter variation due to manufacturing imperfections. Process variations severely inhibit the reliable operation of circuits, as the operational frequency at the nominal process corner is insufficient to suppress timing violations across the entire variability spectrum. To avoid variability-induced timing errors, previous efforts impose pessimistic and performance-degrading timing guardbands atop the operating frequency. In this work, we employ approximate computing principles and propose a circuit-agnostic automated framework for generating variability-aware approximate circuits that eliminate process-induced timing guardbands. Variability effects are accurately portrayed with the creation of variation-aware standard cell libraries, fully compatible with standard EDA tools. The underlying transistors are fully calibrated against industrial measurements from Intel 14nm FinFET in which both electrical characteristics of transistors and variability effects are accurately captured. In this work, we explore the design space of approximate variability-aware designs to automatically generate circuits of reduced variability and increased performance without the need for timing guardbands. Experimental results show that by introducing negligible functional error of merely$\boldsymbol {5.3 \times 10^{-3}}$, our variability-aware approximate circuits can be reliably operated under process variations without sacrificing the application performance.
Konstantinos Balaskas, Florian Klemme, Georgios Zervakis 0001, Kostas Siozios, Hussam Amrouch, Jörg Henkel
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Approximate Computing for ML: State-of-the-art, Challenges and Visions
abstract
In this paper, we present our state-of-the-art approximate techniques that cover the main pillars of approximate computing research. Our analysis considers both static and reconfigurable approximation techniques as well as operation-specific approximate components (e.g., multipliers) and generalized approximate highlevel synthesis approaches. As our application target, we discuss the improvements that such techniques bring on machine learning and neural networks. In addition to the conventionally analyzed performance and energy gains, we also evaluate the improvements that approximate computing brings in the operating temperature.
Georgios Zervakis 0001, Hassaan Saadat, Hussam Amrouch, Andreas Gerstlauer, Sri Parameswaran, Jörg Henkel
ASP-DAC1
2021 Control Variate Approximation for DNN Accelerators
abstract
In this work, we introduce a control variate approximation technique for low error approximate Deep Neural Network (DNN) accelerators. The control variate technique is used in Monte Carlo methods to achieve variance reduction. Our approach significantly decreases the induced error due to approximate multiplications in DNN inference, without requiring time-exhaustive retraining compared to state-of-the-art. Leveraging our control variate method, we use highly approximated multipliers to generate power-optimized DNN accelerators. Our experimental evaluation on six DNNs, for Cifar-10 and Cifar100 datasets, demonstrates that, compared to the accurate design, our control variate approximation achieves same performance and 24% power reduction for a merely 0.16% accuracy loss.
Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel
DAC1
2021 Reliability-Aware Quantization for Anti-Aging NPUs
Sami Salamin, Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Jörg Henkel, Hussam Amrouch
DATE2
2021 FeFET and NCFET for Future Neural Networks: Visions and Opportunities
abstract
The goal of this special session paper is to introduce and discuss different emerging technologies for logic circuitry and memory as well as new lightweight architectures for neural networks. We demonstrate how the ever-increasing complexity in Artificial Intelligent (AI) applications, resulting in an immense increase in the computational power, necessitates inevitably employing innovations starting from the underlying devices all the way up to the architectures. Two different promising emerging technologies will be presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new beyond-CMOS technology with advantages for offering low power and/or higher accuracy for neural network inference. (ii) Ferroelectric FET (FeFET) as a novel non-volatile, area-efficient and ultra-low power memory device. In addition, we demonstrate how Binarized Neural Networks (BNNs) offer a promising alternative for traditional Deep Neural Networks (DNNs) due to its lightweight hardware implementation. Finally, we present the challenges from combining FeFET-based NVM with NNs and summarize our perspectives for future NNs and the vital role that emerging technologies may play.
Mikail Yayla, Kuan-Hsun Chen, Georgios Zervakis 0001, Jörg Henkel, Jian-Jia Chen, Hussam Amrouch
DATE3
2021 Positive/Negative Approximate Multipliers for DNN Accelerators
abstract
Recent Deep Neural Networks (DNNs) manage to deliver superhuman accuracy levels on many AI tasks. DNN accelerators are becoming integral components of modern systems-on-chips. DNNs perform millions of arithmetic operations per inference and DNN accelerators integrate thousands of multiply-accumulate units leading to increased energy requirements. To lower the energy consumption of DNN accelerators, approximate computing principles are employed. However, complex DNNs can be increasingly sensitive to approximation. In this work, we present a dynamically configurable approximate multiplier that supports three operation modes, i.e., exact, positive error, and negative error. In addition, we propose a filter-oriented approximation method to map the weights to the appropriate modes of the approximate multiplier. Our mapping algorithm balances the positive with the negative errors due to the approximate multiplications, aiming at maximizing the energy reduction while minimizing the overall convolution error. We evaluate our approach on multiple DNNs and datasets against state-of-the-art approaches, where our method achieves 18.33% energy gains on average across 7 NNs on 4 different datasets for a maximum accuracy drop of only 1%.
Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel
ICCAD2
2021 Automated Design Approximation to Overcome Circuit Aging
abstract
Transistor aging phenomena manifest themselves as degradations in the main electrical characteristics of transistors. Over time, they result in a significant increase of cell propagation delay, leading to errors due to timing violations, since the operating frequency becomes unsustainable as the circuit ages. Conventional techniques employ timing guardbands to mitigate aging-induced delay increase, which leads to considerable performance losses from the beginning of the circuit’s lifetime. Leveraging the inherent error resilience of a vast number of application domains, approximate computing was recently introduced as an aging mitigation mechanism. In this work, we present the first automated framework for generatingaging-aware approximate circuits. Our framework, by applying directed gate-level netlist approximation, induces a small functional error and recovers the delay degradation due to aging. As a result, our optimized circuits eliminate aging-induced timing errors. Experimental evaluation over a variety of arithmetic circuits and image processing benchmarks demonstrates that for an average error of merely$5\times 10^{-3}$, our framework completely eliminates aging-induced timing guardbands. Compared to the respective baseline circuits without timing guardbands (i.e., iso-performance evaluation), the error of the circuits generated by our framework is$1208\times $smaller.
Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Jörg Henkel, Kostas Siozios
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 On the Resiliency of NCFET Circuits Against Voltage Over-Scaling
abstract
Approximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds.
Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 PROTON: Post-Synthesis Ferroelectric Thickness Optimization for NCFET Circuits
abstract
For the first time, we demonstrate an optimization technique to synthesize circuits in the Negative Capacitance FET (NCFET) technology. NCFET is a rapidly emerging technology to replace the currently employed CMOS technology due to its profound ability to overcome the fundamental limit in scaling along with its full compatibility with the existing fabrication process. This is achieved by replacing the traditional transistor gate dielectric with a ferroelectric layer that manifests itself as a Negative Capacitance (NC), which magnifies the electric field. As a result, NCFET-based circuits can operate at a higher clock frequency without the need to increase the operating voltage. NC breaks one of the fundamental laws in physics in which the total capacitance of two capacitors connected in series becomes larger–instead of smaller in ordinary capacitors– than each of them. This could lead to sub-optimal netlists, suffering from significant increase in dynamic power and IR-drops. To suppress that, we employ the relation between delay decrease and capacitance increase of gates w.r.t ferroelectric thickness. Our technique takes an optimized netlist, obtained from commercial EDA tools, and then selectively determines the optimal ferroelectric thickness for each gate in the netlist, so that the maximum performance provided by NCFET is still achieved while the dynamic power is considerably decreased (45% on average),i.e., no trade-offs. Particularly, our technique enables the full exploitation of the performance benefits originating by NCFET, at a significantly lower (power) cost. Compared to state of the art, our technique decreases the energy-delay-product of circuits by 25% on average and reduces the deleterious effects of IR-drop by 56%. Hence, efficiency and reliability of circuits are improved without any loss in the obtained performance from NCFET.
Sami Salamin, Georgios Zervakis 0001, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 NPU Thermal Management
abstract
Neural processing units (NPUs) are becoming an integral part in all modern computing systems due to their substantial role in accelerating neural networks (NNs). The significant improvements in cost-energy-performance stem from the massive array of multiply accumulate (MAC) units that remarkably boosts the throughput of NN inference. In this work, we are the first to investigate the thermal challenges that NPUs bring, revealing how MAC arrays, which form the heart of any NPU, impose serious thermal bottlenecks to on-chip systems due to their excessive power densities. For the first time, we explore: 1) the effectiveness of precision scaling and frequency scaling (FS) in temperature reductions and 2) how advanced on-chip cooling using superlattice thin-film thermoelectric (TE) open doors for new tradeoffs between temperature, throughput, cooling cost, and inference accuracy in NPU chips. Our work unveils that hybrid thermal management, which composes different means to reduce the NPU temperature, is a key. To achieve that, we propose and implement PFS-TE technique that couples precision and FS together with superlattice TE cooling for effective NPU thermal management. Using commercial signoff tools, we obtain accurate power and timing analysis of MAC arrays after a full-chip design is performed based on 14-nm Intel FinFET technology. Then, multiphysics simulations using finite-element methods are carried out for accurate heat simulations in the presence and absence of on-chip cooling. Afterward, comprehensive design-space exploration is presented to demonstrate the Pareto frontier and the existing tradeoffs between temperature reductions, power overheads due to cooling, throughput, and inference accuracy. Using a wide range of NNs trained for image classification, experimental results demonstrate that our novel NPU thermal management increases the inference efficiency (TOPS/Joule) by 1.33×, 1.87×, and 2× under different temperature constraints; 105 °C, 85 °C, and 70 °C, respectively, while the average accuracy drops merely from 89.0% to 85.5%.
Hussam Amrouch, Georgios Zervakis 0001, Sami Salamin, Hammam Kattan, Iraklis Anagnostopoulos, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Co-design Implications of Cost-effective On-demand Acceleration for Cloud Healthcare Analytics: The AEGLE approach
abstract
Nowadays, big data and machine learning are transforming the way we realize and manage our data. Even though the healthcare domain has recognized big data analytics as a prominent candidate, it has not yet fully grasped their promising benefits that allow medical information to be converted to useful knowledge. In this paper, we introduce AEGLE's big data infrastructure provided as a Platform as a Service. Utilizing the suite of genomic analytics from the Chronic Lymphocytic Leukaemia (CLL) use case, we show that on-demand acceleration is profitable w.r.t a pure software cloud-based solution. However, we further show that on-demand acceleration is not offered as a "free-lunch" and we provide an in-depth analysis and lessons learnt on the co-design implications to be carefully considered for enabling cost-effective acceleration at the cloud-level.
Dimosthenis Masouros, Konstantina Koliogeorgi, Georgios Zervakis 0001, Alexandra Kosvyra, Achilleas Chytas, Sotirios Xydis, Ioanna Chouvarda, Dimitrios Soudris
DATE3
2019 VADER: Voltage-Driven Netlist Pruning for Cross-Layer Approximate Arithmetic Circuits
abstract
Leveraging the inherent error resilience of a large number of application domains, approximate computing is established as an efficient design alternative to improve their energy profile. In this brief, we design energy optimal cross-layer approximate arithmetic circuits by enabling the efficient application of voltage overscaling (VOS). Departing from the conventional approaches followed today, we introduce the voltage-driven functional approximation and present the VoltAge-Driven nEtlist pRuning (VADER) framework. VADER is an automated synthesis framework that can be seamlessly integrated in any hardware design flow and implements a voltage-driven gate-level netlist pruning. Experimental evaluation shows that VADER reduces the error of the VOS application by 52% on average and delivers on average designs with 34% higher energy savings compared to state-of-the-art approximate adders and multipliers.
Georgios Zervakis 0001, Konstantina Koliogeorgi, Dimitrios Anagnostos, Nikolaos Zompakis, Kostas Siozios
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Approximate Hybrid High Radix Encoding for Energy-Efficient Inexact Multipliers
abstract
Approximate computing forms a design alternative that exploits the intrinsic error resilience of various applications and produces energy-efficient circuits with small accuracy loss. In this paper, we propose an approximate hybrid high radix encoding for generating the partial products in signed multiplications that encodes the most significant bits with the accurate radix-4 encoding and the least significant bits with an approximate higher radix encoding. The approximations are performed by rounding the high radix values to their nearest power of two. The proposed technique can be configured to achieve the desired energy-accuracy tradeoffs. Compared with the accurate radix-4 multiplier, the proposed multipliers deliver up to 56% energy and 55% area savings, when operating at the same frequency, while the imposed error is bounded by a Gaussian distribution with near-zero average. Moreover, the proposed multipliers are compared with state-of-the-art inexact multipliers, outperforming them by up to 40% in energy consumption, for similar error values. Finally, we demonstrate the scalability of our technique.
Vasileios Leon, Georgios Zervakis 0001, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.2
2018 VOSsim: A Framework for Enabling Fast Voltage Overscaling Simulation for Approximate Computing Circuits
Georgios Zervakis 0001, Fotios Ntouskas, Sotirios Xydis, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Pre-Encoded Multipliers Based on Non-Redundant Radix-4 Signed-Digit Encoding
abstract
In this paper, we introduce an architecture of pre-encoded multipliers for digital signal processing applications based on off-line encoding of coefficients. To this extend, the Non-Redundant radix-4 Signed-Digit (NR4SD) encoding technique, which uses the digit values$\lbrace -1,0,+1,+2 \rbrace$or$\lbrace -2,-1,0,+1 \rbrace$, is proposed leading to a multiplier design with less complex partial products implementation. Extensive experimental analysis verifies that the proposed pre-encoded NR4SD multipliers, including the coefficients memory, are more area and power efficient than the conventional Modified Booth scheme.
Kostas Tsoumanis, Nicholas Axelos, Nikolaos Moschopoulos, Georgios Zervakis 0001, Kiamal Z. Pekmestzi
IEEE Trans. Computers4
2016 Flexible DSP Accelerator Architecture Exploiting Carry-Save Arithmetic
abstract
Hardware acceleration has been proved an extremely promising implementation strategy for the digital signal processing (DSP) domain. Rather than adopting a monolithic application-specific integrated circuit design approach, in this brief, we present a novel accelerator architecture comprising flexible computational units that support the execution of a large set of operation templates found in DSP kernels. We differentiate from previous works on flexible accelerators by enabling computations to be aggressively performed with carry-save (CS) formatted data. Advanced arithmetic design concepts, i.e., recoding techniques, are utilized enabling CS optimizations to be performed in a larger scope than in previous approaches. Extensive experimental evaluations show that the proposed accelerator architecture delivers average gains of up to 61.91% in area-delay product and 54.43% in energy consumption compared with the state-of-art flexible datapaths.
Kostas Tsoumanis, Sotirios Xydis, Georgios Zervakis 0001, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Design-Efficient Approximate Multiplication Circuits Through Partial Product Perforation
abstract
Approximate computing has received significant attention as a promising strategy to decrease power consumption of inherently error tolerant applications. In this paper, we focus on hardware-level approximation by introducing the partial product perforation technique for designing approximate multiplication circuits. We prove in a mathematically rigorous manner that in partial product perforation, the imposed errors are bounded and predictable, depending only on the input distribution. Through extensive experimental evaluation, we apply the partial product perforation method on different multiplier architectures and expose the optimal architecture-perforation configuration pairs for different error constraints. We show that, compared with the respective exact design, the partial product perforation delivers reductions of up to 50% in power consumption, 45% in area, and 35% in critical delay. In addition, the product perforation method is compared with the state-of-the-art approximation techniques, i.e., truncation, voltage overscaling, and logic approximation, showing that it outperforms them in terms of power dissipation and error.
Georgios Zervakis 0001, Kostas Tsoumanis, Sotirios Xydis, Dimitrios Soudris, Kiamal Z. Pekmestzi
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Approximate Multiplier Architectures Through Partial Product Perforation: Power-Area Tradeoffs Analysis
abstract
Approximate computing has received significant attention as a promising strategy to decrease power consumption of inherently error-tolerant applications. Hardware approximation mainly targets arithmetic units, e.g. adders and multipliers. In this paper, we design new approximate hardware multipliers and propose the Partial Product Perforation technique, which omits a number of consecutive partial products by perforating their generation. Through extensive experimental evaluation, we apply the partial product perforation method on different multiplier architectures and expose the optimal configurations for different error values. We show that the partial product perforation delivers reductions of up to 50% in power consumption, 45% in area and 35% in critical delay. Also, the product perforation method is compared with state-of-the-art works on approximate computing that consider the Voltage Over-Scaling (VOS) and logic approximation (i.e. design of approximate compressors) techniques, outperforming them in terms of power dissipation by up to 17% and 20% on average respectively. Finally, with respect to the aforementioned gains, the error value delivered by the proposed product perforation method is smaller by 70% and 99% than the VOS and logic approximation methods respectively.
Georgios Zervakis 0001, Kostas Tsoumanis, Sotirios Xydis, Nicholas Axelos, Kiamal Z. Pekmestzi
ACM Great Lakes Symposium on VLSI1
2015 Hybrid approximate multiplier architectures for improved power-accuracy trade-offs
abstract
Approximate computing forms a promising design alternative for inherently error resilient applications, trading accuracy for power savings. In this paper, we exploit multi-level approximation, i.e. at the algorithmic, the logic and the circuit level, to design low power approximate arithmetic architectures for hardware multipliers. Motivated from the limited power savings that approximation techniques can achieve in isolation, we explore hybrid methods that apply simultaneously more than one techniques from different layers. We introduce the concept of perforation for approximate arithmetic circuit design and we explore the newly defined design space of hybrid designs showing that it leads to lower power consumption at every examined error range. To address the increased complexity of the target design space, we introduce an heuristic optimization technique and the corresponding design framework that automatically generates hybrid low-power approximate multipliers requiring a small number of design evaluations, i.e. synthesis, simulation, power and timing analysis. Through extensive experimentation, we show that the proposed techniques converge towards optimal solutions and deliver approximate designs that are always more efficient with respect to state-of-art approaches. Power savings of 11% are reported for small error bounds and more than 30% in case of more relaxed error constraints.
Georgios Zervakis 0001, Sotirios Xydis, Kostas Tsoumanis, Dimitrios Soudris, Kiamal Z. Pekmestzi
ISLPED1
2014 A segmentation-based BISR scheme
abstract
With memory estate increasing in System-On-Chips and highly integrated products, memory defects and wearout effects are the determining factor in the chip's yield loss and reliability. In this paper, a multiple cache-based Built-in Self-Repair scheme is proposed that is able to repair from the word level down to the bit level. Moreover, it is proved that the level of segmentation does not affect the repair efficiency. An exploration is then conducted to find the optimal scheme in terms of area overhead.
Georgios Zervakis 0001, Nikolaos Eftaxiopoulos-Sarris, Kostas Tsoumanis, Nicholas Axelos, Kiamal Z. Pekmestzi
ASP-DAC1
2014 FF-DICE: An 8T soft-error tolerant cell using Independent Dual Gate SOI FinFETs
abstract
In this paper we present FF-DICE (Footless-FinFET-DICE), an 8T footless storage element that exhibits soft error resilience characteristics to Single Event Upsets. The proposed cell utilises IDG (Independent Dual Gate) FinFETs to merge the functions of a typical cell's NMOS drivers and NMOS access transistors, thereby saving 33% of the transistors required for a typical DICE cell. Given the IDG FinFET dual gate mutual coupling and excellent control over the transistor conductive channel, simulations show that the proposed cell can operate as a regular memory cell, with Static Voltage Noise Margin of 341mV, Static Current Noise Margin of 15uA and provides soft-error immunity to particle strikes on single nodes, as well as considerable area savings compared to similar designs.
Nicholas Axelos, Nikolaos Eftaxiopoulos-Sarris, Georgios Zervakis 0001, Kostas Tsoumanis, Kiamal Z. Pekmestzi
IOLTS3
2013 A radiation tolerant and self-repair memory cell
abstract
In this paper a new radiation tolerant memory cell is proposed. As CMOS technologies scale down, existing hardened cells, which are technology dependent, become more and more vulnerable to radiation effects. We tested some of the most known and effective hardened cells using the SPICE simulator LTspice but no one proved to ensure data integrity. The proposed cell consists of three standard 6T cells and proves to be 100% radiation tolerant in any technology, having however an expected area and power overhead comparing to the 6T and DICE cells. According to simulation results, these overheads are proportional to the number of transistors used, but the read time when no error has occurred and the write time are shorter.
Nikolaos Eftaxiopoulos-Sarris, Georgios Zervakis 0001, Kostas Tsoumanis, Kiamal Z. Pekmestzi
IOLTS2