EDBT 2026 Demo / reviewers in the wild / expert
Andrea Calimera
dblp:36/291
· DBLP profile ↗
82ranked-venue papers
15as first author
21since 2021 · last 2026
0000-0001-5881-3811ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 73 · 15 first-author · 12 since 2021Software engineering, systems software and programming languages · 17 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 4 since 2021Computer networks · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Runtime Feature Compression for Adaptive Keyword Spotting on Embedded SystemsabstractVoice user interfaces rely on keyword spotting (KWS) to detect wake-word commands, enabling low-power devices to switch from drowsy to active states and initiate more complex tasks. In embedded systems, KWS combines handcrafted acoustic features extraction with lightweight neural network classifiers to achieve accurate detection within strict resource constraints. Adapting KWS to time-varying energy budgets requires optimization strategies that operate at runtime. Most existing approaches adjust the complexity of the neural model but overlook that a substantial amount of latency, and thus energy consumption, is due to feature extraction, which remains unaffected by model scaling. This work introducesRuntime Feature Compression(RFC), a dynamic rescaling strategy that modulates the workload of the entire KWS pipeline. RFC promotes thehop-lengthparameter of the Short-Time Fourier Transform as a runtime control knob to adjust the number of time frames in speech features, allowing a single model to operate across multiple latency modes. To support this flexibility, we introduce two training-time techniques:HopAugment, a data augmentation scheme that exposes the model to variable hop lengths during training, andMasked Layers, which preserve consistent activation statistics during training and inference under compressed feature settings. Evaluations on four KWS datasets using the TC-ResNet model family show that RFC outperforms model scaling techniques, offering a wider range of latency-accuracy trade-offs. RFC achieves up to 31.8% lower latency without accuracy degradation, or up to 0.30% higher accuracy within equivalent latency bounds. That proves RFC improves adaptability in energy-constrained IoT speech interfaces. A set of ablation studies further demonstrates the robustness of RFC by evaluating the role of its training components, batching strategies, ability to preserve accuracy with a shared weight set, scalability across operating modes, and applicability to different model architectures. Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi |
IEEE Internet Things J. | 2 |
| 2026 | FedAGF: Adaptive Concurrency via Gradient Feedback for Mitigating Extreme Label Skew in Budget-Constrained Federated LearningabstractCross-device Federated Learning (FL) enables large fleets of distributed edge devices to collaboratively train a global classification model without sharing their data. The resulting training quality is strongly limited byextreme label skew, a condition where each device holds samples from a subset of the target classes. In such cases, local updates become biased toward local distributions, slowing down convergence and degrading the global model accuracy. These effects get critical when devices operate under constrained energy budgets that restrict their participation to a limited number of synchronization rounds, further reducing the achievable accuracy. To overcome these limitations, we introduce Federated Learning with Adaptive Concurrency via Gradient Feedback (FedAGF), a control policy that dynamically adjusts the number of devices selected for synchronization throughout training. FedAGF adapts to training dynamics by monitoring global model updates and elevating participation whenever progress slows. This adaptive mechanism balances accuracy improvement with efficient use of devices’ energy budgets, allocating resources when they provide the greatest benefit to convergence. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that FedAGF effectively mitigates extreme label skew, achieving up to 14.14% higher accuracy than state-of-the-art FL methods and enabling efficient and scalable training, even under skewed label distributions and resource constraints. Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Privacy-Preserving Federated Learning for Household Characteristic IdentificationabstractThis work presents a privacy-preserving training framework for household characteristic identification from electricity consumption data. The proposed framework integrates two main components: (i) a synthetic data generation pipeline capable of replicating realistic energy traces from diverse family compositions, capturing fine-grained sociodemographic attributes such as household size, employment status, age groups, and home occupancy patterns; (ii) a training strategy based on Federated Learning (FL) secured with homomorphic encryption, enabling collaborative model training while preserving data ownership. Our synthetic dataset enables the performance assessment of different training scenarios, including siloed model training by individual energy utilities and secure collaboration via FL. Experimental results show that siloed training leads to inconsistent and suboptimal performance, while privacy-preserving FL achieves accuracy comparable to conventional centralized training—an ideal yet not viable option due to data regulation constraints. Our findings highlight the effectiveness of FL as a secure solution for collaborative sociodemographic profiling in smart grids. Erich Malan, Claudia De Vizia, Marco Castangia, Valentino Peluso, Andrea Calimera, Enrico Macii |
COMPSAC | 5 |
| 2025 | Gradient-Aware Participation for Energy Reduction in Federated Learning with Extreme Label SkewabstractFederated Learning (FL) enables distributed clients to train a global classification model collaboratively while preserving data privacy. A major challenge in FL is ensuring efficient training with limited computing and communication resources, especially when clients’ datasets contain samples from a restricted subset of target classes, a problem known as extreme label skew. Under such a condition, model updates from clients are biased toward their local data distributions, resulting in slow convergence and increased energy consumption due to the need for additional training rounds. This paper introduces FL with Gradient-Aware Participation (FedGAP), a novel strategy aimed at reducing energy consumption while preserving model accuracy even with extreme label skew. FedGAP dynamically adjusts the cohort size, i.e., the number of participating clients per training round, based on the evolution of the global model’s pseudo-gradient. By detecting stagnant phases where progress toward convergence stalls, FedGAP increases the cohort size to escape suboptimal regions and accelerate learning, thereby minimizing the waste of resources. Experiments on CIFAR-10 and CIFAR-100 demonstrate that FedGAP achieves up to 2.74× greater energy efficiency compared to state-of-the-art methods without compromising accuracy. Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii |
IJCNN | 3 |
| 2024 | Private Tensor Freezing for an Efficient Federated Learning with Homomorphic EncryptionabstractFederated Learning (FL) is a privacy-preserving machine learning strategy where distributed clients share updates of locally trained models with a central server. The server aggregates those updates to refine a global version of the model without accessing the clients' data. Even if the raw data never leaves clients, adversarial attacks on the server side can still extract sensitive information from the transmitted model updates. Homomorphic Encryption (HE) offers a robust solution to this privacy concern: clients send encrypted model updates to the server; the server operates the aggregation of the received updates without having to decrypt. Unfortunately, HE leads to a substantial computational and communication overhead on both the client and the server side, preventing the adoption in practical, real-life applications. In this work, we introduce a training option to mitigate this problem by promoting Private Tensor Freezing (PTF), a progressive and secure gating scheme by which the number of model tensors involved in the training and synchronization stages gradually reduces over time, alleviating (i) the pressure of HE encryption/decryption on the client side, (ii) the communication volumes from/to the server, and (iii) the computing complexity of the aggregation stage on the server side. Experiments on four image classification benchmarks trained within a state-of-the-art FL framework secured with CKKS encryption reveal the effectiveness of PTF: up to 37.4% in data volume reduction, 35.5% less compute time on the client side, and 36.4% less compute time on the server side. Valentino Peluso, Erich Malan, Andrea Calimera, Enrico Macii |
ICCD | 3 |
| 2024 | Automatic Layer Freezing for Communication Efficiency in Cross-Device Federated LearningabstractFederated learning (FL) is a collaborative machine learning paradigm where network-edge clients train a global model under the orchestration of a central server. Unlike traditional distributed learning, each participating client keeps its data locally, ensuring privacy protection by default. However, state-of-the-art FL implementations suffer from massive information exchange between clients and the server. This issue prevents the adoption in constrained environments, typical of the Internet of Things domain, where the communication bandwidth and the energy budget are severely limited. To achieve higher efficiency at scale, the future of FL calls for additional optimizations to reach high-quality learning capability with lower communication pressure. To address this challenge, we propose automatic layer freezing (ALF), an embedded mechanism that gradually drops a growing portion of the model out of the training and synchronization phases of the learning loop, reducing the volume of exchanged data with the central server. ALF monitors the evolution of model updates and identifies layers that have reached a stable representation, where further weight updates would have minimal impact on accuracy. By freezing these layers, ALF achieves substantial savings in communication bandwidth and energy consumption. The proposed implementation of the ALF mechanism is compatible with any FL strategy, requiring minimal effort and without interfering with existing optimizations. The extensive experiments conducted using a representative set of FL strategies applied to two image classification tasks show that ALF improves the communication efficiency of the baseline FL implementations, ensuring up to 83.91% of data volume savings with no or marginal losses of accuracy. Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi |
IEEE Internet Things J. | 3 |
| 2023 | Enabling DVFS Side-Channel Attacks for Neural Network Fingerprinting in Edge Inference ServicesabstractThe Inference-as-a-Service (IaaS) delivery model provides users access to pre-trained deep neural networks while safeguarding network code and weights. However, IaaS is not immune to security threats, like side-channel attacks (SCAs), that exploit unintended information leakage from the physical characteristics of the target device. Exposure to such threats grows when IaaS is deployed on distributed computing nodes at the edge. This work identifies a potential vulnerability of low-power CPUs that facilitates stealing the deep neural network architecture without physical access to the hardware or interference with the execution flow. Our approach relies on a Dynamic Voltage and Frequency Scaling (DVFS) side-channel attack, which monitors the CPU frequency state during the inference stages. Specifically, we introduce a dedicated load-testing methodology that imprints distinguishable signatures of the network on the frequency traces. A machine learning classifier is then used to infer the victim architecture. Experimental results on two commercial ARM Cortex-A CPUs, the A72 and A57, demonstrate the attack can identify the target architecture from a pool of 12 convolutional neural networks with an average accuracy of 98.7% and 92.4% Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii |
ISLPED | 3 |
| 2023 | Dynamic ConvNets on Tiny Devices via Nested SparsityabstractThis work introduces a new training and compression pipeline to build nested sparse convolutional neural networks (ConvNets), a class of dynamic ConvNets suited for inference tasks deployed on resource-constrained devices at the edge of the Internet of Things. A nested sparse ConvNet consists of a single ConvNet architecture, containing$N$sparse subnetworks with nested weights subsets, like a Matryoshka doll, and can trade accuracy for latency at runtime, using the model sparsity as a dynamic knob. To attain high accuracy at training time, we propose a gradient masking technique that optimally routes the learning signals across the nested weight subsets. To minimize the storage footprint and efficiently process the obtained models at inference time, we introduce a new sparse matrix compression format with dedicated compute kernels that fruitfully exploit the characteristic of the nested weights subsets. Tested on image classification and object detection tasks on an off-the-shelf ARM-M7 microcontroller unit (MCU), nested sparse ConvNets outperform variable-latency solutions naively built assembling single sparse models trained as stand-alone instances, achieving 1) comparable accuracy; 2) remarkable storage savings; and 3) high performance. Moreover, when compared to state-of-the-art dynamic strategies, such as dynamic pruning and layer width scaling, nested sparse ConvNets turn out to be Pareto optimal in the accuracy versus latency space. Matteo Grimaldi, Luca Mocerino, Antonio Cipolletta, Andrea Calimera |
IEEE Internet Things J. | 4 |
| 2023 | Efficient Deep Learning Models for Privacy-Preserving People Counting on Low-Resolution Infrared ArraysabstractUltralow-resolution infrared (IR) array sensors offer a low cost, energy efficient, and privacy-preserving solution for people counting, with applications, such as occupancy monitoring and visitor flow analysis in private and public spaces. Previous work has shown that deep learning (DL) can yield superior performance on this task. However, the literature was missing an extensive comparative analysis of various efficient DL architectures for IR array-based people counting, that considers not only their accuracy but also the cost of deploying them on memory- and energy-constrained Internet of Things (IoT) edge nodes. Such analysis is key for system designers, since it helps them select the most appropriate DL model given the constraints of their target hardware. In this work, we address this need by comparing six different DL architectures on a novel data set composed of IR images collected from a commercial$8\times8$array, which we made openly available. With a wide architectural exploration of each model type, we obtain a rich set of Pareto-optimal solutions, spanning cross-validated balanced accuracy scores in the 55.70%–82.70% range. When deployed on a commercial microcontroller (MCU) by STMicroelectronics, the STM32L4A6ZG, these models occupy 0.41–9.28kB of memory, and require 1.10–7.74 ms per inference, while consuming 17.18–$120.43 \mu \text{J}$of energy. Our models are significantly more accurate than a previous deterministic method (up to +39.9%), while being up to$3.53\times $faster and more energy efficient. So, our work serves also as a demonstration that DL can not only achieve higher accuracy but also higher efficiency compared to classic algorithms for this type of task. Further, our models’ accuracy is comparable to state-of-the-art DL solutions on similar resolution sensors, despite a much lower complexity. All our models enable continuous, real-time inference on an MCU-based IoT node, with years of autonomous operation without battery recharging. Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Internet Things J. | 6 |
| 2022 | Privacy-preserving Social Distance Monitoring on Microcontrollers with Low-Resolution Infrared Sensors and CNNsabstractLow-resolution infrared (IR) array sensors offer a low-cost, low-power, and privacy-preserving alternative to optical cameras and smartphones/wearables for social distance monitoring in indoor spaces, permitting the recognition of basic shapes, without revealing the personal details of individuals. In this work, we demonstrate that an accurate detection of social distance violations can be achieved processing the raw output of a 8x8 IR array sensor with a small-sized Convolutional Neural Network (CNN). Furthermore, the CNN can be executed directly on a Microcontroller (MCU)-based sensor node.With results on a newly collected open dataset, we show that our best CNN achieves 86.3% balanced accuracy, significantly outperforming the 61% achieved by a state-of-the-art deterministic algorithm. Changing the architectural parameters of the CNN, we obtain a rich Pareto set of models, spanning 70.5-86.3% accuracy and 0.18-75k parameters. Deployed on a STM32L476RGMCU, these models have a latency of 0.73-5.33ms, with an energy consumption per inference of 9.38-68.57$\mu$J. Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISCAS | 6 |
| 2022 | Energy-Quality Scalable Monocular Depth Estimation on Low-Power CPUsabstractThe recent advancements in deep learning have demonstrated that inferring high-quality depth maps from a single image has become feasible and accurate, thanks to convolutional neural networks (CNNs), but how to process such compute- and memory-intensive models on portable and low-power devices remains a concern. Dynamic energy-quality scaling is an interesting yet less explored option in this field. It can improve efficiency through opportunistic computing policies where performances are boosted only when needed, achieving on average substantial energy savings. Implementing such a computing paradigm encompasses the availability of a scalable inference model, which is the target of this work. Specifically, we describe and characterize the design of an energy-quality scalable pyramidal network (EQPyD-Net), a lightweight CNN capable of modulating at runtime the computational effort with minimal memory resources. We describe the architecture of the network and the optimization flow, covering the important aspects that enable the dynamic scaling, namely, the optimized training procedures, the compression stage via fixed-point quantization, and the code optimization for the deployment on commercial low-power CPUs adopted in the edge segment. To assess the effect of the proposed design knobs, we evaluated the prediction quality on the standard KITTI data set and the energy and memory resources on the ARM Cortex-A53 CPU. The collected results demonstrate the flexibility of the proposed network and its energy efficiency. EQPyD-Net can be shifted across five operating points, ranging from a maximum accuracy of 82.2% with 0.4 Frame/J and up to 92.6% of energy savings with 6.1% of accuracy loss, still keeping a compact memory footprint of 5.2 MB for the weights and 38.3 MB (in the worst case) for the processing. Antonio Cipolletta, Valentino Peluso, Andrea Calimera, Matteo Poggi, Fabio Tosi, Filippo Aleotti, Stefano Mattoccia |
IEEE Internet Things J. | 3 |
| 2022 | Monocular Depth Perception on Microcontrollers for Edge ApplicationsabstractDepth estimation is crucial in several computer vision applications, and a recent trend in this field aims at inferring such a cue from a single camera. Unfortunately, despite the compelling results achieved, state-of-the-art monocular depth estimation methods are computationally demanding, thus precluding their practical deployment in several application contexts characterized by low-power constraints. Therefore, in this paper, we propose a lightweight Convolutional Neural Network based on a shallow pyramidal architecture, referred to as$\mu $PyD-Net, enabling monocular depth estimation on microcontrollers. The network is trained in a peculiar self-supervised manner leveraging proxy labels obtained through a traditional stereo algorithm. Moreover, we propose optimization strategies aimed at performing computations with quantized 8-bit data and map the high-level description of the network to low-level layers optimized for the target microcontroller architecture. Exhaustive experimental results on standard datasets and an in-depth evaluation with a device belonging to the popular Arm Cortex-M family confirm that obtaining sufficiently accurate monocular depth estimation on microcontrollers is feasible. To the best of our knowledge, our proposal is the first one enabling such remarkable achievement, paving the way for the deployment of monocular depth cues onto the tiny end-nodes of distributed sensor networks. Valentino Peluso, Antonio Cipolletta, Andrea Calimera, Matteo Poggi, Fabio Tosi, Filippo Aleotti, Stefano Mattoccia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Human Activity Recognition on Microcontrollers with Quantized and Adaptive Deep Neural NetworksabstractHuman Activity Recognition (HAR) based on inertial data is an increasingly diffused task on embedded devices, from smartphones to ultra low-power sensors. Due to the high computational complexity of deep learning models, most embedded HAR systems are based on simple and not-so-accurate classic machine learning algorithms. This work bridges the gap between on-device HAR and deep learning, proposing a set of efficient one-dimensional Convolutional Neural Networks (CNNs) that can be deployed on general purpose microcontrollers (MCUs). Our CNNs are obtained combining hyper-parameters optimization with sub-byte and mixed-precision quantization, to find good trade-offs between classification results and memory occupation. Moreover, we also leverage adaptive inference as an orthogonal optimization to tune the inference complexity at runtime based on the processed input, hence producing a more flexible HAR system. With experiments on four datasets, and targeting an ultra-low-power RISC-V MCU, we show that (i) we are able to obtain a rich set of Pareto-optimal CNNs for HAR, spanning more than 1 order of magnitude in terms of memory, latency, and energy consumption; (ii) thanks to adaptive inference, we can derive >20 runtime operating modes starting from a single CNN, differing by up to 10% in classification scores and by more than 3× in inference complexity, with a limited memory overhead; (iii) on three of the four benchmarks, we outperform all previous deep learning methods, while reducing the memory occupation by more than 100×. The few methods that obtain better performance (both shallow and deep) are not compatible with MCU deployment; (iv) all our CNNs are compatible with real-time on-device HAR, achieving an inference latency that ranges between 9 μs and 16 ms. Their memory occupation varies in 0.05–23.17 kB, and their energy consumption in 0.05 and 61.59 μJ, allowing years of continuous operation on a small battery supply. Francesco Daghero, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2021 | Ultra-compact binary neural networks for human activity recognition on RISC-V processorsabstractHuman Activity Recognition (HAR) is a relevant inference task in many mobile applications. State-of-the-art HAR at the edge is typically achieved with lightweight machine learning models such as decision trees and Random Forests (RFs), whereas deep learning is less common due to its high computational complexity. In this work, we propose a novel implementation of HAR based on deep neural networks, and precisely on Binary Neural Networks (BNNs), targeting low-power general purpose processors with a RISC-V instruction set. BNNs yield very small memory footprints and low inference complexity, thanks to the replacement of arithmetic operations with bit-wise ones. However, existing BNN implementations on general purpose processors impose constraints tailored to complex computer vision tasks, which result in over-parametrized models for simpler problems like HAR. Therefore, we also introduce a new BNN inference library, which targets ultra-compact models explicitly. With experiments on a single-core RISC-V processor, we show that BNNs trained on two HAR datasets obtain higher classification accuracy compared to a state-of-the-art baseline based on RFs. Furthermore, our BNN reaches the same accuracy of a RF with either less memory (up to 91%) or more energy-efficiency (up to 70%), depending on the complexity of the features extracted by the RF. Francesco Daghero, Daniele Jahier Pagliari, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino |
CF | 7 |
| 2021 | On The Efficiency of Sparse-Tiled Tensor Graph Processing For Low Memory UsageabstractThe memory space taken to host and process large tensor graphs is a limiting factor for embedded ConvNets. Even though many data-driven compression pipelines have proven their efficacy, this work shows there is still room for optimization at the intersection with compute-oriented optimizations. We demonstrate that tensor pruning via weight sparsification can cooperate with a model-agnostic tiling strategy, leading ConvNets towards a new feasible region of the solution space. The collected results show for the first time fast versions of MobileNets deployed at full scale on an ARM M7 core with 512KB of RAM and 2MB of FLASH memory. Antonio Cipolletta, Andrea Calimera |
DAC | 2 |
| 2021 | Dataflow Restructuring for Active Memory Reduction in Deep Neural NetworksabstractThe volume reduction of the activation maps produced by the hidden layers of a Deep Neural Network (DNN) is a critical aspect in modern applications as it affects the on-chip memory utilization, the most limited and costly hardware resource. Despite the availability of many compression methods that leverage the statistical nature of deep learning to approximate and simplify the inference model, e.g., quantization and pruning, there is room for deterministic optimizations that instead tackle the problem from a computational view. This work belongs to this latter category as it introduces a novel method for minimizing the active memory footprint. The proposed technique, which is data-, model-, compiler-, and hardware-agnostic, does implement a functional-preserving, automated graph restructuring where the memory peaks are suppressed and distributed over time, leading to flatter profiles with less memory pressure. Results collected on a representative class of Convolutional DNNs with different topologies, from Vgg16 and SqueezeNetV1.1 to the recent MobileNetV2, ResNet18, and InceptionV3, provide clear evidence of applicability, showing remarkable memory savings (62.9% on average) with low computational overhead (8.6% on average). Antonio Cipolletta, Andrea Calimera |
DATE | 2 |
| 2021 | ACME: An Energy-Efficient Approximate Bus Encoding for I2CabstractIn ultra low power systems with many peripherals, off-chip serial interconnects contribute significantly to the total energy budget. Leveraging the error-resilience characteristics of many embedded applications, the approximate computing paradigm has been applied to serial bus encodings to reduce interconnect consumption. However, the power model considered in previous works was purely capacitive. Accordingly, the objective of these approximate encodings was simply to reduce the transition count. While this works well for most bus standards, one notable exception is represented by I2C, whose open-drain physical connection makes the static energy consumed by logic-0 values on the bus extremely relevant. In this work, we propose ACME, the first approximate serial bus encoding targeting specifically I2C connections. With a simple encoding/decoding scheme, ACME concurrently reduces both the static and dynamic energy on the bus by maximizing the number of logic-1 values in codewords, while simultaneously reducing transitions. Using an accurate bus model and realistic capacitance and resistance values selected according to the I2C standard, we show that our encoding outperforms state-of-the-art solutions and reduces the total energy consumption on the bus by 57% on average, with an error smaller than 0.1%. Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2021 | Adaptive Random Forests for Energy-Efficient Inference on MicrocontrollersabstractRandom Forests (RFs) are widely used Machine Learning models in low-power embedded devices, due to their hardware friendly operation and high accuracy on practically relevant tasks. The accuracy of a RF often increases with the number of internal weak learners (decision trees), but at the cost of a proportional increase in inference latency and energy consumption. Such costs can be mitigated considering that, in most applications, inputs are not all equally difficult to classify. Therefore, a large RF is often necessary only for (few) hard inputs, and wasteful for easier ones. In this work, we propose an early-stopping mechanism for RFs, which terminates the inference as soon as a high-enough classification confidence is reached, reducing the number of weak learners executed for easy inputs. The early-stopping confidence threshold can be controlled at runtime, in order to favor either energy saving or accuracy. We apply our method to three different embedded classification tasks, on a single-core RISC-V microcontroller, achieving an energy reduction from 38% to more than 90% with a drop of less than 0.5% in accuracy. We also show that our approach outperforms previous adaptive ML methods for RFs. Francesco Daghero, Alessio Burrello, Luca Benini, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
VLSI-SoC | 5 |
| 2021 | AdapTTA: Adaptive Test-Time Augmentation for Reliable Embedded ConvNetsabstractConvolutional Neural Networks (ConvNets) are trained offline using the few available data and may therefore suffer from substantial accuracy loss when ported on the field, where unseen input patterns received under unpredictable external conditions can mislead the model. Test-Time Augmentation (TTA) techniques aim to alleviate such common side effect at inference-time, first running multiple feed-forward passes on a set of altered versions of the same input sample, and then computing the main outcome through a consensus of the aggregated predictions. Unfortunately, the implementation of TTA on embedded CPUs introduces latency penalties that limit its adoption on edge applications. To tackle this issue, we propose AdapTTA, an adaptive implementation of TTA that controls the number of feed-forward passes dynamically, depending on the complexity of the input. Experimental results on state-of-the-art ConvNets for image classification deployed on a commercial ARM Cortex-A CPU demonstrate AdapTTA reaches remarkable latency savings, from $1.40 \times$ to $2.21 \times$, and hence a higher frame rate compared to static TTA, still preserving the same accuracy gain. Luca Mocerino, Roberto Giorgio Rizzo, Valentino Peluso, Andrea Calimera, Enrico Macii |
VLSI-SoC | 4 |
| 2021 | Manufacturing as a Data-Driven Practice: Methodologies, Technologies, and ToolsabstractIn recent years, the introduction and exploitation of innovative information technologies in industrial contexts have led to the continuous growth of digital shop floor environments. The new Industry 4.0 model allows smart factories to become very advanced IT industries, generating an ever-increasing amount of valuable data. As a consequence, the necessity of powerful and reliable software architectures is becoming prominent along with data-driven methodologies to extract useful and hidden knowledge supporting the decision-making process. This article discusses the latest software technologies needed to collect, manage, and elaborate all data generated through innovative Internet-of-Things (IoT) architectures deployed over the production line, with the aim of extracting useful knowledge for the orchestration of high-level control services that can generate added business value. This survey covers the entire data life cycle in manufacturing environments, discussing key functional and methodological aspects along with a rich and properly classified set of technologies and tools, useful to add intelligence to data-driven services. Therefore, it serves both as a first guided step toward the rich landscape of the literature for readers approaching this field and as a global yet detailed overview of the current state of the art in the Industry 4.0 domain for experts. As a case study, we discuss, in detail, the deployment of the proposed solutions for two research project demonstrators, showing their ability to mitigate manufacturing line interruptions and reduce the corresponding impacts and costs. Tania Cerquitelli, Daniele Jahier Pagliari, Andrea Calimera, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Massimo Poncino |
Proc. IEEE | 3 |
| 2021 | Fast and Accurate Inference on Microcontrollers With Boosted Cooperative Convolutional Neural Networks (BC-Net)abstractArithmetic precision scaling is mandatory to deploy Convolutional Neural Networks (CNNs) on resource-constrained devices such as microcontrollers (MCUs), and quantization via fixed-point or binarization are the most adopted techniques today. Despite being born by the same concept of bit-width lowering, these two strategies differ substantially each other, and hence are often conceived and implemented separately. However, their joint integration is feasible and, if properly implemented, can bring to large savings and high processing efficiency. This work elaborates on this aspect introducing a boosted collaborative mechanism that pushes CNNs towards higher performance and more predictive capability. Referred as BC-Net, the proposed solution consists of a self-adaptive conditional scheme where a lightweight binary net and an 8-bit quantized net are trained to cooperate dynamically. Experiments conducted on four different CNN benchmarks deployed on off-the-shelf boards powered with the MCUs of the Cortex-M family by ARM show that BC-Nets outperform classical quantization and binarization when applied as separate techniques (up to 81.49% speed-up and up to 3.8% of accuracy improvement). The comparative analysis with a previously proposed cooperative method also demonstrates BC-Nets achieve substantial savings in terms of both performance (+19%) and accuracy (+3.45%). Luca Mocerino, Andrea Calimera |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | Optimization Tools for ConvNets on the EdgeabstractThe shift of Convolutional Neural Networks (ConvNets) into low-power devices with limited compute and memory resources calls for cross-layer strategies spanning from hardware to software optimization. This work answers to this need, presenting a collection of tools for efficient deployment of ConvNets on the edge, Valentino Peluso, Enrico Macii, Andrea Calimera |
VLSI-SOC | 3 |
| 2020 | Corrigendum to"Approximate error detection-correction for efficient adaptive voltage Over-Scaling"[Integration 63 (2018) 220-231]
Roberto Giorgio Rizzo, Andrea Calimera, Jun Zhou 0017 |
Integr. | 2 |
| 2020 | Logic Synthesis of Pass-Gate Logic Circuits With Emerging Ambipolar TechnologiesabstractEmerging devices and new ultrascaled silicon transistors have shown disruptive electrical and functional properties that might bring digital hardware to the next level. The key issue today concerns their integration. Even though the classical complementary logic style is the most intuitive option, other strategies such as pass-transistors that were discarded in the past because they did not fit silicon MOSFETs logic should be reconsidered. Obviously, the assessment of such alternatives requires customized CAD tools and optimization engines. The objective of this paper is to introduce a synthesis and optimization flow for pass-gate logic circuits mapped onto emerging ambipolar technologies. As main contributions we propose: 1) a novel EXNOR-based decomposition technique that fully exploits do not care conditions to generate compact logic function representations and 2) a dedicated one-pass synthesis flow where optimization and technology mapping are concurrently run on a common data structure, the reduced ordered pass-diagram. Experimental results demonstrate that the proposed flow outperforms existing synthesis tools by achieving more compact circuit representations with 8.5× less devices and about 8× shallower structures (on average), while still yielding lower CPU times. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Energy-Efficient Convolutional Neural Networks via Recurrent Data ReuseabstractDeep learning (DL) algorithms have substantially improved in terms of accuracy and efficiency. Convolutional Neural Networks (CNNs) are now able to outperform traditional algorithms in computer vision tasks such as object classification, detection, recognition, and image segmentation. They represent an attractive solution for many embedded applications which may take advantage from machine-learning at the edge. Needless to say, the key to success lies under the availability of efficient hardware implementations which meet the stringent design constraints.Inspired by the way human brains process information, this paper presents a method that improves the processing efficiency of CNNs leveraging their repetitiveness. More specifically, we introduce (i) a clustering methodology that maximizes weights/activation reuse, and (ii) the design of a heterogeneous processing element which integrates a Floating-Point Unit (FPU) with an associative memory that manages recurrent patterns. The experimental analysis reveals that the proposed method achieves substantial energy savings with low accuracy loss, thus providing a practical design option that might find application in the growing segment of edge-computing. Luca Mocerino, Valerio Tenace, Andrea Calimera |
DATE | 3 |
| 2019 | Enabling Energy-Efficient Unsupervised Monocular Depth Estimation on ARMv7-Based PlatformsabstractThis work deals with the implementation of energy-efficient monocular depth estimation using a low-cost CPU for low-power embedded systems. It first describes the PyD-Net depth estimation network, which consists of a lightweight CNN able to approach state-of-the-art accuracy with ultra-low resource usage. Then it proposes an accuracy-driven complexity reduction strategy based on a hardware-friendly fixed-point quantization. Finally, it introduces the low-level optimization enabling effective use of integer neural kernels. The objective is threefold: (i) prove the efficiency of the new quantization flow on a depth estimation network, that is, the capability to retaining the accuracy reached by floating-point arithmetic using 16- and 8-bit integers, (ii) demonstrate the portability of the quantized model into a general-purpose 32-bit RISC architecture of the ARM Cortex family, (iii) quantify the accuracy-energy tradeoff of unsupervised monocular estimation to establish its use in the embedded domain. The experiments have been run on a Raspberry PI board powered by a Broadcom BCM2837 chipset. A parametric analysis conducted over the KITTI date-set shows marginal accuracy loss with 16-bit (8-bit) integers and energy savings up to 6.55× (9.23×) w.r.t. floating-point. Compared to high-end CPU and GPU the proposed solution improves scalability. Valentino Peluso, Antonio Cipolletta, Andrea Calimera, Matteo Poggi, Fabio Tosi, Stefano Mattoccia |
DATE | 3 |
| 2019 | SAID: A Supergate-Aided Logic Synthesis Flow for Memristive CrossbarsabstractA Memristor is a two-terminal device that can serve as a non-volatile memory element with built-in logic capabilities. Arranged in a crossbar structure, memristive arrays allow to represent complex Boolean logic functions that adhere to the logic-in-memory paradigm, where data and logic gates are glued together on the same piece of hardware. Needless to say, novel and ad-hoc CAD solutions are required to achieve practical and feasible hardware implementations. Existing techniques aim at optimal mapping strategies that account for Boolean logic functions described by means of 2-input NOR and NOT gates, thus overlooking the optimization capabilities that a smart and dedicated technology-aware logic synthesis can provide. In this paper, we introduce a novel library-free supergate-aided (SAID) logic synthesis approach with a dedicated mapping strategy tailored on MAGIC crossbars. Supergates are obtained with a Look-Up Table (LUT)-based synthesis that splits a complex logic network into smaller Boolean functions. Those functions are then mapped on the crossbar array as to minimize latency. The proposed SAID flow allows to (i) maximize supergate-level parallelism, thus reducing the total number of computing cycles, and (ii) relax mapping constraints, allowing an easy and fast mapping of Boolean functions on memristive crossbars. Experimental results obtained on several benchmarks from ISCAS'85 and IWLS'93 suites demonstrate that our solution is capable to outperform other state-of-the-art techniques in terms of speedup (3.89× in the best case), at the expense of a very low area overhead. Valerio Tenace, Roberto Giorgio Rizzo, Debjyoti Bhattacharjee, Anupam Chattopadhyay, Andrea Calimera |
DATE | 5 |
| 2019 | Arbitrary-Precision Convolutional Neural Networks on Low-Power IoT ProcessorsabstractThe deployment of Convolutional Neural Networks (CNNs) on resource-constrained IoT devices calls for accurate model re-sizing and optimization. Among the proposed compression strategies, n-ary fixed-point quantization has proven effective in reducing both computational effort and memory footprint with no (or limited) accuracy loss. However, its use requires custom components and special memory allocation strategies which are not available and burdensome to implement on low-power/low-cost cores. In order to bridge this gap, this work introduces Virtual Quantization (VQ), a hardware-friendly compression method which allows to implement equivalent n ary CNNs on general purpose instruction-set architectures. The proposed VQ framework is validated for the IoT family of ARM MCUs (ARM Cortex-M) and tested with three different real-life applications (i.e. Image Classification, Keyword Spotting, Facial Expression Recognition). Valentino Peluso, Matteo Grimaldi, Andrea Calimera |
VLSI-SoC | 3 |
| 2018 | All-digital embedded meters for on-line power estimationabstractModern low power designs use multiple knobs for concurrent dynamic and leakage power optimization; supply voltage and threshold voltage are the most adopted. An efficient control of these knobs needs management policies aware of the power breakdown. This implies the availability of smart on-chip strategies for dynamic and leakage power estimation at runtime. In this paper, we address this issue proposing the implementation of embedded dynamic/static power meters that use an optimized regression model fed with data collected from in-situ activity monitors. The number of sensors, their bitwidth and optimal placement are obtained through an automated design flow. The methodology works for general logic and applies not just to processor cores, but also to application-specific designs. We apply our solution to a representative class of benchmarks, showing that it can achieve an average estimation error smaller than 3%, with limited area and power overheads. Daniele Jahier Pagliari, Valentino Peluso, Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2018 | Energy-performance design exploration of a low-power microprogrammed deep-learning acceleratorabstractThis paper presents the design space exploration of a novel microprogrammable accelerator in which PEs are connected with a Network-on-Chip and benefit from low-power features enabled through a practical implementation of a Dual-Vddassignment scheme. An analytical model, fitted with postlayout data obtained with a 28nm FDSOI design kit, returns implementations with optimal energy-performance tradeoff by taking into consideration all the key design-space variables. The obtained Pareto analysis helps us infer optimization rules aimed at improving quality of design. Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto |
DATE | 4 |
| 2018 | Scalable-effort ConvNets for multilevel classificationabstractThis work introduces the concept of scalable-effort Convolutional Neural Networks (ConvNets), an effort-accuracy scalable model for classification of data at multilevel abstraction. Scalable-effort ConvNets are able to adapt at run-time to the complexity of the classification problem, i.e. the level of abstraction defined by the application (or context), and reach a given classification accuracy with minimal computational effort. The mechanism is implemented using a single-weight scalable-precision model rather than an ensemble of quantized weight models; this makes the proposed strategy highly flexible and particularly suited for embedded architectures with limited resource availability. The paper describes (i) a hardware/software vertical implementation of scalable-precision multiply&accumulate arithmetic, (ii) an accuracy-constrained heuristic that delivers near-optimal layer-by-layer precision mapping at a predefined level of abstraction. It also reports the validation for three state-of-the-art nets, i.e. AlexNet, SqueezeNet and MobileNet, trained and tested with ImageNet. Collected results show scalable-effort ConvNets guarantee flexibility and substantial savings: 47.07% computational effort reduction at minimum accuracy, or 30.6% accuracy improvement at maximum effort w.r.t. standard flat ConvNets (average over the three benchmarks for high-level classification). Valentino Peluso, Andrea Calimera |
ICCAD | 2 |
| 2018 | Weak-MAC: Arithmetic Relaxation for Dynamic Energy-Accuracy Scaling in ConvNetsabstractThis work introduces Weak-MAC (WeMAC), a precision scaling strategy for fixed point Multiply&Accumulate operations. By leveraging an algorithmic relaxation of multiplication, it improves the efficiency of deep convolutional neural networks. WeMAC can be run on a software-programmable 8-bit MAC unit and is retraining free, namely, it makes ConvNet achieving reasonable inference accuracy even without retraining. Simulation results conducted on three ConvNets trained over well recognized data-set (SVHN, CIFAR-10 and CIFAR-100) demonstrate WeMAC enables new Pareto points in the energy-accuracy tradeoff of HW accelerators. With 25% more energy efficiency w.r.t. full precision, WeMAC represents a middle way between the low energy consumption of 8-bit arithmetic and the high accuracy of 16-bit arithmetic. Valentino Peluso, Andrea Calimera |
ISCAS | 2 |
| 2018 | Multiplication by Inference using Classification Trees: A Case-Study AnalysisabstractInspired by cognitive functions of the human brain, machine learning-driven synthesis flows can map Boolean functions as Classification Trees that work like statistical inference engines. Circuits of this kind infer output values by evaluating the key features of the function learned during the training stage. We propose this idea for arithmetic circuits and, more specifically, for the design of aninferential8-by-8 bit unsigned multiplier. Using as case-study an error-resilient image blending application, we quantify the most representative figures of merit, also giving comparison against a classical radix-4 multi-level implementation. Experimental results demonstrate theinferentialmultiplier guarantees 76% average accuracy, 22% less area, and 2× latency reduction that can be used for power optimization. Roberto Giorgio Rizzo, Valerio Tenace, Andrea Calimera |
ISCAS | 3 |
| 2018 | Design-Space Exploration of Pareto-Optimal Architectures for Deep Learning with DVFSabstractSpecialized computing engines are required to accelerate the execution of Deep Learning (DL) algorithms in an energy-efficient way. To adapt the processing throughput of these accelerators to the workload requirements while saving power, Dynamic Voltage and Frequency Scaling (DVFS) seems the natural solution. However, DL workloads need to frequently access the off-chip memory, which tends to make the performance of these accelerators memory-bound rather than computation-bound, hence reducing the effectiveness of DVFS. In this work we use a performance-power analytical model fitted on a parametrized implementation of a DL accelerator in a 28-nm FDSOI technology to explore a large design space and to obtain the Pareto points that maximize the effectiveness of DVFS in the sub-space of throughput and energy efficiency. In our model we consider the impact on performance and power of the off-chip memory using real data of a commercial low-power DRAM. Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto |
ISCAS | 4 |
| 2018 | Energy-Driven Precision Scaling for Fixed-Point ConvNetsabstractData precision scaling is a well-known technique for power/energy minimization in error-resilient applications. It has proven particularly suited for embedded Convolutional Neural Networks (ConvNets) made run on fixed-point arithmetic coprocessors. The key observation is that methods that only account for accuracy during the precision assignment process may lead to sub-optimal energy minimization. This work introduces an energy-driven optimization that delivers per-layer quantization under a user-defined accuracy constraint. The tool is conceived for accelerators that dynamically adapt their energy and accuracy through software-programmable multiprecision Multiply&Accumulate (MAC) units. Simulation results collected on different ConvNets trained with public data-set show substantial energy savings and improved energy-accuracy tradeoffs w.r.t. conventional fixed-point methods. Valentino Peluso, Andrea Calimera |
VLSI-SoC | 2 |
| 2018 | Inferential Logic: a Machine Learning Inspired Paradigm for Combinational CircuitsabstractMachine learning (ML) theories and tools suggest alternative forms to conceive and represent relationships among data. The same theories find their application in the Boolean domain, where logic functions can be described as inference rules. This paper introduces Inferential Logic, a novel paradigm that leverages the ML concept of statistical inference for the design of combinational logic circuits, the Inferential Logic Circuits (ILCs). This new design concept is conceived for low-power circuits that run quasi-exact computation in error-resilient applications, but it also provides an exact run-mode that can be dynamically enabled when accuracy scaling is not an option. Valerio Tenace, Andrea Calimera |
VLSI-SoC | 2 |
| 2018 | Approximate Error Detection-Correction for efficient Adaptive Voltage Over-Scaling
Roberto Giorgio Rizzo, Andrea Calimera, Jun Zhou 0017 |
Integr. | 2 |
| 2018 | Quasi-exact logic functions through classification trees
Valerio Tenace, Andrea Calimera |
Integr. | 2 |
| 2017 | Early bird sampling: A short-paths free error detection-correction strategy for data-driven VOSabstractRazor is a milestone in the field of Error Detection&Correction strategies for low-power operation. Despite the impressive level of maturity, its application on circuits other than pipelined processors still remains an open issue. Firstly, the error detection mechanism relies on special flip-flops (FFs), the Razor-FFs, whose use imposes heavy hold-time fixing and large circuit area/power overheads; secondly, the error correction is performed through instruction replay, a practice that is not available (or very expensive to implement) in generic circuits. This work introduces Early Bird Sampling (EBS), a Razor variant that applies to low-power sequential circuits. The EBS allows to (i) solve the problem of short-path races bypassing tedious holdtime fixing design stages, (ii) reduce design overhead exploiting a local logic-masking mechanism for error correction. As a key feature, EBS enables Data-Driven Voltage Over-Scaling (DD-VOS), an aggressive dynamic voltage scaling strategy particularly suited for ultra-low power error-resilient applications. Simulation runs on a representative set of circuits provide a fair comparison with a standard Razor strategy. The collected results show EBS reduces area overheads (3.6% against 71.6% for Razor) and improves the voltage scaling profile achieving lower energy-per-operation (savings w.r.t. Razor range from 19.1% to 53.1%). Roberto Giorgio Rizzo, Valentino Peluso, Andrea Calimera, Jun Zhou 0017, Xin Liu 0015 |
VLSI-SoC | 3 |
| 2016 | Graphene-PLA (GPLA): a Compact and Ultra-Low Power Logic Array ArchitectureabstractThe key characteristics of the next generation of ICs for wearable applications include high integration density, small area, low power consumption, high energy-efficiency, reliability and enhanced mechanical properties like stretchability and transparency. The proper mix of new materials and novel integration strategies is the enabling factor to achieve those design specifications. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | Enabling quasi-adiabatic logic arrays for silicon and beyond-silicon technologiesabstractAdiabatic logic aims at mimicking an adiabatic (i.e., without energy exchange) charging process in digital circuits. Although regarded as a mostly theoretical computation style, research on the topic has been constantly active over the years, providing several demonstrations of working implementations [1]. The interest in adiabatic circuits recently increased with the introduction of emerging devices, e.g., Nanoelectromechanicals switches (NEMs) [2] and graphene p-n junctions [3], which have been proven to be good technological vehicles for adiabatic computing. Despite their energy efficiency, adiabatic logic faced severe limitations in reaching large scale integration due to the difficulty in logic pipelining and the lack of CAD tools able to cope with today's design complexity. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
ISCAS | 2 |
| 2016 | Ultra-Fine Grain Vdd-Hopping for energy-efficient Multi-Processor SoCsabstractThis paper introduces Ultra-Fine Grain Vdd-Hopping (FINE-VH), an extension of Dynamic Voltage-Frequency Scaling (DVFS) for energy efficient Multi-Processor SoCs (MPSoCs). The proposed technique leverages the working principle of Vdd-Hopping applied at ultra-fine granularity, i.e., within the core, by means of a layout-assisted, level-shifter free, dynamic dual-Vdd control strategy where leakage currents are minimized through an optimal timing-driven poly-bias assignment procedure. Valentino Peluso, Andrea Calimera, Enrico Macii, Massimo Alioto |
VLSI-SoC | 2 |
| 2016 | Multi-function logic synthesis of silicon and beyond-silicon ultra-low power pass-gates circuitsabstractPass-gates logic is known to be intrinsically more energy efficient than static CMOS. This feature attracted the research interest over the years and many working implementations have been demonstrated. Recent works, in particular, have shown that pass-gates logic is well suited for ultra-low power adiabatic circuits mapped on emerging technologies. Despite the progress made, several design issues still prevent pass-gates logic circuits reaching large scale integration. In this work we deal with the lack of synthesis tools and methodologies. We propose a multi-function decomposition engine that yields (i) an efficient abstract circuit modeling through a more compact data-structure, the Multi-Function Pass Diagram (MFPD) and (ii) an effective multi-gate area/delay-driven low-power synthesis&optimization flow. Simulation results conducted on different technologies, i.e., silicon and graphene, demonstrate that logic circuits synthesized with the proposed tool are smaller in size and depth, hence less power consuming and faster than circuits obtained through conventional synthesis flows based on Binary Decision Diagrams. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
VLSI-SoC | 2 |
| 2015 | One-pass logic synthesis for graphene-based Pass-XNOR logic circuitsabstractElectrostatically controlled graphene P-N junctions are devices built on a single layer graphene sheet that can be turned-ON/OFF via external potential difference. Their electrical behavior resembles a CMOS transmission gate with an embedded XNOR Boolean functionality. Recent works presented an efficient design style, the Pass-XNOR logic (PXL), which allows the implementation of adiabatic logic circuits with ultra low-power features. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
DAC | 2 |
| 2015 | Characterizing the Activity Factor in NBTI Aging Models for Embedded CoresabstractIn deeply scaled CMOS technologies, device aging causes cores performance parameters to degrade over time. While accurate models to efficiently assess these degradation exist for devices and circuits, no reliable model for processor cores has gained strong acceptance in the literature. In this work, we propose a methodology for deriving an NBTI aging model for embedded cores. Based on an accurate characterization on the netlist of the core, we were able to (1) prove the independence of the aging on the workload (i.e., executed instructions), and (2) calculate an equivalent average constant aging factor that justifies the use of the baseline model template. We derived and assessed the proposed model by using a RISC-like processor core implemented in a 45nm process technology as a reference architecture, achieving a maximum error of 2.2% against simulated data on the core netlist. Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | Exploiting the Expressive Power of Graphene Reconfigurable Gates via Post-Synthesis OptimizationabstractAs an answer to the new electronics market demands, semiconductor industry is looking for different materials, new process technologies and alternative design solutions that can support Silicon replacement in the VLSI domain. The recent introduction of graphene, together with the option of electrostatically controlling its doping profile, has shown a possible way to implement fast and power efficient Reconfigurable Gates (RGs). Also, and this is the most important feature considered in this work, those graphene RGs show higher expressive power, i.e., they implement more complex functions, like Majority, MUX, XOR, with less area w.r.t. CMOS counterparts. Unfortunately, state-of-the-art synthesis tools, which have been customized for standard NAND/NOR CMOS gates, do not exploit the aforementioned feature of graphene RGs. Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino, Luca G. Amarù, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Design and Characterization of Analog-to-Digital Converters using Graphene P-N JunctionsabstractElectrostatically controlled graphene p-n junctions are devices built on single-layer graphene sheets whose in-to-out resistance can be dynamically tuned through external voltage potentials. Roberto Giorgio Rizzo, Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | An automated design flow for approximate circuits based on reduced precision redundancyabstractReduced Precision Redundancy (RPR) is a popular Approximate Computing technique, in which a circuit operated in Voltage Over-Scaling (VOS) is paired to a reduced-bitwidth and faster replica so that VOS-induced timing errors are partially recovered by the replica, and their impact is mitigated. Previous works have provided various examples of effective implementations of RPR, which however suffer from three limitations: first, these circuits are designed using ad-hoc procedures, and no generalization is provided; second, error impact analysis is carried out statistically, thus neglecting issues like non-elementary data distribution and temporal correlation. Last, only dynamic power was considered in the optimization. In this work we propose a new generalized approach to RPR that allows to overcome all these limitations, leveraging the capabilities of state-of-the-art synthesis and simulation tools. By sacrificing theoretical provability in favor of an empirical input-based analysis, we build a design tool able to automatically add RPR to a preexisting gate-level netlist. Thanks to this method, we are able to confute some of the conclusions drawn in previous works, in particular those related to statistical assumptions on inputs; we show that a given inputs distribution may yield extremely different results depending on their temporal behavior. Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino |
ICCD | 2 |
| 2014 | Pass-XNOR logic: A new logic style for P-N junction based graphene circuitsabstractIn this work we introduce a new logic style for p-n junctions based digital graphene circuits: the pass-XNOR logic style. The latter enables the realization of compact, energy efficient circuits that better exploit the characteristics of graphene. We first show how a single p-n junction can be conceived as a pass-XNOR gate, i.e., a transmission gate with embedded logic functionality, the XNOR Boolean operator. Secondly, we propose a smart integration strategy in which series/parallel connections of pass-XNOR gates allow to implement AND/OR logical conjunctions, and, therefore, all possible truth tables. Experimental results conducted on a set of representative logic functions show the superior of pass-XNOR logic circuits w.r.t. standard CMOS circuits and graphene circuits that use p-n junctions in a complementary-like structure. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 2 |
| 2014 | Ultra Low-Power Computation via Graphene-Based Adiabatic Logic GatesabstractAs an answer to the difficulties in improving the figures of merit of deeply-scaled CMOS devices, researchers have looked at alternative materials and technologies for implementing new devices that can overcome the limitations of CMOS. Graphene has emerged as one of the most promising candidates among these new materials, recent works have demonstrated the implementation of electrostatically-controlled p-n junctions that can serve as the basic primitive for a new class of compact, fast and energy-efficient graphene-based logic gates. In this work we revisit those gates from a different perspective, namely, as devices that can operate adiabatically, that is, that are able to reuse the dissipated energy. We show how to build the basic logic gates by appropriately interconnecting graphene based p-n junctions and characterize those adiabatic gates for power and performance. The comparison between these adiabatic gates and both adiabatic CMOS and their non-adiabatic graphene-based counterpart shows that the former can operate with 2X to 3X less average power, and about 4X better power-delay product. Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
DSD | 2 |
| 2014 | Modeling of Physical Defects in PN Junction Based Graphene Devices
Sandeep Miryala, Matheus Oleiro, Letícia Maria Veiras Bolzani, Andrea Calimera, Enrico Macii, Massimo Poncino |
J. Electron. Test. | 4 |
| 2014 | Dynamic Indexing: Leakage-Aging Co-Optimization for CachesabstractTraditional implementations of low-power states based on voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused by negative bias temperature instability (NBTI), which can be explained in terms of the intuitive correlation between the idleness and the reduced workload of a system. Such a joint benefit has been exploited only partially because of the different nature of energy and aging as cost functions: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually result in limited aging reductions. In this paper, we address this problem in the context of power-managed caches, which represent a critical target for NBTI-reduced aging: given their symmetric structure, SRAM structures are, in particular, sensitive to NBTI effects because they cannot take advantage of the value-dependent recovery typical of NBTI. We propose a strategy called dynamic indexing, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the various power managed units (e.g., lines). This distribution allows fully using the leakage optimization potential and extending the lifetime of a cache. We explore various alternatives, in particular different granularities of the power managed units as well as different reindexing functions. Experimental analysis shows that it is possible to simultaneously reduce leakage power and aging in caches, with minimal power consumption overhead. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Energy-optimal SRAM supply voltage scheduling under lifetime and error constraintsabstractThis work addresses the energy efficiency of the memory architecture in safety-critical systems that have to guarantee a given level of service and a minimum lifetime. We specifically target SRAM structures in which decreased reliability manifests itself in terms of the aging induced by NBTI (Negative Bias Temperature Instability), and in which the level of service is represented by the bit-error rate (BER). Andrea Calimera, Enrico Macii, Massimo Poncino |
DAC | 1 |
| 2013 | A verilog-a model for reconfigurable logic gates based on graphene pn-junctionsabstractSingle layer sheets of graphene show special electrical properties that can enable the next generation of smart ICs. Recent works have proven the availability of an electrostatically controlled pn-junction upon which it is possible to design multi-function reconfigurable logic devices that naturally behave as multiplexers. In this work we introduce a stable large-signal Verilog-A model that mimics the behavior of the aforementioned devices. The proposed model, validated through the SPICE characterization of a MUX-based standard cell library we designed as benchmark, represents a first step towards the integration of Electronic Design Automation tools that can support the design of all-graphene ICs. Sandeep Miryala, Mehrdad Montazeri, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2013 | Delay model for reconfigurable logic gates based on graphene PN-junctionsabstractIn this paper we address the problem of modeling the timing behavior of a new class of reconfigurable logic gates based on electrostatically controlled graphene pn-junctions. These gates naturally behave as a 2-to-1 multiplexer in which the polarity of the input select line can be dynamically reconfigured. Interconnection of multiple gates and proper assignments of the inputs signals allow to implement all the basic Boolean logic functions, and, at a larger scale, any digital circuit. Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Layout-Driven Post-Placement Techniques for Temperature Reduction and Thermal Gradient MinimizationabstractWith the continuing scaling of CMOS technology, on-chip temperature and thermal-induced variations have become a major design concern. To effectively limit the high temperature in a chip equipped with a cost-effective cooling system, thermal specific approaches, besides low power techniques, are necessary at the chip design level. The high temperature in hotspots and large thermal gradients are caused by the high local power density and the nonuniform power dissipation across the chip. With the objective of reducing power density in hotspots, we propose two placement techniques that spread cells in hotspots over a larger area. Increasing the area occupied by the hotspot directly reduces its power density, leading to a reduction in peak temperature and thermal gradient. To minimize the introduced overhead in delay and dynamic power, we maintain the relative positions of the coupling cells in the new layout. We compare the proposed methods in terms of temperature reduction, timing, and area overhead to the baseline method, which enlarges the circuit area uniformly. The experimental results showed that our methods achieve a larger reduction in both peak temperature and thermal gradient than the baseline method. The baseline method, although reducing peak temperature in most cases, has little impact on thermal gradient. Wei Liu 0016, Andrea Calimera, Alberto Macii, Enrico Macii, Alberto Nannarelli, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | IR-drop analysis of graphene-based power distribution networksabstractElectromigration (EM) has been indicated as the killer effect for copper interconnects. ITRS projections show that for future technologies (22nm and beyond) the on-chip current demand will exceed the physical limit copper metal wires can tolerate. This represents a serious limitation for the design of power distribution networks of next generation ICs. New carbon nanomaterials, governed by ballistic transport, have shown higher immunity to EM, thereby representing potential candidate to replace copper. In this paper we make use of compact conductance models to benchmark Graphene Nanoribbons (GNRs) against copper. The two materials have been used to route a state-of-the-art multi-level power-grid architecture obtained through an industrial 45nm physical design flow. Although the adopted design style is optimized for metal grids, results obtained using our simulation framework show that GNRs, if properly sized, can outperform copper, thus allowing the design of reliable circuits with reduced IR-drop penalties. Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 2 |
| 2012 | Investigating the effects of Inverted Temperature Dependence (ITD) on clock distribution networksabstractThe aggressive scaling of CMOS technology toward nanometer lengths contributed to the surfacing of many effects that were not appreciable at the micrometer regime. Among them, Inverted Temperature Dependence (ITD) is certainly the most unusual. It manifests itself as a speed up of CMOS gates when the temperature increases, resulting in a reversal of the worst-case condition, i.e., CMOS gates show the largest delay at low temperatures. On the other hand, for metal interconnects an high temperature still holds as worst case condition. The two contrasting behaviors may invalidate the results obtained through standard design flow which do not consider temperature as an explicit variable in their optimizations. In this paper we focus on the impact of ITD on clock distribution networks (CDN), whose function is vital to guarantee the synchronization among physically spaced sequential components of digital circuits. Using our simulation framework, we characterized the thermal behavior of a clock tree mapped onto an industrial 65nm CMOS technology and obtained using a standard synthesis tool. Results demonstrate the presence of ITD at low operating voltages and open new potential research scenarios into the EDA field. Alessandro Sassone, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino, Richard Goldman, Vazgen Melikyan, Eduard Babayan, Salvatore Rinaudo |
DATE | 2 |
| 2012 | NBTI effects on tree-like clock distribution networksabstractNegative Bias Temperature Instability (NBTI) is considered one of the most critical device reliability concerns in nanometer CMOS technologies, because it causes devices to exhibit a temporal drift of performance over time. Wei Liu 0016, Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2012 | Energy-optimal caches with guaranteed lifetimeabstractThis work addresses the aging of the memory sub-system due to NBTI (Negative Bias Temperature Instability) in systems that have to provide a guaranteed level of service, and specifically, a guaranteed lifetime. Mirko Loghi, Haroon Mahmood, Andrea Calimera, Massimo Poncino, Enrico Macii |
ISLPED | 3 |
| 2012 | NBTI-Aware Data Allocation Strategies for Scratchpad Based Embedded Systems
Cesare Ferri, Dimitra Papagiannopoulou, R. Iris Bahar, Andrea Calimera |
J. Electron. Test. | 4 |
| 2011 | Partitioned cache architectures for reduced NBTI-induced agingabstractConventional power management knobs such as voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused Negative Bias Temperature Instability (NBTI). Such a benefit can be especially exploited in SRAM memories, which are particularly sensitive to NBTI effects: given their symmetric structure, they cannot in fact take advantage of value-dependent recovery. We propose an architectural solutions that is based on the idea of partitioning a memory into multiple banks of identical size. While this organization has been widely used for reducing both dynamic and static power, its exploitation for aging benefits requires proper management of the existing idleness of the various banks. This can be achieved by means of a sort of time-varying addressing scheme in which addresses are mapped to different banks over time in such a way that the idleness is uniformly distributed over all the banks. Experimental analysis shows that it is possible to simultaneously reducing leakage power and aging in caches, with minimal overhead and without modifying the internal structure of the SRAM arrays. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
DATE | 1 |
| 2011 | Moving to Green ICT: From stand-alone power-aware IC design to an integrated approach to energy efficient design for heterogeneous electronic systemsabstractEnergy efficiency is one of the most critical aspects of todays information society. The most obvious benefits of being Green are reduced environmental impact and cost savings. Reducing energy consumption of electronic devices, circuits and heterogeneous systems, however, is not trivial. This requires the development of innovative energy-aware vertical design solutions and EDA technologies for next generations' nanoelectronics circuits and systems, and the related energy generation, conversion and management systems. Salvatore Rinaudo, Giuliana Gangemi, Andrea Calimera, Alberto Macii, Massimo Poncino |
DATE | 3 |
| 2011 | Buffering of frequent accesses for reduced cache agingabstractPrevious works have shown that typical power management knobs such as voltage scaling or power gating can also be exploited to reduce aging phenomena caused by Negative Bias Temperature Instability (NBTI). We propose a scheme for power-managed caches that allows to significantly improving the aging of the cache thanks to the use of a small buffer that stores a copy of the lines that are most critical for aging, that is, the ones with the least opportunity of being power-managed; by using the buffer instead of the cache when accessing these critical lines, the original cache is preserved and its lifetime is significantly prolonged. As a side effect, this scheme improves total power since the less energy-hungry buffer is accessed most of the time. Experimental analysis shows this scheme allows to achieve significant (>3x on average) lifetime extensions for the cache, with a concurrent energy saving between 18 and 24%, depending on cache size. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 1 |
| 2010 | Post-placement temperature reduction techniquesabstractWith technology scaled to deep submicron era, temperature and temperature gradient have emerged as important design criteria. We propose two post-placement techniques to reduce peak temperature by intelligently allocating whitespace in the hotspots. Both methods are fully compliant with commercial technologies, and can be easily integrated with state-of-the-art thermal-aware design flow. Experiments in a set of tests on circuits implemented in STM 65nm technologies show that our methods achieve better peak temperature reduction than directly increasing circuit's area. Wei Liu 0016, Alberto Nannarelli, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2010 | An integrated thermal estimation framework for industrial embedded platformsabstractNext generation industrial embedded platforms require the development of complex power and thermal management solutions. Indeed, an increasingly fine and intrusive thermal control is required because of temperature impact on leakage and reliability. To be effective, the implementation of these policies involves decisions that must be taken during various phases along the design process, to enable the development of architectural level countermeasures and the required hardware knobs, such as power modes, power supply regulation granularity and the number of on-chip temperature sensors. As a consequence, a framework allowing thermal estimation exploiting design-time information is desirable. Andrea Acquaviva, Andrea Calimera, Alberto Macii, Massimo Poncino, Enrico Macii, Matteo Giaconia, Claudio Parrella |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Aging effects of leakage optimizations for cachesabstractBesides static power consumption, sub-90nm devices have to account for NBTI effects, which are one of the major concerns about system reliability. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 1 |
| 2010 | Analysis of NBTI-induced SNM degradation in power-gated SRAM cellsabstractTemporal reliability degradation mechanism, and NBTI in particular, are especially critical for SRAM cells. In fact, unlike logic gates, which under some conditions can be forced into an NBTI-immune state, SRAM cells are always subject to aging, whatever value they are storing. In this work, we first quantify the aging, in terms of degradation of the signal-to-noise margin (SNM), of an SRAM cell as a function of the value stored in the cell, on a 45nm industrial technology. Then, we show how it is possible, by applying power gating to the memory cell, to further reduce the SNM degradation. Finally, we study the joint effect of power gating and bit control techniques. Andrea Calimera, Enrico Macii, Massimo Poncino |
ISCAS | 1 |
| 2010 | Dynamic indexing: concurrent leakage and aging optimization for cachesabstractPrevious works have shown that the traditional implementations of power management (i.e., using power gating or voltage scaling) can also mitigate the aging effect induced by Negative Bias Temperature Instability (NBTI), due to the partial recovery that occurs during the idle intervals used by power management. However, such a potential has been exploited only partially because of the different nature of energy and aging: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually turn into limited aging reductions. We address this problem in the context of caches, for which idleness is related to their access pattern. We propose a dynamic indexing scheme, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the cache lines. In this way it is possible to fully use the leakage optimization potential and to extend the lifetime of a cache. Experimental analysis shows that it is possible to obtain caches that are effectively aging-free, without any penalty in leakage energy reduction. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
ISLPED | 1 |
| 2010 | NBTI-Aware Clustered Power GatingabstractThe emergence of Negative Bias Temperature Instability (NBTI) as the most relevant source of reliability in sub-90nm technologies has led to a new facet of the traditional trade-off between power and reliability. NBTI effects in fact manifest themselves as an increase of the propagation delay of the devices over time, which adds up to the delay penalty incurred by most low-power design solutions. This implies that, given a desired lifetime of a circuit (i.e., a given performance target at some point in time), a power-managed component will fail earlier than a nonpower-managed one. In this work, we show how it is possible to partially overcome this conflict, by leveraging the benefits in terms of aging provided by power-gating (i.e., by using switches that disconnect a logic block from the ground). Thanks to some electrical properties, it is possible to nullify aging effects during standby periods. Based on this important property, we propose a methodology for a NBTI-aware power gating that allows synthesizing low-leakage circuits with maximum lifetime. Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2010 | Temperature-Insensitive Dual- Vth Synthesis for Nanometer CMOS Technologies Under Inverse Temperature DependenceabstractWith the scaling of CMOS technologies, the gap between nominal supply voltage and threshold voltage has decreased significantly. This trend is further amplified in low-power nanometer libraries, which feature cells with identical size and functionality, but different threshold voltages. As a consequence, different cells may have different delay behaviors as the temperature varies within a circuit. For instance, cells with low-threshold devices may experience an increase in delay when temperature increases, whereas cells using high-threshold devices may experience the opposite behavior. The latter effect, also known as inverse temperature dependence (ITD), poses new challenges to circuit designers. Besides making timing analysis more difficult, ITD has important and unforeseeable consequences for power-aware logic synthesis. This paper describes the impact that ITD may have on the design of nanometer circuits. We also provide a threshold voltage assignment algorithm for dual threshold voltage synthesis, which guarantees temperature-insensitive operation of the circuits, together with a significant reduction of both leakage and total power consumption. Experiments performed on a set of standard benchmarks show timing compliance at any operating temperature, and an average leakage reduction around 28% compared to circuits synthesized with a standard synthesis flow that does not take ITD into account. We also apply our proposed synthesis algorithm to a realistic case study consisting of a 32-bit, IEEE-754 floating point unit. Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Enabling concurrent clock and power gating in an industrial design flowabstractClock-gating and power-gating have proven to be very effective solutions for reducing dynamic and static power, respectively. The two techniques may be coupled in such a way that the clock-gating information can be used to drive the control signal of the power-gating circuitry, thus providing additional leakage minimization conditions w.r.t. those manually inserted by the designer. This conceptual integration, however, poses several challenges when moved to industrial design flows. Although both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. This paper presents a layout-oriented synthesis flow which integrates the two techniques and that relies on leading-edge, commercial EDA tools. Starting from a gated-clock netlist, we partition the circuit in a number of clusters that are implicitly determined by the groups of cells that are clock-gated by the same register. Using a row-based granularity, we achieve runtime leakage reduction by inserting dedicated sleep transistors for each cluster. The entire flow has been benchmarked on a industrial design mapped onto a commercial, 65 nm CMOS technology library. Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 2 |
| 2009 | NBTI-aware sleep transistor design for reliable power-gatingabstractNegative Bias Temperature Instability (NBTI) has been regarded as most important source of reliability of CMOS devices, and specifically pMOS transistors. Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 1 |
| 2009 | Placement-aware Clustering for Integrated Clock and Power GatingabstractClock-gating and power-gating are the most widely used solutions for reducing dynamic and static power. They can be potentially integrated so that clock-gating conditions can be used to control the power-gating circuitry thus also reducing static power. This integration becomes however difficult when applied in an industrial design flow. Even if both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 2 |
| 2009 | NBTI-aware power gating for concurrent leakage and aging optimizationabstractPower and reliability are known to be intrinsically conflicting metrics: traditional solutions to improve reliability such as redundancy, increase of voltage levels, and up-sizing of critical devices do contrast with traditional low-power solutions, which rely on small devices and scaled supply voltages. The emergence of Negative Bias Temperature Instability (NBTI) as the most relevant source of unreliability in sub-90nm technologies has even exacerbated this incompatibility of the two metrics: NBTI manifests itself as an increase of the propagation delay over time, which adds up to the delay penalty introduced by most low-power design solutions. In this work, we show how the most widely adopted leakage reduction solution, that is, power-gating, can overcome this conflict, and how it can be used to naturally reduce the effects of NBTI on delay. Based on this important property, we present a methodology for NBTI-aware power gating that allows synthesizing low-leakage circuits with maximum lifetime. Andrea Calimera, Enrico Macii, Massimo Poncino |
ISLPED | 1 |
| 2008 | Optimal MTCMOS Reactivation Under Power Supply Noise and Performance ConstraintsabstractSleep transistor insertion is one of today's most promising and widely adopted solutions for controlling stand-by leakage power in nanometer circuits. Although single-cycle power mode transition reduces wake-up latency, it originates large discharge current spikes, thereby causing IR-drop and inductive ground bounce for the surrounding circuit blocks. We propose a new reactivation solution which helps in controlling power supply fluctuations and in achieving minimum reactivation times. Our structure limits the turn-on current below a given threshold through sequential activation of the sleep transistors, which are connected in parallel and are sized using a novel optimal sizing algorithm. The proposed methodology is validated using HSPICE simulations of several benchmark circuits, which have been synthesized onto a commercial 65 nm CMOS technology library. Andrea Calimera, Luca Benini, Enrico Macii |
DATE | 1 |
| 2008 | Integrating Clock Gating and Power Gating for Combined Dynamic and Leakage Power Optimization in Digital CMOS CircuitsabstractClock Gating and Power Gating are two of the most effective techniques that are applied today for reducing dynamic and leakage power, respectively, in digital CMOS circuits. The combined use of the two solutions, however, poses some challenges in terms of practical integration of the required control logic and the power/timing overhead associated to it. This paper presents an analysis methodology and a prototype CAD tool that support the designer in understanding when the joint application of Clock Gating and Power Gating may result in significant power savings. Enrico Macii, Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Massimo Poncino |
DSD | 3 |
| 2008 | Temperature-insensitive synthesis using multi-vt librariesabstractTemperature fluctuations can alter the delay in MOS circuits. However, increases in temperature do not always lead to a corresponding increase in circuit delay, specifically when operating at low supply voltages. Instead a temperature inversion effect can be observed on the delay of MOS devices under certain conditions, where the delay actually decreases as temperature increases. Given these non-monotonic effects, guaranteeing timing correctness can no longer be achieved simply by characterizing the design under worst case (i.e., high temperature) conditions. In this paper, we present a synthesis methodology in which multi-Vth design is used to generate temperature-insensitive circuits, while minimizing leakage power dissipation as a side-effect. Our experiments with ISCAS benchmark circuits demonstrate the promise of this approach and show that significant reduction in static power is also possible. Andrea Calimera, Enrico Macii, Massimo Poncino, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 1 |
| 2008 | On quantifying the figures of merit of power-gating for leakage power minimization in nanometer CMOS circuitsabstractPower-gating has proved to be one of the most effective solutions for reducing stand-by leakage power in nanometer-scale CMOS circuits, and different strategies and algorithms for its application have been proposed recently. Unfortunately, power- gating comes with its own set of costs: Performance degradation, area increase, dynamic power increase and routing congestion. When a decision to power-gate a design has to be taken, pros and cons of power-gating have to be properly weighted to achieve optimal results. In this paper, we define "Figures of Merit" (FoMs) for power-gating, which can be used by designers to better understand the benefits and costs of power-gating, thereby allowing them to achieve optimal results. We then quantify the FoMs by applying a state-of-the-art, industry-strength power- gating flow on a set of designs implemented onto an industrial 65 nm CMOS process, and provide insightful discussion on how optimum power-gating can be achieved. Ashoka Visweswara Sathanur, Andrea Calimera, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 2 |
| 2008 | Reducing leakage power by accounting for temperature inversion dependence in dual-Vt synthesized circuitsabstractThe effects of temperature on delay depend on several parameters, such as cell size, load, supply voltage, and threshold voltage. In particular, variations in Vth can yield a temperature inversion effect causing a decreases of cell delay as temperature increases. This phenomenon, besides affecting timing analysis of a design, has important and unforeseeable consequences on power optimization techniques. In this paper, we focus on the impact of such effects on multi-Vt design; in particular, we show how traditional dual-Vt optimization may yield timing errors in circuits by ignoring temperature effects. Moreover, we present a temperature-aware dual-Vt optimization technique that reduces leakage power and can guarantee that the circuit is timing feasible at the boundary temperatures provided by the technology library. Our experiments show an average 27% leakage reduction with respect to a non temperature-aware design flow. Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino |
ISLPED | 1 |
| 2007 | Interactive presentation: Efficient computation of discharge current upper bounds for clustered sleep transistor sizing
Ashoka Visweswara Sathanur, Andrea Calimera, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 2 |
| 2007 | Design of a family of sleep transistor cells for a clustered power-gating flow in 65nm technologyabstractClustered sleep transistor insertion is an effective leakage power reduction technique that is well-suited for integration in an automated design flow and offers a flexible tradeoff between area, delay overhead and turn-on transition time. In this work, we focus on the design of a family of sleep transistor cells, fully compatible with the physical design rules of a commercial 65nm CMOS library. We describe circuit-level and layout optimizations, as well as the cell characterization procedure required to support automated sleep transistor cell selection and instantiation in a clustered power-gating insertion flow. Andrea Calimera, Antonio Pullini, Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 1 |