Enrico Macii

dblp:97/6648 · DBLP profile ↗
← Back
321ranked-venue papers
18as first author
64since 2021 · last 2026
0000-0001-9046-5618ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 270 · 18 first-author · 36 since 2021Software engineering, systems software and programming languages · 67 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 61 · 19 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Computer networks · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Late Breaking Results: CHESSY: Coupled Hybrid Emulation with SystemC-FPGA Synchronization
abstract
The growing complexity of cyber-physical systems (CPSs) calls for early prototyping tools that combine accuracy, speed, and usability. Virtual Platforms (VPs) provide fast functional simulation, but hybrid co-emulation solutions, in which key digital components are deployed on FPGA, become necessary when accurate timing modelling is required and RTL simulation is too costly. However, existing hybrid emulation tools are mostly proprietary, and rely on vendor-specific FPGA features. To address this gap, we introduce an open-source framework that connects SystemC-based VPs with FPGA emulation, enabling full-system co-emulation of digital and non-digital components. The FPGA accelerates the execution of main digital subsystems, while a wrapper coordinates timing and communication with the VP through JTAG, maintaining synchronization with simulated peripherals. Evaluations using a RISC-V SoC, with an example in the biosignals processing domain, show up to 2500× speedup compared to RTL simulation, while maintaining less than 2× total simulation time relative to pure FPGA emulation.
Lorenzo Ruotolo, Giovanni Pollo, Mohamed Amine Hamdi, Matteo Risso, Yukai Chen, Enrico Macii, Massimo Poncino, Sara Vinco, Alessio Burrello, Daniele Jahier Pagliari
DATE6
2026 Variable-precision neuromorphic state space model for on-edge activity classification
abstract
Neuromorphic computing is rising as a promising paradigm for efficient AI, leveraging event-driven computation to achieve low-power and high-performance computing. Due to the real-time processing required by edge devices with minimal power consumption, optimizing neuromorphic models for on-edge applications can be crucial to address the issue of power efficiency and resource-constraint devices. This work explores the definition of a neuromorphic state space model and its deployment on non-dedicated hardware. Structured sparsity and quantization techniques are leveraged to enhance the model’s efficiency. By compressing synaptic operations and memory footprint, we demonstrate how neuromorphic models can be adapted for on-edge deployment, ensuring low-latency and memory efficient inference. This study highlights the potential of neuromorphic models as a scalable solution for real-world embedded systems with limited resources.
Benedetto Leto, Gianvito Urgese, Enrico Macii, Vittorio Fra
Future Gener. Comput. Syst.3
2026 End-to-end Automated Deep Neural Network Optimization for PPG-based Blood Pressure Estimation on Wearables
abstract
Photoplethysmography-based Blood Pressure (BP) estimation is a challenging task, particularly on resource-constrained wearable devices. However, fully on-board processing is desirable to ensure user data confidentiality. Recent Deep Neural Networks (DNNs) have achieved high BP estimation accuracy by reconstructing BP waveforms or directly regressing BP values, but their large memory, computation, and energy requirements hinder deployment on wearables. This work introduces a fully automated DNN design pipeline that combines hardware-aware Neural Architecture Search, pruning, and Mixed-Precision Search to generate accurate yet compact BP prediction models optimized for ultra-low-power multi-core Systems-on-Chip (SoCs). Starting from state-of-the-art baseline models on four public datasets, our optimized networks achieve up to 7.99% lower error with a 7.5 \(\times\) parameter reduction, or up to 83 \(\times\) fewer parameters with negligible accuracy loss. All models fit within 512 kB of memory on our target SoC (GreenWaves’ GAP8), requiring less than 55 kB and achieving an average inference latency of 142 ms and energy consumption of 7.25 mJ. Patient-specific fine-tuning further improves accuracy by up to 64%, enabling fully autonomous, low-cost BP monitoring on wearables.
Francesco Carlucci, Giovanni Pollo, Xiaying Wang, Massimo Poncino, Enrico Macii, Luca Benini, Sara Vinco, Alessio Burrello, Daniele Jahier Pagliari
ACM Trans. Comput. Heal.5
2026 Runtime Feature Compression for Adaptive Keyword Spotting on Embedded Systems
abstract
Voice user interfaces rely on keyword spotting (KWS) to detect wake-word commands, enabling low-power devices to switch from drowsy to active states and initiate more complex tasks. In embedded systems, KWS combines handcrafted acoustic features extraction with lightweight neural network classifiers to achieve accurate detection within strict resource constraints. Adapting KWS to time-varying energy budgets requires optimization strategies that operate at runtime. Most existing approaches adjust the complexity of the neural model but overlook that a substantial amount of latency, and thus energy consumption, is due to feature extraction, which remains unaffected by model scaling. This work introducesRuntime Feature Compression(RFC), a dynamic rescaling strategy that modulates the workload of the entire KWS pipeline. RFC promotes thehop-lengthparameter of the Short-Time Fourier Transform as a runtime control knob to adjust the number of time frames in speech features, allowing a single model to operate across multiple latency modes. To support this flexibility, we introduce two training-time techniques:HopAugment, a data augmentation scheme that exposes the model to variable hop lengths during training, andMasked Layers, which preserve consistent activation statistics during training and inference under compressed feature settings. Evaluations on four KWS datasets using the TC-ResNet model family show that RFC outperforms model scaling techniques, offering a wider range of latency-accuracy trade-offs. RFC achieves up to 31.8% lower latency without accuracy degradation, or up to 0.30% higher accuracy within equivalent latency bounds. That proves RFC improves adaptability in energy-constrained IoT speech interfaces. A set of ablation studies further demonstrates the robustness of RFC by evaluating the role of its training components, batching strategies, ability to preserve accuracy with a shared weight set, scalability across operating modes, and applicability to different model architectures.
Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi
IEEE Internet Things J.3
2026 The inNuCE Research Infrastructure and the Neuromorphic MLOps for AIoT Prototyping
abstract
Neuromorphic computing promises significant improvements in latency and energy efficiency for machine intelligence at the edge. However, its adoption in the IoT domain is still limited by the heterogeneity of the HW, the immaturity of the toolchains, and the poor reproducibility of experiments. The present paper sets out the inNuCE RI, a two-pillar facility composed of a physical inNuCE Lab and a cloud-based inNuCE HPP. The purpose of the inNuCE RI is to enable developers to prototype, evaluate and compare neuromorphic and conventional end-to-end digital solutions. From a methodological perspective, we formalize the adaptation of MLOps to event-driven sensing and brain-inspired computation as NMLOps. We illustrate how inNuCE RI instantiates NMLOps through containerized toolchains orchestrated with Kubernetes and Slurm-managed heterogeneous resources (neuromorphic chips, FPGAs, GPUs, MCUs). The approach is analyzed on representative AIoT use cases, including HAR, Braille reading, event-based gesture recognition, Hi-Co semantization of memories, navigation tracking, and constraint satisfaction problems. The development of inNuCE RI has been driven by the need to facilitate the transition from prototype (in nuce) to engineered AIoT systems for lower entry barriers and enforce reproducibility. This paves the way for future system-of-systems engineering.
Gianvito Urgese, Vittorio Fra, Andrea Pignata, Giuseppe Fanuli, Walter Gallego Gomez, Riccardo Pignari, Michelangelo Barocci, Benedetto Leto, Salvatore Tilocca, Nicola Cassetta, Paolo Montuschi, Enrico Macii
IEEE Internet Things J.12
2026 FedAGF: Adaptive Concurrency via Gradient Feedback for Mitigating Extreme Label Skew in Budget-Constrained Federated Learning
abstract
Cross-device Federated Learning (FL) enables large fleets of distributed edge devices to collaboratively train a global classification model without sharing their data. The resulting training quality is strongly limited byextreme label skew, a condition where each device holds samples from a subset of the target classes. In such cases, local updates become biased toward local distributions, slowing down convergence and degrading the global model accuracy. These effects get critical when devices operate under constrained energy budgets that restrict their participation to a limited number of synchronization rounds, further reducing the achievable accuracy. To overcome these limitations, we introduce Federated Learning with Adaptive Concurrency via Gradient Feedback (FedAGF), a control policy that dynamically adjusts the number of devices selected for synchronization throughout training. FedAGF adapts to training dynamics by monitoring global model updates and elevating participation whenever progress slows. This adaptive mechanism balances accuracy improvement with efficient use of devices’ energy budgets, allocating resources when they provide the greatest benefit to convergence. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that FedAGF effectively mitigates extreme label skew, achieving up to 14.14% higher accuracy than state-of-the-art FL methods and enabling efficient and scalable training, even under skewed label distributions and resource constraints.
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Privacy-Preserving Federated Learning for Household Characteristic Identification
abstract
This work presents a privacy-preserving training framework for household characteristic identification from electricity consumption data. The proposed framework integrates two main components: (i) a synthetic data generation pipeline capable of replicating realistic energy traces from diverse family compositions, capturing fine-grained sociodemographic attributes such as household size, employment status, age groups, and home occupancy patterns; (ii) a training strategy based on Federated Learning (FL) secured with homomorphic encryption, enabling collaborative model training while preserving data ownership. Our synthetic dataset enables the performance assessment of different training scenarios, including siloed model training by individual energy utilities and secure collaboration via FL. Experimental results show that siloed training leads to inconsistent and suboptimal performance, while privacy-preserving FL achieves accuracy comparable to conventional centralized training—an ideal yet not viable option due to data regulation constraints. Our findings highlight the effectiveness of FL as a secure solution for collaborative sociodemographic profiling in smart grids.
Erich Malan, Claudia De Vizia, Marco Castangia, Valentino Peluso, Andrea Calimera, Enrico Macii
COMPSAC6
2025 Energy-Aware Error Correction Method for Indoor Positioning and Tracking
abstract
Indoor positioning is crucial for the effective use of drones in smart environments, enabling precise navigation and control in complex indoor spaces where GPS signals are weak or unavailable and wireless communication-based systems must be used. In order to improve positioning accuracy, various distance measurement techniques and related error correction methods have been proposed in the literature. However, these methods are mostly focused on accuracy and often require a significant amount of computational resources, which is quite inefficient when deployed on battery-operated devices like small robots or drones because of their limited battery capacity. Moreover, conventional error correction methods are little effective for the tracking of moving objects. In this paper, we first analyze the trade-off between energy consumption and accuracy for the error correction and identify the most energy-efficient error correction method. Based on this analysis in the accuracy/energy space, we introduce a new energy-efficient error correction method that is especially targeted for tracking a moving object. We validated our solution by implementing an Ultra-Wideband based indoor positioning system and demonstrated that the proposed method improves positioning accuracy by 15% and reduces energy consumption by 33% compared to the state-of-the-art method.
Donkyu Baek, Yukai Chen, Enrico Macii, Massimo Poncino
DATE4
2025 Coupling Neural Networks and Physics Equations For Li-Ion Battery State-of-Charge Prediction
abstract
Estimating the evolution of the battery's State of Charge (SoC) in response to its usage is critical for implementing effective power management policies and for ultimately improving the system's lifetime. Most existing estimation methods are either physics-based digital twins of the battery or data-driven models such as Neural Networks (NNs). In this work, we propose two new contributions in this domain. First, we introduce a novel NN architecture formed by two cascaded branches: one to predict the current SoC based on sensor readings, and one to estimate the SoC at a future time as a function of the load behavior. Second, we integrate battery dynamics equations into the training of our NN, merging the physics-based and data-driven approaches, to improve the models' generalization over variable prediction horizons. We validate our approach on two publicly accessible datasets, showing that our Physics-Informed Neural Networks (PINNs) outperform purely data-driven ones while also obtaining superior prediction accuracy with a smaller architecture with respect to the state-of-the-art.
Giovanni Pollo, Alessio Burrello, Enrico Macii, Massimo Poncino, Sara Vinco, Daniele Jahier Pagliari
DATE3
2025 Gradient-Aware Participation for Energy Reduction in Federated Learning with Extreme Label Skew
abstract
Federated Learning (FL) enables distributed clients to train a global classification model collaboratively while preserving data privacy. A major challenge in FL is ensuring efficient training with limited computing and communication resources, especially when clients’ datasets contain samples from a restricted subset of target classes, a problem known as extreme label skew. Under such a condition, model updates from clients are biased toward their local data distributions, resulting in slow convergence and increased energy consumption due to the need for additional training rounds. This paper introduces FL with Gradient-Aware Participation (FedGAP), a novel strategy aimed at reducing energy consumption while preserving model accuracy even with extreme label skew. FedGAP dynamically adjusts the cohort size, i.e., the number of participating clients per training round, based on the evolution of the global model’s pseudo-gradient. By detecting stagnant phases where progress toward convergence stalls, FedGAP increases the cohort size to escape suboptimal regions and accelerate learning, thereby minimizing the waste of resources. Experiments on CIFAR-10 and CIFAR-100 demonstrate that FedGAP achieves up to 2.74× greater energy efficiency compared to state-of-the-art methods without compromising accuracy.
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii
IJCNN4
2025 Building Damage Assessment in Conflict Zones: A Deep Learning Approach Using Geospatial Sub-Meter Resolution Data
abstract
Very High Resolution (VHR) geospatial image analysis is crucial for humanitarian assistance in both natural and anthropogenic crises, as it allows to rapidly identify the most critical areas that need support. Nonetheless, manually inspecting large areas is time-consuming and requires domain expertise. Thanks to their accuracy, generalization capabilities, and highly parallelizable workload, Deep Neural Networks (DNNs) provide an excellent way to automate this task. Nevertheless, there is a scarcity of VHR data pertaining to conflict situations, and consequently, of studies on the effectiveness of DNNs in those scenarios. Motivated by this, our work extensively studies the applicability of a collection of state-of-the-art Convolutional Neural Networks (CNNs) originally developed for natural disasters damage assessment in a war scenario. To this end, we build an annotated dataset with pre- and post-conflict images of the Ukrainian city of Mariupol. We then explore the transferability of the CNN models in both zero-shot and learning scenarios, demonstrating their potential and limitations. To the best of our knowledge, this is the first study to use sub-meter resolution imagery to assess building damage in combat zones.
Matteo Risso, Alessia Goffi, Beatrice Alessandra Motetti, Alessio Burrello, Jean Baptiste Bove, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari, Giuseppe Maffeis
IPAS6
2025 MEbots: Integrating a RISC-V Virtual Platform with a Robotic Simulator for Energy-aware Design
abstract
Virtual Platforms (VPs) enable early software validation of autonomous systems’ electronics, reducing costs and time-to-market. While many VPs support both functional and non-functional simulation (e.g., timing, power), they lack the capability of simulating the environment in which the system operates. In contrast, robotics simulators lack accurate timing and power features. This twofold shortcoming limits the effectiveness of the design flow, as the designer can not fully evaluate the features of the solution under development. This paper presents a novel, fully open-source framework bridging this gap by integrating a robotics simulator (Webots) with a VP for RISC-V-based systems (MESSY). The framework enables a holistic, mission-level, energy-aware co-simulation of electronics in their surrounding environment, streamlining the exploration of design configurations and advanced power management policies.
Giovanni Pollo, Mohamed Amine Hamdi, Matteo Risso, Lorenzo Ruotolo, Pietro Furbatto, Matteo Isoldi, Yukai Chen, Alessio Burrello, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari, Sara Vinco
ISLPED9
2025 Foundation Models for Structural Health Monitoring
abstract
Structural Health Monitoring (SHM) is a critical task for ensuring the safety and reliability of civil infrastructures, typically realized on bridges and viaducts by means of vibration monitoring. In this paper, we propose for the first time the use of Transformer neural networks, with a Masked Auto-Encoder architecture, asFoundation Modelsfor SHM. We demonstrate the ability of these models to learn generalizable representations from multiple large datasets through self-supervised pre-training, which, coupled with task-specific fine-tuning, allows them to outperform state-of-the-art traditional methods on diverse tasks, including Anomaly Detection (AD) and Traffic Load Estimation (TLE). We then extensively explore model size versus accuracy trade-offs and experiment with Knowledge Distillation (KD) to improve the performance of smaller Transformers, enabling their embedding directly into the SHM edge nodes. We showcase the effectiveness of our foundation models using data from three operational viaducts. For AD, we achieve a near-perfect 99.9% accuracy with a monitoring time span of just 15 windows. In contrast, a state-of-the-art method based on Principal Component Analysis (PCA) obtains its first good result (95.03% accuracy), only considering 120 windows. On two different TLE tasks, our models obtain state-of-the-art performance on multiple evaluation metrics (R2score, MAE% and MSE%). On the first benchmark, we achieve an R2score of 0.97 and 0.90 for light and heavy vehicle traffic, respectively, while the best previous approach (a Random Forest) stops at 0.91 and 0.84. On the second one, we achieve an R2score of 0.54 versus the 0.51 of the best competitor method, a Long-Short Term Memory network.
Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta, Yhorman Alexander Bedoya Velez, Andrea Acquaviva, Massimo Poncino, Enrico Macii, Luca Benini, Alessio Burrello
IEEE Trans. Sustain. Comput.7
2024 VARADE: a Variational-based AutoRegressive model for Anomaly Detection on the Edge
abstract
Detecting complex anomalies on massive amounts of data is a crucial task in Industry 4.0, best addressed by deep learning. However, available solutions are computationally demanding, requiring cloud architectures prone to latency and bandwidth issues. This work presents VARADE, a novel solution implementing a light autoregressive framework based on variational inference, which is best suited for real-time execution on the edge. The proposed approach was validated on a robotic arm, part of a pilot production line, and compared with several state-of-the-art algorithms, obtaining the best trade-off between anomaly detection accuracy, power consumption and inference frequency on two different edge platforms.
Alessio Mascolini, Sebastiano Gaiardelli, Francesco Ponzio, Nicola Dall'Ora, Enrico Macii, Sara Vinco, Santa Di Cataldo, Franco Fummi
DAC5
2024 An AI-Enabled Framework for Smart Semiconductor Manufacturing
abstract
With the rise of Machine Learning (ML) and Artificial Intelligence (AI), the semiconductor industry is undergoing a revolution in how it approaches manufacturing. The SMART-IC project (DATE'24 MPP category: initial stage) works in this direction, by proposing an AI-enabled framework to support the smart monitoring and optimization of the semiconductor manufacturing process. An AI-powered engine examines sensor data recording physical parameters during production (like gas flow, temperature, voltage, etc.) as well as test data, with different goals: (1) the identification of anomalies in the production chain, either offline from collected data-traces or online from a continuous stream of sensed data; (2) the forecasting of new data of the future production; and (3) the automatic generation of synthetic traces, to strengthen the data-based algorithms. All such tasks provide valuable information to an advanced Manufacturing Execution System (MES), which reacts by optimizing the production process and management of the equipment maintenance policies. SMART-IC is a 300k€ academic project funded by the Italian Ministry of University and supported by STMicroelectronics and Technoprobe with industrial expertise and real-world applications. This paper shares the view of SMART-IC on the future of semiconductor manufacturing, the preliminary efforts, and the future results that will be reached by the end of the project, in 2025.
Khaled Alamin, Davide Appello, Alessandro Beghi, Nicola Dall'Ora, Fabio Depaoli, Santa Di Cataldo, Franco Fummi, Sebastiano Gaiardelli, Michele Lora, Enrico Macii, Alessio Mascolini, Daniele Pagano, Francesco Ponzio, Gian Antonio Susto, Sara Vinco
DATE10
2024 Model-Driven Feature Engineering for Data-Driven Battery SOH Model
abstract
Accurate State of Health (SoH) estimation is indispensable for ensuring battery system safety, reliability, and run-time monitoring. However, as instantaneous runtime measurement of SoH remains impractical when not unfeasible, appropriate models are required for its estimation. Recently, various data-driven models have been proposed, which solve various weaknesses of traditional models. However, the accuracy of data-driven models heavily depends on the quality of the training datasets, which usually contain data that are easy to measure but that are only partially or weakly related to the physical/chemical mechanisms that determine battery aging. In this study, we propose a novel feature engineering approach, which involves augmenting the original dataset with purpose-designed features that better represent the aging phenomena. Our contribution does not consist of a new machine-learning model but rather in the addition of selected features to an existing model. This methodology consistently demonstrates enhanced accuracy across various machine-learning models and battery chemistries, yielding an approximate 25% SoH estimation accuracy improve-ment. Our work bridges a critical gap in battery research, offering a promising strategy to significantly enhance SoH estimation by optimizing feature selection.
Khaled Alamin, Daniele Jahier Pagliari, Yukai Chen, Enrico Macii, Sara Vinco, Massimo Poncino
DATE4
2024 HW-SW Optimization of DNNs for Privacy-Preserving People Counting on Low-Resolution Infrared Arrays
abstract
Low-resolution infrared (IR) array sensors enable people counting applications such as monitoring the occupancy of spaces and people flows while preserving privacy and minimizing energy consumption. Deep Neural Networks (DNNs) have been shown to be well-suited to process these sensor data in an accurate and efficient manner. Nevertheless, the space of DNNs' archi-tectures is huge and its manual exploration is burdensome and often leads to sub-optimal solutions. To overcome this problem, in this work, we propose a highly automated full-stack optimization flow for DNNs that goes from neural architecture search, mixed-precision quantization, and post-processing, down to the realization of a new smart sensor prototype, including a Microcontroller with a customized instruction set. Integrating these cross-layer optimizations, we obtain a large set of Pareto-optimal solutions in the 3D-space of energy, memory, and accuracy. Deploying such solutions on our hardware platform, we improve the state-of-the-art achieving up to 4.2 x model size reduction, 23.8 x code size reduction, and 15.38 x energy reduction at iso-accuracy.
Matteo Risso, Francesco Daghero, Alessio Burrello, Seyedmorteza Mollaei, Marco Castellano, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
DATE7
2024 Physics-Informed Neural Networks: A Step Towards Data-Driven Optimization of Additive Manufacturing
abstract
Laser powder bed fusion (L-PBF) is the most popular Additive Manufacturing (AM) process for metals. It builds a 3D object layer-by-layer, by spreading metal powder on top of the previous layer and selectively melting it with a laser. Despite its many advantages, large-scale production may be hampered by the large number of process parameters and the challenges associated with their optimization. We propose an automated parameter selection approach based on process signatures extracted from a parameterized simulation of the process. Specifically, we outline a rapid data-driven simulation method based on Physics-Informed Neural Network (PINN). This approach involves training a neural network to solve the partial differential equation describing the process at varying values of a parameter of interest (for example, the laser power), thus eliminating the need for repeated Finite Elements Method (FEM) simulations. Our preliminary experiments demonstrate the feasibility of our approach.
Fabio Depaoli, Stefano Felicioni, Francesco Ponzio, Alessandro Aliberti, Enrico Macii, Federica Bondioli, Elisa Padovano, Santa Di Cataldo
ETFA5
2024 Natively Neuromorphic LMU Architecture for Encoding-Free SNN-Based HAR on Commercial Edge Devices
Vittorio Fra, Benedetto Leto, Andrea Pignata, Enrico Macii, Gianvito Urgese
ICANN (10)4
2024 Private Tensor Freezing for an Efficient Federated Learning with Homomorphic Encryption
abstract
Federated Learning (FL) is a privacy-preserving machine learning strategy where distributed clients share updates of locally trained models with a central server. The server aggregates those updates to refine a global version of the model without accessing the clients' data. Even if the raw data never leaves clients, adversarial attacks on the server side can still extract sensitive information from the transmitted model updates. Homomorphic Encryption (HE) offers a robust solution to this privacy concern: clients send encrypted model updates to the server; the server operates the aggregation of the received updates without having to decrypt. Unfortunately, HE leads to a substantial computational and communication overhead on both the client and the server side, preventing the adoption in practical, real-life applications. In this work, we introduce a training option to mitigate this problem by promoting Private Tensor Freezing (PTF), a progressive and secure gating scheme by which the number of model tensors involved in the training and synchronization stages gradually reduces over time, alleviating (i) the pressure of HE encryption/decryption on the client side, (ii) the communication volumes from/to the server, and (iii) the computing complexity of the aggregation stage on the server side. Experiments on four image classification benchmarks trained within a state-of-the-art FL framework secured with CKKS encryption reveal the effectiveness of PTF: up to 37.4% in data volume reduction, 35.5% less compute time on the client side, and 36.4% less compute time on the server side.
Valentino Peluso, Erich Malan, Andrea Calimera, Enrico Macii
ICCD4
2024 A comparative analysis of Machine Learning Techniques for short-term grid power forecasting and uncertainty analysis of Wave Energy Converters
abstract
Wave Energy is one of the renewable sources with greatest potential. Since power coming from waves fluctuates, the grid integration of wave energy involves several power conditioning stages to comply with grid quality requirements. However, to ensure full integration of wave energy in a smart grid scenario and unlock advanced monitoring and control techniques (e.g. Demand/Response), it is crucial to forecast the output power. This work proposes a methodology to forecast in short-term horizons (i.e. 15 min to 240 min) the power delivered to the grid of the Inertial Sea Wave Energy Converter (ISWEC), a device that harnesses wave power through the inertial effect of a gyroscope. Therefore, we designed, optimized and compared the performance of five known machine learning techniques for time series point forecasting: Random Forest, Support Vector Regression, Long Short-Term Memory Neural Network, Transformer Neural Network and 1 Dimensional Convolutional Neural Network. Additionally, we studied the efficacy of downsampling technique aggregating original dataset sampled every 0.1 s in time steps of 1min, 3min, 5min and 15min to compare the performance behaviour of the different machine learning models for these datasets. Furthermore, we implemented Prediction Intervals (PIs) to calculate the inherent uncertainties associated with the previously mentioned machine learning techniques. These PIs were built based on the Non-Parametric Kernel Density Estimator technique. The point forecasting and the PIs results showed that models’ performance improved as the downsampling increased. Moreover, the Random Forest model was the worst-performing in all cases. Finally, none of the other models can be considered the best overall.
Rafael Natalio Fontana Crespo, Alessandro Aliberti, Lorenzo Bottaccioli, Edoardo Pasta, Sergej Antonello Sirigu, Enrico Macii, Giuliana Mattiazzo, Edoardo Patti
Eng. Appl. Artif. Intell.6
2024 Dynamic Decision Tree Ensembles for Energy-Efficient Inference on IoT Edge Nodes
abstract
With the increasing popularity of Internet of Things (IoT) devices, there is a growing need for energy-efficient machine learning (ML) models that can run on constrained edge nodes. Decision tree ensembles, such as random forests (RFs) and gradient boosting (GBTs), are particularly suited for this task, given their relatively low complexity compared to other alternatives. However, their inference time and energy costs are still significant for edge hardware. Given that said costs grow linearly with the ensemble size, this article proposes the use of dynamic ensembles, that adjust the number of executed trees based both on a latency/energy target and on the complexity of the processed input, to tradeoff computational cost and accuracy. We focus on deploying these algorithms on multicore low-power IoT devices, designing a tool that automatically converts a Python ensemble into optimized C code, and exploring several optimizations that account for the available parallelism and memory hierarchy. We extensively benchmark both static and dynamic RFs and GBTs on three state-of-the-art IoT-relevant data sets, using an 8-core ultralow-power System-on-Chip (SoC), GAP8, as the target platform. Thanks to the proposed early stopping mechanisms, we achieve an energy reduction of up to 37.9% with respect to static GBTs (8.82 uJ versus 14.20 uJ per inference) and 41.7% with respect to static RFs (2.86 uJ versus 4.90 uJ per inference), without losing accuracy compared to the static model.
Francesco Daghero, Alessio Burrello, Enrico Macii, Paolo Montuschi, Massimo Poncino, Daniele Jahier Pagliari
IEEE Internet Things J.3
2024 Automatic Layer Freezing for Communication Efficiency in Cross-Device Federated Learning
abstract
Federated learning (FL) is a collaborative machine learning paradigm where network-edge clients train a global model under the orchestration of a central server. Unlike traditional distributed learning, each participating client keeps its data locally, ensuring privacy protection by default. However, state-of-the-art FL implementations suffer from massive information exchange between clients and the server. This issue prevents the adoption in constrained environments, typical of the Internet of Things domain, where the communication bandwidth and the energy budget are severely limited. To achieve higher efficiency at scale, the future of FL calls for additional optimizations to reach high-quality learning capability with lower communication pressure. To address this challenge, we propose automatic layer freezing (ALF), an embedded mechanism that gradually drops a growing portion of the model out of the training and synchronization phases of the learning loop, reducing the volume of exchanged data with the central server. ALF monitors the evolution of model updates and identifies layers that have reached a stable representation, where further weight updates would have minimal impact on accuracy. By freezing these layers, ALF achieves substantial savings in communication bandwidth and energy consumption. The proposed implementation of the ALF mechanism is compatible with any FL strategy, requiring minimal effort and without interfering with existing optimizations. The extensive experiments conducted using a representative set of FL strategies applied to two image classification tasks show that ALF improves the communication efficiency of the baseline FL implementations, ensuring up to 83.91% of data volume savings with no or marginal losses of accuracy.
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii, Paolo Montuschi
IEEE Internet Things J.4
2023 Comparative analysis of neural networks techniques to forecast Airfare Prices
abstract
With the growth of tourism industry, airplanes have became an affordable choice for medium- and long-distance travels. Accurate forecasting of flights tickets helps the aviation industry to match demand, supply flexibly and optimize aviation resources. Airline companies use dynamic pricing strategies to determine the price of airline tickets to maximize profits. Passengers want to purchase tickets at the lowest selling price for the flight of their choice. However, airline tickets are a special commodity that is time-sensitive and scarce, and the price of airline tickets is affected by various factors.Our research work provides a systematic comparison of various traditional machine learning methods (i.e., Ridge Regression, Lasso Regression, K-Nearest Neighbor, Decision Tree, XGBoost, Random Forest) and deep learning methods (e.g., Fully Connected Networks, Convolutional Neural Networks, Transformer) to address the problem of airfare prediction, by keeping the consumers’ needs. Moreover, we proposed innovative Bayesian neural networks, which represent the first exploitation attempt of Bayesian Inference for the airfare prediction task, to the best of our knowledge. Therefore, we evaluate the performance of our implemented and optimized models on an open dataset. The experimental results show that deep learning-based methods achieve better results on average than traditional ones, while Bayesian neural networks can achieve better performance among the other machine learning methods. However, taking into account both prediction performance and computational time, the Random Forest turns out to be the best choice to apply in this scenario.
Alessandro Aliberti, Yao Xin, Alessio Viticchié, Enrico Macii, Edoardo Patti
COMPSAC4
2023 LSTM for Grid Power Forecasting in Short-Term from Wave Energy Converters
abstract
In recent times, the consistent growth of wave energy makes it one of the most promising forms of renewable energy. Due to the intermittency and non-stationary nature of waves, the grid integration of these renewable energy sources involves a series of complex power conditioning stages to deliver grid electric power that meets the corresponding quality standards. Furthermore, to enable optimal management and operation of a smart grid power system, forecasting the wave power delivered to the grid is essential. In this paper, we present a novel approach based on Long Short-Term Memory Neural Network to forecast the wave power delivered to the grid of a Wave Energy Converter (WEC) - the ISWEC, which is a device able to harvest sea energy by exploiting the inertial effect of a gyroscope - in short-time horizons (e.g. 1min). The data for the analysis was obtained from a simulator that combines a model of the ISWEC device and the power conditioning grid integration for this particular WEC. In addition, to investigate the effectiveness of downsampling, we compared the performance behavior of the raw dataset and downsampled versions of it. The results showed that as the downsampling increases, so does the forecasting accuracy: the forecasting performance of the raw dataset returned the worst results, while the one of the dataset with the biggest downsampling studied returned the best.
Rafael Natalio Fontana Crespo, Alessandro Aliberti, Lorenzo Bottaccioli, Enrico Macii, Giorgio Fighera, Edoardo Patti
COMPSAC4
2023 Energy-efficient Wearable-to-Mobile Offload of ML Inference for PPG-based Heart-Rate Estimation
abstract
Modern smartwatches often include photoplethysmographic (PPG) sensors to measure heartbeats or blood pressure through complex algorithms that fuse PPG data with other signals. In this work, we propose a collaborative inference approach that uses both a smartwatch and a connected smartphone to maximize the performance of heart rate (HR) tracking while also maximizing the smartwatch's battery life. In particular, we first analyze the trade-offs between running on-device HR tracking or offloading the work to the mobile. Then, thanks to an additional step to evaluate the difficulty of the upcoming HR prediction, we demonstrate that we can smartly manage the workload between smartwatch and smartphone, maintaining a low mean absolute error (MAE) while reducing energy consumption. We benchmark our approach on a custom smartwatch prototype, including the STM32WB55 MCU and Bluetooth Low-Energy (BLE) communication, and a Raspberry Pi3 as a proxy for the smartphone. With our Collaborative Heart Rate Inference System (CHRIS), we obtain a set of Pareto-optimal configurations demonstrating the same MAE as State-of-Art (SoA) algorithms while consuming less energy. For instance, we can achieve approximately the same MAE of TimePPG-Small [1] (5.54 BPM MAE vs. 5.60 BPM MAE) while reducing the energy by 2.03×, with a configuration that offloads 80% of the predictions to the phone. Furthermore, accepting a performance degradation to 7.16 BPM of MAE, we can achieve an energy consumption of 179 uJ per prediction, 3.03× less than running TimePPG-Small on the smartwatch, and 1.82× less than streaming all the input data to the phone.
Alessio Burrello, Matteo Risso, Noemi Tomasello, Yukai Chen, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
DATE6
2023 A Distributed Software Platform for Additive Manufacturing
abstract
Additive Manufacturing (AM), a cornerstone of Industry 4.0, is expected to revolutionise production in practically all industries. However, multiple production challenges still exist, preventing its diffusion. In recent years, Machine Learning algorithms have been employed to overcome these hurdles. Nonetheless, the usage of these algorithms is constrained by the scarcity of data together with the challenges associated with accessing and integrating the information generated during the AM pipeline. In this work, we present a vendor-agnostic platform for AM that enables collecting, storing, analysing and linking the heterogeneous data of the complete AM process. We conducted an extensive analysis of the different AM datatypes and identified the most suitable technologies for storing them. Furthermore, we performed an in-depth study of the requirements of different AM stakeholders to develop a rich and intuitive Graphical User Interface. We showcased the specific usage of the platform for Powder Bed Fusion, one of the most popular AM processes, in a real industrial scenario, integrating specific existing modules for in-situ monitoring and real-time defect detection.
Rafael Natalio Fontana Crespo, Davide Cannizzaro, Lorenzo Bottaccioli, Enrico Macii, Edoardo Patti, Santa Di Cataldo
ETFA4
2023 Neuro-Symbolic Empowered Denoising Diffusion Probabilistic Models for Real-Time Anomaly Detection in Industry 4.0: Wild-and-Crazy-Idea Paper
abstract
Industry 4.0 involves the integration of digital technologies, such as IoT, Big Data, and AI, into manufacturing and industrial processes to increase efficiency and productivity. As these technologies become more interconnected and interdependent, Industry 4.0 systems become more complex, which brings the difficulty of identifying and stopping anomalies that may cause disturbances in the manufacturing process. This paper aims to propose a diffusion-based model for real-time anomaly prediction in Industry 4.0 processes. Using a neuro-symbolic approach, we integrate industrial ontologies in the model, thereby adding formal knowledge on smart manufacturing. Finally, we propose a simple yet effective way of distilling diffusion models through Random Fourier Features for deployment on an embedded system for direct integration into the manufacturing process. To the best of our knowledge, this approach has never been explored before.
Luigi Capogrosso, Alessio Mascolini, Federico Girella, Geri Skenderi, Sebastiano Gaiardelli, Nicola Dall'Ora, Francesco Ponzio, Enrico Fraccaroli, Santa Di Cataldo, Sara Vinco, Enrico Macii, Franco Fummi, Marco Cristani
FDL11
2023 Robotic Arm Dataset (RoAD): A Dataset to Support the Design and Test of Machine Learning-Driven Anomaly Detection in a Production Line
abstract
The early detection of anomalous behaviors from a production line is a fundamental aspect of Industry 4.0, facilitated by the collection of massive amounts of data enabled by the Industrial Internet of Things. Nonetheless, the design and validation of anomaly detection algorithms, mostly based on sophisticated Machine Learning models, heavily rely on the availability of annotated datasets of realistic anomalies, which is very difficult to obtain in a real production line. To address this problem, we introduce the Robotic Arm Dataset (RoAD), specifically designed to support the development and validation of Multivariate Time Series Anomaly Detection (MTSAD) algorithms. We collect and annotate a large number of data and metadata to characterize the motion and energy consumption of a collaborative robotic arm in a full-fledged production line and annotate a comprehensive set of healthy as well as realistic anomalies scenarios. To prove the significance of RoAD and encourage future developments, we benchmark several state-of-the-art anomaly detection algorithms on our newly introduced dataset, and we freely release it to the scientific community.
Alessio Mascolini, Sebastiano Gaiardelli, Francesco Ponzio, Nicola Dall'Ora, Enrico Macii, Sara Vinco, Santa Di Cataldo, Franco Fummi
IECON5
2023 Model-Driven Dataset Generation for Data-Driven Battery SOH Models
abstract
Estimating the State of Health (SOH) of batteries is crucial for ensuring the reliable operation of battery systems. Since there is no practical way to instantaneously measure it at run time, a model is required for its estimation. Recently, several data-driven SOH models have been proposed, whose accuracy heavily relies on the quality of the datasets used for their training. Since these datasets are obtained from measurements, they are limited in the variety of the charge/discharge profiles. To address this scarcity issue, we propose generating datasets by simulating a traditional battery model (e.g., a circuit-equivalent one). The primary advantage of this approach is the ability to use a simulatable battery model to evaluate a potentially infinite number of workload profiles for training the data-driven model. Furthermore, this general concept can be applied using any simulatable battery model, providing a fine spectrum of accuracy/complexity tradeoffs. Our results indicate that using simulated data achieves reasonable accuracy in SOH estimation, with a 7.2 % error relative to the simulated model, in exchange for a 27X memory reduction and a$\approx 2000\mathrm{X}$speedup.
Khaled Alamin, Francesco Daghero, Giovanni Pollo, Daniele Jahier Pagliari, Yukai Chen, Enrico Macii, Massimo Poncino, Sara Vinco
ISLPED6
2023 Enabling DVFS Side-Channel Attacks for Neural Network Fingerprinting in Edge Inference Services
abstract
The Inference-as-a-Service (IaaS) delivery model provides users access to pre-trained deep neural networks while safeguarding network code and weights. However, IaaS is not immune to security threats, like side-channel attacks (SCAs), that exploit unintended information leakage from the physical characteristics of the target device. Exposure to such threats grows when IaaS is deployed on distributed computing nodes at the edge. This work identifies a potential vulnerability of low-power CPUs that facilitates stealing the deep neural network architecture without physical access to the hardware or interference with the execution flow. Our approach relies on a Dynamic Voltage and Frequency Scaling (DVFS) side-channel attack, which monitors the CPU frequency state during the inference stages. Specifically, we introduce a dedicated load-testing methodology that imprints distinguishable signatures of the network on the frequency traces. A machine learning classifier is then used to infer the victim architecture. Experimental results on two commercial ARM Cortex-A CPUs, the A72 and A57, demonstrate the attack can identify the target architecture from a pool of 12 convolutional neural networks with an average accuracy of 98.7% and 92.4%
Erich Malan, Valentino Peluso, Andrea Calimera, Enrico Macii
ISLPED4
2023 Precision-aware Latency and Energy Balancing on Multi-Accelerator Platforms for DNN Inference
abstract
The need to execute Deep Neural Networks (DNNs) at low latency and low power at the edge has spurred the development of new heterogeneous Systems-on-Chips (SoCs) encapsulating a diverse set of hardware accelerators. How to optimally map a DNN onto such multi-accelerator systems is an open problem. We propose ODiMO, a hardware-aware tool that performs a fine-grain mapping across different accelerators on-chip, splitting individual layers and executing them in parallel, to reduce inference energy consumption or latency, while taking into account each accelerator's quantization precision to maintain accuracy. Pareto-optimal networks in the accuracy vs. energy or latency space are pursued for three popular dataset/DNN pairs, and deployed on the DIANA heterogeneous ultra-low power edge AI SoC. We show that ODiMO reduces energy/latency by up to 33%/31% with limited accuracy drop (−0.53%/-0.32%) compared to manual heuristic mappings.
Matteo Risso, Alessio Burrello, Giuseppe Maria Sarda, Luca Benini, Enrico Macii, Massimo Poncino, Marian Verhelst, Daniele Jahier Pagliari
ISLPED5
2023 W2WNet: A two-module probabilistic Convolutional Neural Network with embedded data cleansing functionality
Francesco Ponzio, Enrico Macii, Elisa Ficarra, Santa Di Cataldo
Expert Syst. Appl.2
2023 Efficient Deep Learning Models for Privacy-Preserving People Counting on Low-Resolution Infrared Arrays
abstract
Ultralow-resolution infrared (IR) array sensors offer a low cost, energy efficient, and privacy-preserving solution for people counting, with applications, such as occupancy monitoring and visitor flow analysis in private and public spaces. Previous work has shown that deep learning (DL) can yield superior performance on this task. However, the literature was missing an extensive comparative analysis of various efficient DL architectures for IR array-based people counting, that considers not only their accuracy but also the cost of deploying them on memory- and energy-constrained Internet of Things (IoT) edge nodes. Such analysis is key for system designers, since it helps them select the most appropriate DL model given the constraints of their target hardware. In this work, we address this need by comparing six different DL architectures on a novel data set composed of IR images collected from a commercial$8\times8$array, which we made openly available. With a wide architectural exploration of each model type, we obtain a rich set of Pareto-optimal solutions, spanning cross-validated balanced accuracy scores in the 55.70%–82.70% range. When deployed on a commercial microcontroller (MCU) by STMicroelectronics, the STM32L4A6ZG, these models occupy 0.41–9.28kB of memory, and require 1.10–7.74 ms per inference, while consuming 17.18–$120.43 \mu \text{J}$of energy. Our models are significantly more accurate than a previous deterministic method (up to +39.9%), while being up to$3.53\times $faster and more energy efficient. So, our work serves also as a demonstration that DL can not only achieve higher accuracy but also higher efficiency compared to classic algorithms for this type of task. Further, our models’ accuracy is comparable to state-of-the-art DL solutions on similar resolution sensors, despite a much lower complexity. All our models enable continuous, real-time inference on an MCU-based IoT node, with years of autonomous operation without battery recharging.
Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
IEEE Internet Things J.7
2023 Lightweight Neural Architecture Search for Temporal Convolutional Networks at the Edge
abstract
Neural Architecture Search (NAS) is quickly becoming the go-to approach to optimize the structure of Deep Learning (DL) models for complex tasks such as Image Classification or Object Detection. However, many other relevant applications of DL, especially at the edge, are based on time-series processing and require models with unique features, for which NAS is less explored. This work focuses in particular on Temporal Convolutional Networks (TCNs), a convolutional model for time-series processing that has recently emerged as a promising alternative to more complex recurrent architectures. We propose the first NAS tool that explicitly targets the optimization of the most peculiar architectural parameters of TCNs, namely dilation, receptive-field and number of features in each layer. The proposed approach searches for networks that offer good trade-offs between accuracy and number of parameters/operations, enabling an efficient deployment on embedded platforms. Moreover, its fundamental feature is that of being lightweight in terms of search complexity, making it usable even with limited hardware resources. We test the proposed NAS on four real-world, edge-relevant tasks, involving audio and bio-signals: (i) PPG-based Heart-Rate Monitoring, (ii) ECG-based Arrythmia Detection, (iii) sEMG-based Hand-Gesture Recognition, and (iv) Keyword Spotting.Results show that, starting from a single seed network, our method is capable of obtaining a rich collection of Pareto optimal architectures, among which we obtain models with the same accuracy as the seed, and 15.9-152× fewer parameters. Moreover, the NAS finds solutions that Pareto-dominate state-of-the-arthand-tuned models for 3 out of the 4 benchmarks, and are Pareto-optimal on the fourth (sEMG). Compared to three state-of-the-art NAS tools, ProxylessNAS, MorphNet and FBNetV2, our method explores a larger search space for TCNs (up to 1012×) and obtains superior solutions, while requiring low GPU memory and search time. We deploy our NAS outputs on two distinct edge devices, the multicore GreenWaves Technology GAP8 IoT processor and the single-core STMicroelectronics STM32H7 microcontroller. With respect to the state-of-the-art hand-tuned models, we reduce latency and energy of up to 5.5× and 3.8× on the two targets respectively, without any accuracy loss.
Matteo Risso, Alessio Burrello, Francesco Conti 0001, Lorenzo Lamberti, Yukai Chen, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
IEEE Trans. Computers7
2023 Clustering Appliance Operation Modes With Unsupervised Deep Learning Techniques
abstract
In smart grids, consumers can be involved in demand response programs to reduce the total power consumption of their households during the peak hours of the day. Unfortunately, nowadays, utility companies are facing important challenges in the implementation of demand response programs because of their negative impact on the comfort of end-users. In this article, we cluster the different operation modes of household appliances based on the analysis of their power signatures. For this purpose, we implement an autoencoder neural network to create a better data representation of the power signatures. Then, we cluster the different operational programs by using aK-means algorithm fitted to the new data representation. To test our methodology, we study the operation modes of some washing machines and dishwashers whose power signatures were derived from both submeters and nonintrusive load monitoring techniques. Our clustering analysis reveals the existence of multiple working programs showing well-defined features in terms of both average energy consumption and duration. Our results can then be used to improve demand response programs by reducing their impact on the comfort of end-users. Furthermore, end-users can rely on our framework to favor lighter operation modes and reduce their overall energy consumption.
Marco Castangia, Nicola Barletta, Christian Camarda, Stefano Quer, Enrico Macii, Edoardo Patti
IEEE Trans. Ind. Informatics5
2022 Comparative Analysis of Neural Networks Techniques for Lithium-ion Battery SOH Estimation
abstract
Li-ion batteries have become the most important technology for electric mobility. One of the most pressing chal-lenges is the development of reliable methods for battery state-of-health (SOH) diagnosis and estimation of remaining useful life. In electric mobility scenario, battery capacity degradation prediction is crucial to ensure service availability and life duration. This research work provides a comprehensive comparative analysis of neural networks for a data-driven approach suitable for SOH estimation on single cells, stressed under laboratory conditions. For this purpose, different neural networks (i.e., LSTM, GRU, 1D-CNN, CNN-LSTM) are trained and optimized on NASA Randomized Battery Usage dataset. Experimental results demonstrate that data-driven neural networks generally performed well SOH estimation on single cells. In detail, the 1D-CNN best predicts SOH and has the lowest variance in the output. The LSTM have the highest variance in estimating SOH, while GRU and CNN-LSTM tend to overestimate and underestimate the value of SOH, respectively.
Alessandro Aliberti, Filippo Boni, Alessandro Perol, Marco Zampolli, Rémi Jacques Philibert Jaboeuf, Paolo Tosco, Enrico Macii, Edoardo Patti
COMPSAC7
2022 A Nonlinear Two-Parameter Model for the Spatial Analysis of Solar Irradiation
abstract
Nowadays, energy estimation in various application areas is a major research topic. Additionally, various machine learning techniques, especially regression methods and artificial neural networks, have been developed in recent decades to improve the accuracy of such estimates. This article presents a nonlinear compact regression model for estimating the yearly solar irradiation in Africa and Europe by considering only the latitude and mean temperature of the locations as input parameters. The definition of the values of the coefficients is based on the least-square method constrained by the maximum absolute error. The results of 16 conventional regression models, using the same number of predictors, were compared with the result of the model proposed. Our model minimizes the root mean square error by at least 15%.
Alberto Bocca, Alberto Macii, Enrico Macii
COMPSAC3
2022 Bioformers: Embedding Transformers for Ultra-Low Power sEMG-based Gesture Recognition
abstract
Human-machine interaction is gaining traction in rehabilitation tasks, such as controlling prosthetic hands or robotic arms. Gesture recognition exploiting surface electromyographic (sEMG) signals is one of the most promising approaches, given that sEMG signal acquisition is non-invasive and is directly related to muscle contraction. However, the analysis of these signals still presents many challenges since similar gestures result in similar muscle contractions. Thus the resulting signal shapes are almost identical, leading to low classification accuracy. To tackle this challenge, complex neural networks are employed, which require large memory footprints, consume relatively high energy and limit the maximum battery life of devices used for classification. This work addresses this problem with the introduction of the Bioformers. This new family of ultra-small attention-based architectures approaches state-of-the-art performance while reducing the number of parameters and operations of 4.9 ×. Additionally, by introducing a new inter-subjects pre-training, we improve the accuracy of our best Bioformer by 3.39 %, matching state-of-the-art accuracy without any additional inference cost. Deploying our best performing Bioformer on a Parallel, Ultra-Low Power (PULP) microcontroller unit (MCU), the GreenWaves GAP8, we achieve an inference latency and energy of 2.72 ms and 0.14 mJ, respectively, 8.0× lower than the previous state-of-the-art neural network, while occupying just 94.2 kB of memory.
Alessio Burrello, Francesco Bianco Morghet, Moritz Scherer 0001, Simone Benatti, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
DATE6
2022 C-NMT: A Collaborative Inference Framework for Neural Machine Translation
abstract
Collaborative Inference (CI) optimizes the latency and energy consumption of deep learning inference through the inter-operation of edge and cloud devices. Albeit beneficial for other tasks, CI has never been applied to the sequence-to-sequence mapping problem at the heart of Neural Machine Translation (NMT). In this work, we address the specific issues of collaborative NMT, such as estimating the latency required to generate the (unknown) output sequence, and show how existing CI methods can be adapted to these applications. Our experiments show that CI can reduce the latency of NMT by up to 44% compared to a non-collaborative approach.
Yukai Chen, Roberta Chiaro, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
ISCAS3
2022 Privacy-preserving Social Distance Monitoring on Microcontrollers with Low-Resolution Infrared Sensors and CNNs
abstract
Low-resolution infrared (IR) array sensors offer a low-cost, low-power, and privacy-preserving alternative to optical cameras and smartphones/wearables for social distance monitoring in indoor spaces, permitting the recognition of basic shapes, without revealing the personal details of individuals. In this work, we demonstrate that an accurate detection of social distance violations can be achieved processing the raw output of a 8x8 IR array sensor with a small-sized Convolutional Neural Network (CNN). Furthermore, the CNN can be executed directly on a Microcontroller (MCU)-based sensor node.With results on a newly collected open dataset, we show that our best CNN achieves 86.3% balanced accuracy, significantly outperforming the 61% achieved by a state-of-the-art deterministic algorithm. Changing the architectural parameters of the CNN, we obtain a rich Pareto set of models, spanning 70.5-86.3% accuracy and 0.18-75k parameters. Deployed on a STM32L476RGMCU, these models have a latency of 0.73-5.33ms, with an energy consumption per inference of 9.38-68.57$\mu$J.
Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
ISCAS7
2022 Multi-Complexity-Loss DNAS for Energy-Efficient and Memory-Constrained Deep Neural Networks
abstract
Neural Architecture Search (NAS) is increasingly popular to automatically explore the accuracy versus computational complexity trade-off of Deep Learning (DL) architectures. When targeting tiny edge devices, the main challenge for DL deployment is matching the tight memory constraints, hence most NAS algorithms consider model size as the complexity metric. Other methods reduce the energy or latency of DL models by trading off accuracy and number of inference operations. Energy and memory are rarely considered simultaneously, in particular by low-search-cost Differentiable NAS (DNAS) solutions.
Matteo Risso, Alessio Burrello, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
ISLPED4
2022 High Resolution Explanation Maps for CNNs using Segmentation Networks
abstract
Recent developments have resulted in multiple techniques trying to explain how deep neural networks achieve their predictions. The explainability maps provided by such techniques are useful to understand what the network has learned and increase user confidence in critical applications such as the medical field or autonomous driving. Nonetheless, they typically have very low resolutions, severely limiting their capability of identifying finer details or multiple subjects. In this paper we employ an encoder-decoder architecture with skip connection known as U-Net, originally developed for segmenting medical images, as an image classifier and we show that state of the art explainable techniques applied to U-Net can generate pixel level explanation maps for images of any resolution.
Alessio Mascolini, Francesco Ponzio, Enrico Macii, Elisa Ficarra, Santa Di Cataldo
VL/HCC3
2022 Solar radiation forecasting with deep learning techniques integrating geostationary satellite images
Raimondo Gallo, Marco Castangia, Alberto Macii, Enrico Macii, Edoardo Patti, Alessandro Aliberti
Eng. Appl. Artif. Intell.4
2022 Human Activity Recognition on Microcontrollers with Quantized and Adaptive Deep Neural Networks
abstract
Human Activity Recognition (HAR) based on inertial data is an increasingly diffused task on embedded devices, from smartphones to ultra low-power sensors. Due to the high computational complexity of deep learning models, most embedded HAR systems are based on simple and not-so-accurate classic machine learning algorithms. This work bridges the gap between on-device HAR and deep learning, proposing a set of efficient one-dimensional Convolutional Neural Networks (CNNs) that can be deployed on general purpose microcontrollers (MCUs). Our CNNs are obtained combining hyper-parameters optimization with sub-byte and mixed-precision quantization, to find good trade-offs between classification results and memory occupation. Moreover, we also leverage adaptive inference as an orthogonal optimization to tune the inference complexity at runtime based on the processed input, hence producing a more flexible HAR system. With experiments on four datasets, and targeting an ultra-low-power RISC-V MCU, we show that (i) we are able to obtain a rich set of Pareto-optimal CNNs for HAR, spanning more than 1 order of magnitude in terms of memory, latency, and energy consumption; (ii) thanks to adaptive inference, we can derive >20 runtime operating modes starting from a single CNN, differing by up to 10% in classification scores and by more than 3× in inference complexity, with a limited memory overhead; (iii) on three of the four benchmarks, we outperform all previous deep learning methods, while reducing the memory occupation by more than 100×. The few methods that obtain better performance (both shallow and deep) are not compatible with MCU deployment; (iv) all our CNNs are compatible with real-time on-device HAR, achieving an inference latency that ranges between 9 μs and 16 ms. Their memory occupation varies in 0.05–23.17 kB, and their energy consumption in 0.05 and 61.59 μJ, allowing years of continuous operation on a small battery supply.
Francesco Daghero, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
ACM Trans. Embed. Comput. Syst.7
2021 Ultra-compact binary neural networks for human activity recognition on RISC-V processors
abstract
Human Activity Recognition (HAR) is a relevant inference task in many mobile applications. State-of-the-art HAR at the edge is typically achieved with lightweight machine learning models such as decision trees and Random Forests (RFs), whereas deep learning is less common due to its high computational complexity. In this work, we propose a novel implementation of HAR based on deep neural networks, and precisely on Binary Neural Networks (BNNs), targeting low-power general purpose processors with a RISC-V instruction set. BNNs yield very small memory footprints and low inference complexity, thanks to the replacement of arithmetic operations with bit-wise ones. However, existing BNN implementations on general purpose processors impose constraints tailored to complex computer vision tasks, which result in over-parametrized models for simpler problems like HAR. Therefore, we also introduce a new BNN inference library, which targets ultra-compact models explicitly. With experiments on a single-core RISC-V processor, we show that BNNs trained on two HAR datasets obtain higher classification accuracy compared to a state-of-the-art baseline based on RFs. Furthermore, our BNN reaches the same accuracy of a RF with either less memory (up to 91%) or more energy-efficiency (up to 70%), depending on the complexity of the features extracted by the RF.
Francesco Daghero, Daniele Jahier Pagliari, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino
CF8
2021 Forecasting the Grid Power Demand of Charging Stations from EV Drivers' Attitude
abstract
In recent years there has been a significant increase in the production of electric vehicles (EVs), in the global strive to reduce polluting gases produced by conventional fossil-fuel driven vehicles. Therefore, many optimization algorithms have been proposed for EV mobility and the charging of battery packs in the stations connected to power grids. However, there are situations in which experimental results are not sufficient, and simulations are needed. In this work, we address the effects of the charge demands of an EV fleet on the grid by considering the attitude of EV drivers, and especially their range anxiety. This influences their decision of when to recharge the battery pack. To this end, an agent-based model has been developed for the simulation of a power grid considering different scenarios based mainly on the state of charge (SOC) of battery packs at the time of the charging requests of EVs at service stations. The results indicate that in general a high battery SOC at the beginning of charging increases the probability of reaching higher power peaks on the grid.
Alberto Bocca, Alberto Macii, Enrico Macii
COMPSAC3
2021 Design of District-level Photovoltaic Installations for Optimal Power Production and Economic Benefit
abstract
PhotoVoltaic (PV) installations are a widespread source of renewable energy, and are quite common urban buildings’ roofs. To soften both the initial investment and the recurrent maintenance costs, the current market trends delegate the construction of PV installations to Energy Aggregators, i.e., grouping of consumers and producers that act as a single entity to satisfy local energy demand and to sell the surplus energy to the grid. In this perspective, PV installations can be designed with a larger perspective, i.e., at district level, to maximize power production not of a single building but rather of a number of blocks of a city. This implies new challenges, including efficient data management (the covered area can be squared kilometers wide) and optimal PV installation (the number of PV modules can be in the order of hundreds or even thousands). This paper proposes a framework to combine detailed geographic and irradiance information to determine an optimal PV installation over a district, by maximizing both power production and economic convenience. Our simulation results run on a real-world district prove that the framework allows an advanced evaluation of costs and benefit, that can be used by Energy Aggregators to design a new PV installation, and demonstrate an improvement on power generation up to 20% w.r.t. standard installations.
Matteo Orlando, Lorenzo Bottaccioli, Sara Vinco, Enrico Macii, Massimo Poncino, Edoardo Patti
COMPSAC4
2021 Pruning In Time (PIT): A Lightweight Network Architecture Optimizer for Temporal Convolutional Networks
abstract
Temporal Convolutional Networks (TCNs) are promising Deep Learning models for time-series processing tasks. One key feature of TCNs is time-dilated convolution, whose optimization requires extensive experimentation. We propose an automatic dilation optimizer, which tackles the problem as a weight pruning on the time-axis, and learns dilation factors together with weights, in a single training. Our method reduces the model size and inference latency on a real SoC hardware target by up to 7.4× and 3×, respectively with no accuracy drop compared to a network without dilation. It also yields a rich set of Pareto-optimal TCNs starting from a single model, outperforming hand-designed solutions in both size and accuracy.
Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari, Francesco Conti 0001, Lorenzo Lamberti, Enrico Macii, Luca Benini, Massimo Poncino
DAC6
2021 Image analytics and machine learning for in-situ defects detection in Additive Manufacturing
abstract
In the context of Industry 4.0, metal Additive Manufacturing (AM) is considered a promising technology for medical, aerospace and automotive fields. However, the lack of assurance of the quality of the printed parts can be an obstacle for a larger diffusion in industry. To this date, AM is most of the times a trial-and-error process, where the faulty artefacts are detected only after the end of part production. This impacts on the processing time and overall costs of the process. A possible solution to this problem is the in-situ monitoring and detection of defects, taking advantage of the layer-by-layer nature of the build. In this paper, we describe a system for in-situ defects monitoring and detection for metal Powder Bed Fusion (PBF), that leverages an off-axis camera mounted on top of the machine. A set of fully automated algorithms based on Computer Vision and Machine Learning allow the timely detection of a number of powder bed defects and the monitoring of the object's profile for the entire duration of the build.
Davide Cannizzaro, Antonio Giuseppe Varrella, Stefano Paradiso, Roberta Sampieri, Enrico Macii, Edoardo Patti, Santa Di Cataldo
DATE5
2021 A Bayesian approach to Expert Gate Incremental Learning
abstract
Incremental learning involves Machine Learning paradigms that dynamically adjust their previous knowledge whenever new training samples emerge. To address the problem of multi-task incremental learning without storing any samples of the previous tasks, the so-called Expert Gate paradigm was proposed, which consists of a Gate and a downstream network of task-specific CNNs, a.k.a. the Experts. The gate forwards the input to a certain expert, based on the decision made by a set of autoencoders. Unfortunately, as a CNN is intrinsically incapable of dealing with inputs of a class it was not specifically trained on, the activation of the wrong expert will invariably end into a classification error. To address this issue, we propose a probabilistic extension of the classic Expert Gate paradigm. Exploiting the prediction uncertainty estimations provided by Bayesian Convolutional Neural Networks (B-CNNs), the proposed paradigm is able to either reduce, or correct at a later stage, wrong decisions of the gate. The goodness of our approach is shown by experimental comparisons with state-of-the-art incremental learning methods.
Valerio Mieuli, Francesco Ponzio, Alessio Mascolini, Enrico Macii, Elisa Ficarra, Santa Di Cataldo
IJCNN4
2021 Robust and Energy-Efficient PPG-Based Heart-Rate Monitoring
abstract
A wrist-worn PPG sensor coupled with a lightweight algorithm can run on a MCU to enable non-invasive and comfortable monitoring, but ensuring robust PPG-based heart-rate monitoring in the presence of motion artifacts is still an open challenge. Recent state-of-the-art algorithms combine PPG and inertial signals to mitigate the effect of motion artifacts. However, these approaches suffer from limited generality. Moreover, their deployment on MCU-based edge nodes has not been investigated. In this work, we tackle both the aforementioned problems by proposing the use of hardware-friendly Temporal Convolutional Networks (TCN) for PPG-based heart estimation. Starting from a single "seed" TCN, we leverage an automatic Neural Architecture Search (NAS) approach to derive a rich family of models. Among them, we obtain a TCN that outperforms the previous state-of-the- art on the largest PPG dataset available (PPGDalia), achieving a Mean Absolute Error (MAE) of just 3.84 Beats Per Minute (BPM). Furthermore, we tested also a set of smaller yet still accurate (MAE of 5.64 - 6.29 BPM) networks that can be deployed on a commercial MCU (STM32L4) which require as few as 5k parameters and reach a latency of 17.1 ms consuming just 0.21 mJ per inference.
Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari, Simone Benatti, Enrico Macii, Luca Benini, Massimo Poncino
ISCAS5
2021 TCN Mapping Optimization for Ultra-Low Power Time-Series Edge Inference
abstract
Temporal Convolutional Networks (TCNs) are emerging lightweight Deep Learning models for Time Series analysis. We introduce an automated exploration approach and a library of optimized kernels to map TCNs on Parallel Ultra-Low Power (PULP) microcontrollers. Our approach minimizes latency and energy by exploiting a layer tiling optimizer to jointly find the tiling dimensions and select among alternative implementations of the causal and dilated 1D-convolution operations at the core of TCNs. We benchmark our approach on a commercial PULP device, achieving up to $103 \times $ lower latency and $20.3 \times $ lower energy than the Cube-AI toolkit executed on the STM32L4 and from $2.9 \times $ to $26.6 \times $ lower energy compared to commercial closed-source and academic open-source approaches on the same hardware target.
Alessio Burrello, Alberto Dequino, Daniele Jahier Pagliari, Francesco Conti 0001, Marcello Zanghieri, Enrico Macii, Luca Benini, Massimo Poncino
ISLPED6
2021 ACME: An Energy-Efficient Approximate Bus Encoding for I2C
abstract
In ultra low power systems with many peripherals, off-chip serial interconnects contribute significantly to the total energy budget. Leveraging the error-resilience characteristics of many embedded applications, the approximate computing paradigm has been applied to serial bus encodings to reduce interconnect consumption. However, the power model considered in previous works was purely capacitive. Accordingly, the objective of these approximate encodings was simply to reduce the transition count. While this works well for most bus standards, one notable exception is represented by I2C, whose open-drain physical connection makes the static energy consumed by logic-0 values on the bus extremely relevant. In this work, we propose ACME, the first approximate serial bus encoding targeting specifically I2C connections. With a simple encoding/decoding scheme, ACME concurrently reduces both the static and dynamic energy on the bus by maximizing the number of logic-1 values in codewords, while simultaneously reducing transitions. Using an accurate bus model and realistic capacitance and resistance values selected according to the I2C standard, we show that our encoding outperforms state-of-the-art solutions and reduces the total energy consumption on the bus by 57% on average, with an error smaller than 0.1%.
Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino
ISLPED4
2021 Adaptive Random Forests for Energy-Efficient Inference on Microcontrollers
abstract
Random Forests (RFs) are widely used Machine Learning models in low-power embedded devices, due to their hardware friendly operation and high accuracy on practically relevant tasks. The accuracy of a RF often increases with the number of internal weak learners (decision trees), but at the cost of a proportional increase in inference latency and energy consumption. Such costs can be mitigated considering that, in most applications, inputs are not all equally difficult to classify. Therefore, a large RF is often necessary only for (few) hard inputs, and wasteful for easier ones. In this work, we propose an early-stopping mechanism for RFs, which terminates the inference as soon as a high-enough classification confidence is reached, reducing the number of weak learners executed for easy inputs. The early-stopping confidence threshold can be controlled at runtime, in order to favor either energy saving or accuracy. We apply our method to three different embedded classification tasks, on a single-core RISC-V microcontroller, achieving an energy reduction from 38% to more than 90% with a drop of less than 0.5% in accuracy. We also show that our approach outperforms previous adaptive ML methods for RFs.
Francesco Daghero, Alessio Burrello, Luca Benini, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari
VLSI-SoC6
2021 AdapTTA: Adaptive Test-Time Augmentation for Reliable Embedded ConvNets
abstract
Convolutional Neural Networks (ConvNets) are trained offline using the few available data and may therefore suffer from substantial accuracy loss when ported on the field, where unseen input patterns received under unpredictable external conditions can mislead the model. Test-Time Augmentation (TTA) techniques aim to alleviate such common side effect at inference-time, first running multiple feed-forward passes on a set of altered versions of the same input sample, and then computing the main outcome through a consensus of the aggregated predictions. Unfortunately, the implementation of TTA on embedded CPUs introduces latency penalties that limit its adoption on edge applications. To tackle this issue, we propose AdapTTA, an adaptive implementation of TTA that controls the number of feed-forward passes dynamically, depending on the complexity of the input. Experimental results on state-of-the-art ConvNets for image classification deployed on a commercial ARM Cortex-A CPU demonstrate AdapTTA reaches remarkable latency savings, from $1.40 \times$ to $2.21 \times$, and hence a higher frame rate compared to static TTA, still preserving the same accuracy gain.
Luca Mocerino, Roberto Giorgio Rizzo, Valentino Peluso, Andrea Calimera, Enrico Macii
VLSI-SoC5
2021 Solar radiation forecasting based on convolutional neural network and ensemble learning
Davide Cannizzaro, Alessandro Aliberti, Lorenzo Bottaccioli, Enrico Macii, Andrea Acquaviva, Edoardo Patti
Expert Syst. Appl.4
2021 A compound of feature selection techniques to improve solar radiation forecasting
Marco Castangia, Alessandro Aliberti, Lorenzo Bottaccioli, Enrico Macii, Edoardo Patti
Expert Syst. Appl.4
2021 Enhancing manufacturing intelligence through an unsupervised data-driven methodology for cyclic industrial processes
Tania Cerquitelli, Francesco Ventura, Daniele Apiletti, Elena Baralis, Enrico Macii, Massimo Poncino
Expert Syst. Appl.5
2021 Leading Information and Communication Technologies for Smart Manufacturing: Facing the New Challenges and Opportunities of the 4th Industrial Revolution
abstract
The first three industrial revolutions came about as a result of mechanization, electricity, and information technology (IT), respectively. Now, the introduction of the Internet of Things and Services into the manufacturing environment is fostering a 4th industrial revolution (Industry 4.0), where the smart optimization and computerization of all the actors and phases of the manufacturing process, including the conceptualization and design of a product, as well as its production and transaction, are taking a leading role.
Santa Di Cataldo, Sukhan Lee 0001, Enrico Macii, Birgit Vogel-Heuser
Proc. IEEE3
2021 Optimizing Quality Inspection and Control in Powder Bed Metal Additive Manufacturing: Challenges and Research Directions
abstract
One of the key targets of Industry 4.0 and digital production, in general, is the support of faster, cleaner, and increasingly customizable manufacturing processes. Additive manufacturing (AM) is a natural fit in this context, as it offers the possibility to produce complex parts without the design constraints of traditional manufacturing routes, typically reducing both material waste and time to market. Nonetheless, the lack of repeatability of the manufacturing process, which typically translates into a lack of reproducibility and reliability of the quality of the final products compared to traditional subtractive technologies, is currently one of the major barriers to the widespread adoption of AM in mass production. To overcome this limitation, there are growing efforts in recent years toward better integration of advanced information technologies into AM, exploiting the layer-by-layer nature of the build. The consequence of these efforts is twofold: 1) the integration of advanced sensing technologies into the AM systems, making possible the in situ monitoring of huge amounts of data at multiple time scales and resolutions and 2) the ever-increasing role of data-driven approaches [especially machine learning (ML)] in the analysis of such data to provide real-time quality monitoring and process optimization. This article introduces and reviews the key technological developments of this phenomenon, with a special focus on metal powder bed fusion (PBF) technologies that are attracting the highest attention by the industrial AM community. After introducing the main manufacturing quality issues and needs that have to be developed and optimized, we provide a wide overview of the latest progress of in situ monitoring and control in metal PBF, with special regards to sensing technologies and ML approaches. Finally, we identify the open challenges and future research directions in this field.
Santa Di Cataldo, Sara Vinco, Gianvito Urgese, Flaviana Calignano, Elisa Ficarra, Alberto Macii, Enrico Macii
Proc. IEEE7
2021 CRIME: Input-Dependent Collaborative Inference for Recurrent Neural Networks
abstract
The excellent accuracy of Recurrent Neural Networks (RNNs) for time-series and natural language processing comes at the cost of computational complexity. Therefore, the choice between edge and cloud computing for RNN inference, with the goal of minimizing response time or energy consumption, is not trivial. An edge approach must deal with the aforementioned complexity, while a cloud solution pays large time and energy costs for data transmission. Collaborative inference is a technique that tries to obtain the best of both worlds, by splitting the inference task among a network of collaborating devices. While already investigated for other types of neural networks, collaborative inference for RNNs poses completely new challenges, such as the strong influence of input length on processing time and energy, and is greatly unexplored. In this paper, we introduce a Collaborative RNN Inference Mapping Engine(CRIME), which automatically selects the best inference device for each input. CRIME is flexible with respect to the connection topology among collaborating devices, and adapts to changes in the connections statuses and in the devices loads. With experiments on several RNNs and datasets, we show that CRIME can reduce the execution time (or end-node energy) by more than 25% compared to any single-device approach.
Daniele Jahier Pagliari, Roberta Chiaro, Enrico Macii, Massimo Poncino
IEEE Trans. Computers3
2021 Supporting Telecommunication Alarm Management System With Trouble Ticket Prediction
abstract
Fault alarm data emanated from heterogeneous telecommunication network services and infrastructures are exploding with network expansions. Managing and tracking the alarms with trouble tickets using manual or expert rule-based methods have become challenging due to increase in the complexity of alarm management systems and demand for deployment of highly trained experts. As the size and complexity of networks hike immensely, identifying semantically identical alarms, generated from heterogeneous network elements from diverse vendors, with data-driven methodologies, has become imperative to enhance efficiency. In this article, data-driven trouble ticket prediction models are proposed to leverage alarm management systems. To improve performance, feature extraction, using a sliding time window and feature engineering, from related history alarm streams, is also introduced. The models were trained and validated with a data set provided by the largest telecommunication provider in Italy. The experimental results showed the promising efficacy of the proposed approach in suppressing false positive alarms with trouble ticket prediction.
Mulugeta Weldezgina Asres, Million Abayneh Mengistu, Pino Castrogiovanni, Lorenzo Bottaccioli, Enrico Macii, Edoardo Patti, Andrea Acquaviva
IEEE Trans. Ind. Informatics5
2021 A Microservices-Based Framework for Smart Design and Optimization of PV Installations
abstract
The design of photovoltaic (PV) installations mostly relies on rule-of-thumb criteria and on gross estimates of the shading patterns, and the few optimized approaches are generally focused on the problem of identifying the most suitable surfaces (e.g., roofs) in a larger geographic area (e.g., city or district). This article proposes a framework to address the design and the optimization of PV installations through a set of microservices focusing on the different variables of the design: identification of the target surfaces, elaboration of weather data, modeling of the PV panel, and floorplanning of the panel on the surface. The microservices architecture ensures extensibility and generality, as the user may execute only a subset of the proposed services or provide novel algorithms to extend the existing ones. Additionally, the framework provides a set of built-in models that allow sensitivity to the distribution of shades and accurate modeling of the power production over time. We show the many benefits of the proposed framework on two different use cases.
Sara Vinco, Daniele Jahier Pagliari, Lorenzo Bottaccioli, Edoardo Patti, Enrico Macii, Massimo Poncino
IEEE Trans. Sustain. Comput.5
2020 GAMES: A General-Purpose Architectural Model for Multi-energy System Engineering Applications
abstract
The growing interest in Multi-Energy Systems (MES) leads the scientific community to implement innovative technologies to analyse and simulate these complex systems. Two main research trends are identified in such analysis: i) improve the usability and capability of preexisting reference architectures in the energy field to cope with high-level use case descriptions, and ii) study the interoperability of such reference architectures in order to increase systematic and functional analysis of MES use cases. GAMES is a a general-purpose architectural model for MES engineering application. The aim is twofold: i) GAMES implements an extension of Smart Grid Architecture Model (SGAM) to cope with MES use case descriptions, and ii) it offers a methodology to deal with a systemic description of the use case through a combination of UML and SysML integrated in the proposed architectural model. Furthermore, GAMES will allow the implementation of Domain Specific Language (DSL) and hardware configuration for the specific components described by UML/SysML diagrams. Compared to other solutions, GAMES allows to assess both research trends in a single hierarchical ICT infrastructure.
Luca Barbierato, Daniele Salvatore Schiera, Edoardo Patti, Enrico Macii, Enrico Pons, Ettore Bompard, Andrea Lanzini, Romano Borchiellini, Lorenzo Bottaccioli
COMPSAC4
2020 Optimal Configuration and Placement of PV Systems in Building Roofs with Cost Analysis
abstract
Following the Smart Grid view, current energy generation systems based on fossil fuels will be replaced with renewable energy sources. Photovoltaic (PV) is currently considered the most promising technology, due to decreasing costs of the devices and to the limited invasiveness in existing infrastructures, that make PV installations quite common urban buildings' roofs. To maximise both power production and Return Of Investment (ROI) of PV installations, new techniques and methodologies should be applied to limit sources of inefficiencies, like shading and power losses due to an incorrect installation. In this paper, we propose a novel solution for an optimal configuration and placement of PV systems in buildings' roofs. Given a number of alternative configurations and a roof of interest, it combines detailed geographic and irradiance information to determine the optimal PV installation, by maximizing both power production and ROI. Our simulation results on two real-world roofs demonstrate an improvement on power generation up to 23% w.r.t. standard compact installations. These results also highlight that a cost analysis, often ignored by standard installation strategies, is nonetheless necessary to guarantee optimal results in terms of PV production and revenue.
Matteo Orlando, Lorenzo Bottaccioli, Edoardo Patti, Enrico Macii, Sara Vinco, Massimo Poncino
COMPSAC4
2020 Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks Inference
abstract
Given the computational complexity of Recurrent Neural Networks (RNNs) inference, IoT and mobile devices typically offload this task to the cloud. However, the execution time and energy consumption of RNN inference strongly depends on the length of the processed input. Therefore, considering also communication costs, it may be more convenient to process short input sequences locally and only offload long ones to the cloud. In this paper, we propose a low-overhead runtime tool that performs this choice automatically. Results based on real edge and cloud devices show that our method is able to simultaneously reduce the total execution time and energy consumption of the system compared to solutions that run RNN inference fully locally or fully in the cloud.
Daniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino
DAC5
2020 A Diode-Aware Model of PV Modules from Datasheet Specifications
abstract
Semi-empirical models of photovoltaic (PV) modules based only on datasheet information are popular in electrical energy systems (EES) simulation because they can be built without measurements and allow quick exploration of alternative devices. One key limitation of these models, however, is the fact that they cannot model the presence of bypass diodes, which are inserted across a set of series-connected cells in a PV module to mitigate the impact of partial shading; datasheet information refer in fact to the operations of the module under uniform irradiance. Neglecting the effect of bypass diodes may incur in significant underestimation of the extracted power.This paper proposes a semi-empirical model for a PV module, that, by taking into account the only available information about bypass diodes in a datasheet, i.e., its number, by a first downscaling the model to a single PV cell and a subsequent upscaling to the level of a substring and of a module, allows to take into accout the diode effect as much accurately as allowed by the datasheet information.Experimental results show that, in a typical PV array on a roof, using a diode-agnostic model can signifantly underestimate the output power production.
Sara Vinco, Yukai Chen, Enrico Macii, Massimo Poncino
DATE3
2020 An Engineering Process model for managing a digitalised life-cycle of products in the Industry 4.0
abstract
The Internet of Things (IoT), and more specifically the industrial IoT, is revolutionising industry. This technology has catalyzed the fourth industrial revolution and inspired movements such as Industry 4.0, the Industrial Internet Consortium and Society 5.0. Morphing an industrial process or assembly line to aggregate Internet-connected devices and systems does not complete the picture. The concept penetrates all aspects of the engineering process (EP) which encompasses the full life-cycle of the product/solution. Phases of the EP traditionally tended to be sequential but, with the IoT, can now evolve and influence other phases throughout the product/solution life-cycle. The EU-funded Arrowhead Tools project aims to promote a service-oriented architecture (SOA) to allow tools within each phase of the engineering process to interact with each other. This paper, applies the proposed EP model to a real value chain composed of multiple stakeholders adopting different EPs for the life-cycle management of a Smart Boiler System.
Gianvito Urgese, Paolo Azzoni, Jan van Deventer, Jerker Delsing, Enrico Macii
NOMS5
2020 Optimization Tools for ConvNets on the Edge
abstract
The shift of Convolutional Neural Networks (ConvNets) into low-power devices with limited compute and memory resources calls for cross-layer strategies spanning from hardware to software optimization. This work answers to this need, presenting a collection of tools for efficient deployment of ConvNets on the edge,
Valentino Peluso, Enrico Macii, Andrea Calimera
VLSI-SOC2
2020 Logic Synthesis of Pass-Gate Logic Circuits With Emerging Ambipolar Technologies
abstract
Emerging devices and new ultrascaled silicon transistors have shown disruptive electrical and functional properties that might bring digital hardware to the next level. The key issue today concerns their integration. Even though the classical complementary logic style is the most intuitive option, other strategies such as pass-transistors that were discarded in the past because they did not fit silicon MOSFETs logic should be reconsidered. Obviously, the assessment of such alternatives requires customized CAD tools and optimization engines. The objective of this paper is to introduce a synthesis and optimization flow for pass-gate logic circuits mapped onto emerging ambipolar technologies. As main contributions we propose: 1) a novel EXNOR-based decomposition technique that fully exploits do not care conditions to generate compact logic function representations and 2) a dedicated one-pass synthesis flow where optimization and technology mapping are concurrently run on a common data structure, the reduced ordered pass-diagram. Experimental results demonstrate that the proposed flow outperforms existing synthesis tools by achieving more compact circuit representations with 8.5× less devices and about 8× shallower structures (on average), while still yielding lower CPU times.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Modeling and Simulation of Cyber-Physical Electrical Energy Systems With SystemC-AMS
abstract
Modern cyber-physical electrical energy systems (CPEES) are characterized by wider adoption of sustainable energy sources and by an increased attention to optimization, with the goal of reducing pollution and wastes. This imposes a need for instruments supporting the design flow, to simulate and validate the behavior of system components and to apply additional optimization and exploration steps. Additionally, each system might be tested with a number of management policies, to evaluate their economic impact. It is thus evident that simulation is a key ingredient in the design flow of CPEES. This paper proposes a framework for CPEES modeling and simulation, that relies on the open-source standard SystemC-AMS. The paper formalizes the information and energy flow in a generic CPEES, by focusing on both AC and DC components, and by including support for mechanical and physical models that represent multiple energy sources and loads. Experimental results, applied to a complex CPEES case study, will prove the effectiveness of the proposed solution, in terms of accuracy, speed up w.r.t. the current state-of-the-art Matlab/Simulink, and support for the design flow.
Yukai Chen, Sara Vinco, Daniele Jahier Pagliari, Paolo Montuschi, Enrico Macii, Massimo Poncino
IEEE Trans. Sustain. Comput.5
2019 Code Mapping in Heterogeneous Platforms Using Deep Learning and LLVM-IR
abstract
Modern heterogeneous platforms require compilers capable of choosing the appropriate device for the execution of program portions. This paper presents a machine learning method designed for supporting mapping decisions through the analysis of the program source code represented in LLVM assembly language (IR) for exploiting the advantages offered by this generalised and optimised representation. To evaluate our solution, we trained an LSTM neural network on OpenCL kernels compiled in LLVM-IR and processed with our tokenizer capable of filtering less-informative tokens. We tested the network that reaches an accuracy of 85% in distinguishing the best computational unit.
Francesco Barchi, Gianvito Urgese, Enrico Macii, Andrea Acquaviva
DAC3
2019 Low-Overhead Power Trace Obfuscation for Smart Meter Privacy
abstract
Smart meters communicate to the utility provider fine-grain information about a user's energy consumption, which could be used to infer the user's habits and pose thus a critical privacy risk. State-of-the-art solutions try to obfuscate the readings of a meter either by using a large re-chargeable battery to filter the trace or by adding random noise to alter it. Both solutions, however, have significant drawbacks: large batteries are prohibitively expensive, whereas digitally added noise implies that the user entrusts the utility provider to protect his/her privacy.
Daniele Jahier Pagliari, Sara Vinco, Enrico Macii, Massimo Poncino
DAC3
2019 Irradiance-Driven Partial Reconfiguration of PV Panels
abstract
Adaptive reconfiguration of a photo-voltaic (PV) panel by means of a switch network is a well-known approach to tackle shading issues dynamically and with a reasonable cost. Most of these approaches assume however that the entire panel is reconfigurable, resulting in high installation costs due to the large wiring overhead required by this solution. In this work we propose an architecture in which only a portion of the panel is made reconfigurable, while minimizing the loss in the extracted power with respect to a fully reconfigurable solution. The key feature of our approach is the use of environmental (irradiance and temperature) data to determine the reconfigurable subset at design time. Simulation results show that, by reconfiguring only about 50-70% of a panel, it is possible to achieve up to 45% power increase with respect to a static topology, while losing less than 5% power with respect to full reconfiguration.
Daniele Jahier Pagliari, Sara Vinco, Enrico Macii, Massimo Poncino
DATE3
2019 Dynamic Beam Width Tuning for Energy-Efficient Recurrent Neural Networks
abstract
Recurrent Neural Networks (RNNs) are state-of-the-art models for many machine learning tasks, such as language modeling and machine translation. Executing the inference phase of a RNN directly in edge nodes, rather than in the cloud, would provide benefits in terms of energy consumption, latency and network bandwidth, provided that models can be made efficient enough to run on energy-constrained embedded devices.
Daniele Jahier Pagliari, Francesco Panini, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2019 Battery-Aware Electric Truck Delivery Route Planner
abstract
Finding the energy-optimal route in the context of parcel delivery with electric vehicles (EVs) is more complicated than for conventional internal combustion engine (ICE) vehicles, where the energy cost of a path is mostly determined by the total traveled distance. In the case of EV delivery, the total energy consumption strongly depends on the order of delivery because the efficiency of the EV is affected by how the transported weight changes over time as it directly affects the battery efficiency. This makes impossible to find an optimal solution using traditional routing algorithms such as the traveling salesman problem (TSP) using a static quantity (e.g., distance) as a metric.In this paper, we propose a solution for the least-energy delivery problem using EVs; we implement an electric truck simulator and evaluate different static metrics to assess their quality on small size instances for which the optimal solution can be computed exhaustively. A greedy algorithm using the empirically best metric (namely, distance × residual weight) provides significant reductions (up to 33%) with respect to a common-sense heaviest first package delivery route determined using a metric suggested by the battery properties, and is sensibly faster than state-of-the-art TSP heuristic algorithms.
Donkyu Baek, Yukai Chen, Enrico Macii, Massimo Poncino, Naehyuck Chang
ISLPED3
2019 CNN-Based Camera-less User Attention Detection for Smartphone Power Management
abstract
The many sensors hosted by mobile electronic devices are commonly used to recognize user activities and context, in order to provide new functionalities, such as tracking physical activity and sleep cycles. Despite its potential, such context recognition is only employed for power management purposes in very specific scenarios (e.g. in-pocket detection). In this work we present a novel context recognition system able to reliably identify whether a mobile device is not being looked at, and to consequently trigger power management actions such as turning off the display and moving to suspended mode. Our method takes as input the readings from common low-power sensors present in virtually all mobile devices and classifies them using a Convolutional Neural Network. Most importantly, the power-hungry camera sub-system is not used, resulting in an extremely energy-efficient detection strategy. Results show that our system is able to identify scenarios in which a device is not being used with 95.6% accuracy, thus reducing the energy overheads by 91% compared to a standard timeout-based power management and by 58% compared to a system relying on the camera.
Daniele Jahier Pagliari, Matteo Ansaldi, Enrico Macii, Massimo Poncino
ISLPED3
2019 Fine-Grain Back Biasing for the Design of Energy-Quality Scalable Operators
abstract
Energy-quality scalable systems are a promising solution to cope with the small energy budgets and high processing demands of mobile and Internet of Things applications. These systems leverage the error resilience of applications to obtain high energy efficiency, at the expense of tolerable reductions in the output quality. Hardware datapath operators able to reconfigure their precision and power consumption at runtime are key components of such systems. However, most implementations of these operators require manual, architecture-specific modifications and tend to have large power overheads compared to standard designs, when working at maximum precision. One promising design-independent alternative is dynamic voltage and accuracy scaling, whose adoption, however, is hindered by incompatibilities with standard design flows. In this paper, we propose a new methodology for the design of energy-quality scalable operators; our solution leverages runtime tuning of transistors threshold voltages to obtain a fine-grain control of the speed and power consumption of standard-cells within an operator. Thanks to the additional flexibility provided by this fine-grain knob, our method overcomes the main limitations of previous solutions, at the cost of a small area overhead. We demonstrate our approach on a 28 nm FDSOI technology; by exploiting the strong effect of back-gate biasing on threshold voltage, we achieve a power consumption reduction of more than 40% compared to the state-of-the-art, for the same precision.
Daniele Jahier Pagliari, Yves Durand, David Coriat, Edith Beigné, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 SystemC-AMS Thermal Modeling for the Co-simulation of Functional and Extra-Functional Properties
abstract
Temperature is a critical property of smart systems, due to its impact on reliability and to its inter-dependence with power consumption. Unfortunately, the current design flows evaluate thermal evolution ex-post on offline power traces. This does not allow to consider temperature as a dimension in the design loop, and it misses all the complex inter-dependencies with design choices and power evolution. In this article, by adopting the functional language SystemC-AMS (Analog Mixed Signal), we propose a method to enable thermal/power/functional co-simulation. The system thermal model is built by using state-of-the-art circuit equivalent models, by exploiting the support for electrical linear networks intrinsic of SystemC-AMS. The experimental results will show that the choice of SystemC-AMS is a winning strategy for building a simultaneous simulation of multiple functional and extra-functional properties of a system. The generated code exposes an accuracy comparable to that of the reference thermal simulator HotSpot. Additionally, the initial overhead due to the general purpose nature of SystemC-AMS is compensated by the surprisingly high performance of transient simulation, with speedups as high as two orders of magnitude.
Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino
ACM Trans. Design Autom. Electr. Syst.3
2019 A Cross-level Verification Methodology for Digital IPs Augmented with Embedded Timing Monitors
abstract
Smart systems are characterized by the integration in a single device of multi-domain subsystems of different technological domains, namely, analog, digital, discrete and power devices, MEMS, and power sources. Such challenges, emerging from the heterogeneous nature of the whole system, combined with the traditional challenges of digital design, directly impact on performance and on propagation delay of digital components. This article proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems with the automatic insertion of delay sensors, which can detect and correct timing failures. The article then proposes a methodology to verify such added features at system level. The augmented model is abstracted to SystemC TLM, which is automatically injected with mutants (i.e., code mutations) to emulate delays and timing failures. The resulting TLM model is finally simulated to identify timing failures and to verify the correctness of the inserted delay monitors. Experimental results demonstrate the applicability of the proposed design and verification methodology, thanks to an efficient sensor-aware abstraction methodology, by applying the flow to three complex case studies.
Sara Vinco, Nicola Bombieri, Daniele Jahier Pagliari, Franco Fummi, Enrico Macii, Massimo Poncino
ACM Trans. Design Autom. Electr. Syst.5
2018 Impact of graph partitioning on SNN placement for a multi-core neuromorphic architecture: work-in-progress
Francesco Barchi, Gianvito Urgese, Enrico Macii, Andrea Acquaviva
CASES3
2018 Multiple alignment of packet sequences for efficient communication in a many-core neuromorphic system: work-in-progress
Gianvito Urgese, Luca Peres, Francesco Barchi, Enrico Macii, Andrea Acquaviva
CASES4
2018 All-digital embedded meters for on-line power estimation
abstract
Modern low power designs use multiple knobs for concurrent dynamic and leakage power optimization; supply voltage and threshold voltage are the most adopted. An efficient control of these knobs needs management policies aware of the power breakdown. This implies the availability of smart on-chip strategies for dynamic and leakage power estimation at runtime. In this paper, we address this issue proposing the implementation of embedded dynamic/static power meters that use an optimized regression model fed with data collected from in-situ activity monitors. The number of sensors, their bitwidth and optimal placement are obtained through an automated design flow. The methodology works for general logic and applies not just to processor cores, but also to application-specific designs. We apply our solution to a representative class of benchmarks, showing that it can achieve an average estimation error smaller than 3%, with limited area and power overheads.
Daniele Jahier Pagliari, Valentino Peluso, Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE5
2018 GIS-based optimal photovoltaic panel floorplanning for residential installations
abstract
Shading is a crucial issue for the placement of PV installations, as it heavily impacts power production and the corresponding return of investment. Nonetheless, residential rooftop installations still rely on rule-of-thumb criteria and on gross estimates of the shading patterns, while more optimized approaches focus solely on the identification of suitable surfaces (e.g., roofs) in a larger geographic area (e.g., city or district). This work addresses the challenge of identifying an optimal (with respect to the overall energy production) placement of PV panels on a roof. The novel aspect of the proposed solution lies in the possibility of having a sparse, irregular placement of individual modules so as to better exploit the variance of solar data. The latter are represented in terms of the distribution of irradiance and temperature values over the roof, as elaborated from historical traces and Geographical Information System (GIS) data. Experimental results will prove the effectiveness of the algorithm through three real world case studies, and that the generated optimal solutions allow to increase power production by up to 28% with respect to rule-of-thumb solutions.
Sara Vinco, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Enrico Macii, Massimo Poncino
DATE5
2018 Battery-aware Design Exploration of Scheduling Policies for Multi-sensor Devices
abstract
Lifetime maximization is a key challenge in battery-powered multi-sensor devices. Battery-aware power management strategies combine task scheduling with dynamic voltage scaling (DVS), accounting for the fact that the power drawn by the device is different from that provided by the battery due to its many non-idealities. However, state-of-the-art techniques in this field do not take into account several important aspects, such as the impact of sensing tasks on the overall power demand, the (operating point dependent) losses due to multiple DC-DC conversions, and the dynamic modifications in battery efficiency caused by different distributions of the currents in the temporal and in the frequency domains. In this work, we propose a novel approach to identify optimal power management solutions, that addresses all these limitations. Specifically, using advanced battery and DC-DC converter models, we propose methods to explore the scheduling space both statically (at design time) and dynamically (at runtime), accounting not only for computation tasks, but also for communication and sensing. With this method, we show that the battery lifetime can be increased by as much as 23.36% if an optimal power management strategy is adopted.
Yukai Chen, Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2018 Optimal Topology-Aware PV Panel Floorplanning with Hybrid Orientation
abstract
Despite of being one of the most widespread green energy sources, the efficiency of PV rooftop installations is still repressed by shading and by the absence of a rigorous irradiance-aware placement approach. The goal of this work is to reach optimal energy production via an irregular placement of PV modules, by considering two degrees of freedom: orientation of each PV module and topology. Experimental results will prove the effectiveness of the proposed solution onto two real world case studies, with an increase of power production of up to 40%.
Sara Vinco, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2018 Battery-Aware Energy Model of Drone Delivery Tasks
abstract
Drones are becoming increasingly popular in the commercial market for various package delivery services. In this scenario, the mostly adopted drones are quad-rotors (i.e., quadcopters). The energy consumed by a drone may become an issue, since it may affect (i) the delivery deadline (quality of service), (ii) the number of packages that can be delivered (throughput) and (iii) the battery lifetime (number of recharging cycles). It is thus fundamental try to find the proper compromise between the energy used to complete the delivery and the speed at which the quadcopter flies to reach the destination. In order to achieve this, we have to consider that the energy required by the drone for completing a given delivery task does not exactly correspond to the energy requested to the battery, since the latter is a non-ideal power supply that is able to deliver power with different efficiencies depending on its state of charge. In this paper, we demonstrate that the proposed battery-aware delivery scheduling algorithm carries more packages than the traditional delivery model with the same battery capacity. Moreover, the battery-aware delivery model is 17% more accurate than the traditional delivery model for the same delivery scheme, which prevents the unexpected drone landing.
Donkyu Baek, Yukai Chen, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED5
2018 Dynamic Bit-width Reconfiguration for Energy-Efficient Deep Learning Hardware
abstract
Deep learning models have reached state of the art performance in many machine learning tasks. Benefits in terms of energy, bandwidth, latency, etc., can be obtained by evaluating these models directly within Internet of Things end nodes, rather than in the cloud. This calls for implementations of deep learning tasks that can run in resource limited environments with low energy footprints. Research and industry have recently investigated these aspects, coming up with specialized hardware accelerators for low power deep learning. One effective technique adopted in these devices consists in reducing the bit-width of calculations, exploiting the error resilience of deep learning. However, bit-widths are tipically set statically for a given model, regardless of input data. Unless models are retrained, this solution invariably sacrifices accuracy for energy efficiency.
Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino
ISLPED2
2018 Directed Graph Placement for SNN Simulation into a multi-core GALS Architecture
abstract
In this paper, we present a methodology for effi-ciently mapping neural networks over a neuromorphic computing architecture. The target architecture is a globally asynchronous locally synchronous (GALS) multi-core designed for simulating spiking neural networks (SNN) in real-time, that is spike timings should be the same as in the human brain. The SNN is implemented as a set of concurrent tasks modelling the behaviour of biological neurons, which are executed on the processing cores and communicate through spikes travelling on a network-on-chip. The problem of neuron-to-core mapping is relevant as a non-efficient allocation may impact real-time and reliability of the neural network execution. We designed a task placement pipeline capable of analysing the network of neurons and producing a placement configuration that enables a reduction of communication between computational nodes. The neuron-to-core mapping problem has been formalised as a problem of minimisation of synaptic elongation. Intuitively, this metric represents the cumulative distance that spikes generated by neurons running on a specific core have to travel to reach their destination core. The proposed placement methodology allows using different techniques to solve the problem. In this work Spectral Analysis, Multilevel Static Mapping, and Simulated Annealing were compared evaluating the overall post-placement synaptic elongation. Results point out that mapping solutions taking into account the directionality of the SNN provide a better placement and quantify this impact. Between all techniques considered only the Simulated Annealing was able to overcome an improvement of 25% compared to a random placement.
Francesco Barchi, Gianvito Urgese, Andrea Acquaviva, Enrico Macii
VLSI-SoC4
2018 A Parallel Hardware Architecture For Quantum Annealing Algorithm Acceleration
abstract
Quantum Annealing (QA) is an emerging technique, derived from Simulated Annealing, providing metaheuristics for multivariable optimisation problems. Studies have shown that it can be applied to solve NP-hard problems with faster convergence and better quality of result than other traditional heuristics, with potential applications in a variety of fields, from transport logistics to circuit synthesis and optimisation. In this paper, we present a hardware architecture implementing a QA-based solver for the Multidimensional Knapsack Problem, designed to improve the performance of the algorithm by exploiting parallelised computation. We synthesised the architecture using as a target an Altera FPGA board and simulated the execution for solving a set of benchmarks available in the literature. Simulation results show that the proposed implementation is about 100 times faster than a single-thread general-purpose CPU without impact on the accuracy of the solution.
Evelina Forno, Andrea Acquaviva, Yuki Kobayashi, Enrico Macii, Gianvito Urgese
VLSI-SoC4
2018 LAPSE: Low-Overhead Adaptive Power Saving and Contrast Enhancement for OLEDs
abstract
Organic Light Emitting Diode (OLED) display panels are becoming increasingly popular especially in mobile devices; one of the key characteristics of these panels is that their power consumption strongly depends on the displayed image. In this paper we propose LAPSE, a new methodology to concurrently reduce the energy consumed by an OLED display and enhance the contrast of the displayed image, that relies on image-specific pixel-by-pixel transformations. Unlike previous approaches, LAPSE focuses specifically on reducing the overheads required to implement the transformation at runtime. To this end, we propose a transformation that can be executed in real time, either in software, with low time overhead, or in a hardware accelerator with a small area and low energy budget. Despite the significant reduction in complexity, we obtain comparable results to those achieved with more complex approaches in terms of power saving and image quality. Moreover, our method allows to easily explore the full quality-versus-power tradeoff by acting on a few basic parameters; thus, it enables the runtime selection among multiple display quality settings, according to the status of the system.
Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino
IEEE Trans. Image Process.2
2018 Thermal Management of Batteries Using Supercapacitor Hybrid Architecture With Idle Period Insertion Strategy
Donghwa Shin, Massimo Poncino, Enrico Macii
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Building Energy Modelling and Monitoring by Integration of IoT Devices and Building Information Models
abstract
In recent years, the research about energy waste and CO2 emission reduction has gained a strong momentum, also pushed by European and national funding initiatives. The main purpose of this large effort is to reduce the effects of greenhouse emission, climate change to head for a sustainable society. In this scenario, Information and Communication Technologies (ICT) play a key role. From one side, advances in physical and environmental information sensing, communication and processing, enabled the monitoring of energy behaviour of buildings in real-time. The access to this information has been made easy and ubiquitous thank to Internet-of-Things (IoT) devices and protocols. From the other side, the creation of digital repositories of buildings and districts (i.e. Building Information Models - BIM) enabled the development of complex and rich energy models that can be used for simulation and prediction purposes. As such, an opportunity is emerging of mixing these two information categories to either create better models and to detect unwanted or inefficient energy behaviours. In this paper, we present a software architecture for management and simulation of energy behaviours in buildings that integrates heterogeneous data such as BIM, IoT, GIS (Geographical Information System) and meteorological services. This integration allows: i) (near-) real-time visualisation of energy consumption information in the building context and ii) building performance evaluation through energy modelling and simulation exploiting data from the field and real weather conditions. Finally, we discuss the experimental results obtained in a real-world case-study.
Lorenzo Bottaccioli, Alessandro Aliberti, Francesca Maria Ugliotti, Edoardo Patti, Anna Osello, Enrico Macii, Andrea Acquaviva
COMPSAC (1)6
2017 A circuit-equivalent battery model accounting for the dependency on load frequency
abstract
Circuit-equivalent battery models are considered defacto standard for modeling and simulation of digital systems due to many practical advantages. In spite of the many variants of models proposed in the literature, none of them accounts for one important feature of the battery dynamics, namely, the dependency on the frequency of current load profile. For a given average current value, current loads with different spectral distributions may have quite different impacts on the battery discharge. This is a very well-know issue in the design of hybrid energy storage systems, where different types of storages devices are used, each with different storage efficiency for different load frequency ranges. We propose a basic modification to a state-of-the-art model that incorporates this load frequency dependency, as well as a methodology to identify the frequency-sensitive parameters of the model from publicly available data (e.g., datasheets). The results show that frequency-agnostic models can significantly overestimate the battery state-of-charge, and that this effect is far from being negligible.
Yukai Chen, Enrico Macii, Massimo Poncino
DATE2
2017 A methodology for the design of dynamic accuracy operators by runtime back bias
abstract
Mobile and IoT applications must balance increasing processing demands with limited power and cost budgets. Approximate computing achieves this goal leveraging the error tolerance features common in many emerging applications to reduce power consumption. In particular, adequate (i.e., energy/quality-configurable) hardware operators are key components in an error tolerant system. Existing implementations of these operators require significant architectural modifications, hence they are often design-specific and tend to have large overheads compared to accurate units. In this paper, we propose a methodology to design adequate data-path operators in an automatic way, which uses threshold voltage scaling as a knob to dynamically control the power/accuracy tradeoff. The method overcomes the limitations of previous solutions based on supply voltage scaling, in that it introduces lower overheads and it allows fine-grain regulation of this tradeoff. We demonstrate our approach on a state-of-the-art 28nm FDSOI technology, exploiting the strong effect of back biasing on threshold voltage. Results show a power consumption reduction of as much as 39% compared to solutions based only on supply voltage scaling, at iso-accuracy.
Daniele Jahier Pagliari, Yves Durand, David Coriat, Anca Mariana Molnos, Edith Beigné, Enrico Macii, Massimo Poncino
DATE6
2017 Workload-driven frequency-aware battery sizing
abstract
Despite the wide body of literature on the sizing of energy storage devices available in the domain of electrical energy systems, the problem has not drawn much attention in the area of battery-powered electronic systems. It is well-known that the straightforward method of sizing battery as the product of an expected duration and the average load current always underestimates the actual capacity that the battery can supply. The variability of the workload and of its spectral distribution will in fact affect the effective capacity of battery that cannot be ignored. This paper proposed a methodology to compute the required capacity of a battery based on the properties of the workload; in particular it accounts for both the impact of the distribution of the current load and of its frequencies, and determines corrective factors for both effects to be used for the calculation of the actual capacity. We used a frequency-sensitive circuit-equivalent battery model to validates our method on three synthetic and two real workloads. Simulation results show that even for workload with same average current, the required capacity can be as much as 70% larger than the capacity estimated using a traditional method.
Yukai Chen, Enrico Macii, Massimo Poncino
ISLPED2
2017 A Layered Methodology for the Simulation of Extra-Functional Properties in Smart Systems
abstract
Smart systems represent a broad class of intelligent, miniaturized devices incorporating functionality like sensing, actuation, and control. In order to support these functions, they must include sophisticated and heterogeneous components, such as sensors and actuators, multiple power sources and storage devices, digital signal processing, and wireless connectivity. The high degree of heterogeneity typical of smart systems has a heavy impact on their design: the challenges are not in fact restricted to their functionality, but are also related to a number of extra-functional properties, including power consumption, temperature, and aging. Current simulation- or model-based design approaches do not target a smart system as a whole, but rather single domains (digital, analog, power devices, etc.) or properties. This paper tries to overcome this limitation by proposing a framework for the concurrent simulation of both functionality and such extra-functional properties. The latter are modeled as different information flows, managed by dedicated “virtual buses” and formalized through the adoption of IP-XACT. SystemC, through the support of physical and continuous time modeling provided by its analog and mixed signal extension, is used to implement both functional and extra-functional models. Experimental results show the efficiency, accuracy and modularity of the proposed approach on an example case study, in which substantial speedups with respect to standard model-based design tools go along with a very high degree of accuracy (-5%). Furthermore, the case study highlights that the proposed framework allows to easily capture at run time the mutual impact of properties, e.g., in case of power and temperature.
Sara Vinco, Yukai Chen, Franco Fummi, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2017 A Flexible Distributed Infrastructure for Real-Time Cosimulations in Smart Grids
abstract
Due to the increasing penetration of distributed generation, storage, electric vehicles, and new information communication technologies, distribution networks are evolving toward the smart grid paradigm. For this reason, new control strategies, algorithms, and technologies need to be tested and validated before their actual field implementation. In this paper, we present a novel modular distributed infrastructure, based on real-time simulation, for multipurpose smart grid studies. The different components of the infrastructure are described, and the system is applied to a case study based on a real urban district located in northern Italy. The presented infrastructure is shown to be flexible and useful for different and multidisciplinary smart grid studies.
Lorenzo Bottaccioli, Abouzar Estebsari, Enrico Pons, Ettore Bompard, Enrico Macii, Edoardo Patti, Andrea Acquaviva
IEEE Trans. Ind. Informatics5
2017 Approximate Energy-Efficient Encoding for Serial Interfaces
abstract
Serial buses are ubiquitous interconnections in embedded computing systems that are used to interface processing elements with peripherals, such as sensors, actuators, and I/O controllers. Despite their limited wiring, as off-chip connections they can account for a significant amount of the total power consumption of a system-on-chip device. Encoding the information sent on these buses is the most intuitive and affordable way to reduce their power contribution; moreover, the encoding can be made even more effective by exploiting the fact that many embedded applications can tolerate intermediate approximations without a significant impact on the final quality of results, thus trading off accuracy for power consumption. We propose a simple yet very effective approximate encoding for reducing dynamic energy in serial buses. Our approach uses differential encoding as a baseline scheme and extends it with bounded approximations to overcome the intrinsic limitations of differential encoding for data with low temporal correlation. We show that the proposed scheme, in addition to yielding extremely compact codecs, is superior to all state-of-the-art approximate serial encodings over a wide set of traces representing data received or sent from/to sensor or actuators.
Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino
ACM Trans. Design Autom. Electr. Syst.2
2016 Serial T0: approximate bus encoding for energy-efficient transmission of sensor signals
abstract
Off-chip serial buses are common in embedded systems, and due to the long physical lines, can contribute significantly to their energy consumption. However, these buses are often connected to analog sensors, whose data is inherently affected by noise and A/D errors. Thus, communication can tolerate small approximations, without a significant impact on the system outputs quality.
Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino
DAC2
2016 Panel: Looking backwards and forwards
Marco Casale-Rossi, Giovanni De Micheli, Antun Domic, Enrico Macii, Domenico Rossi, Joseph Sawicki
DATE4
2016 Low-overhead adaptive constrast enhancement and power reduction for OLEDs
Daniele Jahier Pagliari, Massimo Poncino, Enrico Macii
DATE3
2016 CONTREX: Design of Embedded Mixed-Criticality CONTRol Systems under Consideration of EXtra-Functional Properties
abstract
The increasing processing power of today's HW/SW platforms leads to the integration of more and more functions in a single device. Additional design challenges arise when these functions share computing resources and belong to different criticality levels. The paper presents the CONTREX European project and its preliminary results. CONTREX complements current activities in the area of predictable computing platforms and segregation mechanisms with techniques to consider the extra-functional properties, i.e., timing constraints, power, and temperature. CONTREX enables energy efficient and cost aware design through analysis and optimization of these properties with regard to application demands at different criticality levels.
Ralph Görgen, Kim Grüttner, Fernando Herrera, Pablo Peñil, Julio L. Medina, Eugenio Villar, Gianluca Palermo, William Fornaciari, Carlo Brandolese, Davide Gadioli, Sara Bocchio, Luca Ceva, Paolo Azzoni, Massimo Poncino, Sara Vinco, Enrico Macii, Salvatore Cusenza, John M. Favaro, Raúl Valencia, Ingo Sander, Kathrin Rosvall, Davide Quaglia
DSD16
2016 IP-XACT for smart systems design: extensions for the integration of functional and extra-functional models
abstract
Smart systems are miniaturized devices integrating computation, communication, sensing and actuation. As such, their design can not focus solely on functional behavior, but it must rather take into account different extra-functional concerns, such as power consumption or reliability. Any smart system can thus be modeled through a number of views, each focusing on a specific concern. Such views may exchange information, and they must thus be simulated simultaneously to reproduce mutual influence of the corresponding concerns. This paper shows how the IP-XACT standard, with some necessary extensions, can effectively support this simultaneous simulation. The extended IP-XACT descriptions allow to model extra-functional properties with a homogeneous format, defined by analysing requirements and characteristic of three main concerns, i.e., power, temperature and reliability. The IP-XACT descriptions are then used to automatically generate a skeleton of the simulation infrastructure in SystemC. The skeleton can be easily populated with models available in the literature, thus reaching simultaneous simulation of multiple concerns.
Sara Vinco, Michele Lora, Enrico Macii, Massimo Poncino
FDL3
2016 Fast Thermal Simulation using SystemC-AMS
abstract
Out of the many options available for thermal simulation of digital electronic systems, those based on solving an RC equivalent circuit of the thermal network are the most popular choice in the EDA community, as they provide a reasonable tradeoff between accuracy and complexity. HotSpot, in particular, has become the de-facto standard in these communities, although other simulators are also popular. These tools have many benefits, but they are relatively inefficient when performing thermal analysis for long simulation times, due to the occurrence of a large number of redundant computations intrinsic in the underlying models.
Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2016 Approximate Differential Encoding for Energy-Efficient Serial Communication
abstract
Embedded computing systems include several off-chip serial links, that are typically used to interface processing elements with peripherals, such as sensors, actuators and I/O controllers. Because of the long physical lines of these connections, they can contribute significantly to the total energy consumption. On the other hand, many embedded applications are error resilient, i.e. they can tolerate intermediate approximations without a significant impact on the final quality of results. This feature can be exploited in serial buses to explore the trade-off between data approximations and energy consumption.
Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2016 Graphene-PLA (GPLA): a Compact and Ultra-Low Power Logic Array Architecture
abstract
The key characteristics of the next generation of ICs for wearable applications include high integration density, small area, low power consumption, high energy-efficiency, reliability and enhanced mechanical properties like stretchability and transparency. The proper mix of new materials and novel integration strategies is the enabling factor to achieve those design specifications.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2016 A Unified Model of Power Sources for the Simulation of Electrical Energy Systems
abstract
Models of power sources are essential elements in the simulation of systems that generate, store and manage energy. In spite of the huge difference in power scale, they perform a common function: converting a primary environmental quantity into power. This paper proposes a unified model of a power source that is applicable to any power scale, and that can be derived solely from data contained in the specification or the datasheet of a device. The key feature of our model is the normalization of the energy generation characteristic of the power source by means of a reduction to a function expressing extracted power vs. the "scavenged" quantity. The proposed model proved to apply to two kinds of power sources, i.e., a wind turbine and a photovoltaic panel, and to provide a good level of accuracy and simulation performance w.r.t. widely adopted models.
Sara Vinco, Yukai Chen, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2016 Enabling quasi-adiabatic logic arrays for silicon and beyond-silicon technologies
abstract
Adiabatic logic aims at mimicking an adiabatic (i.e., without energy exchange) charging process in digital circuits. Although regarded as a mostly theoretical computation style, research on the topic has been constantly active over the years, providing several demonstrations of working implementations [1]. The interest in adiabatic circuits recently increased with the introduction of emerging devices, e.g., Nanoelectromechanicals switches (NEMs) [2] and graphene p-n junctions [3], which have been proven to be good technological vehicles for adiabatic computing. Despite their energy efficiency, adiabatic logic faced severe limitations in reaching large scale integration due to the difficulty in logic pipelining and the lack of CAD tools able to cope with today's design complexity.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
ISCAS3
2016 A Li-Ion Battery Charge Protocol with Optimal Aging-Quality of Service Trade-off
abstract
The reduction of usable capacity of rechargeable batteries can be mitigated during the charge process by acting on some stress factors, namely, the average state-of-charge (SOC) and the charge current. Larger values of these quantities cause an increased degradation of battery capacity, so it would be desirable to keep both as low as possible, which is obviously in contrast with the objective of a fast charge. However, by exploiting the fact that in most battery-powered systems the time during which it is plugged for charging largely exceeds the time required to charge, it is possible to devise appropriate charge protocols that achieve a good balance between fast charge and aging.
Yukai Chen, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED4
2016 Frequency domain characterization of batteries for the design of energy storage subsystems
abstract
The Ragone chart is a pictorial representation to express the well-known the trade-off between available energy vs. power of different classes of energy storage devices (ESDs) like batteries or supercapacitors. Ragone charts, however, do not normally provide information about individual devices, which is an essential requirement for the actual design of the energy storage sub-system.
Yukai Chen, Enrico Macii, Massimo Poncino
VLSI-SoC2
2016 Ultra-Fine Grain Vdd-Hopping for energy-efficient Multi-Processor SoCs
abstract
This paper introduces Ultra-Fine Grain Vdd-Hopping (FINE-VH), an extension of Dynamic Voltage-Frequency Scaling (DVFS) for energy efficient Multi-Processor SoCs (MPSoCs). The proposed technique leverages the working principle of Vdd-Hopping applied at ultra-fine granularity, i.e., within the core, by means of a layout-assisted, level-shifter free, dynamic dual-Vdd control strategy where leakage currents are minimized through an optimal timing-driven poly-bias assignment procedure.
Valentino Peluso, Andrea Calimera, Enrico Macii, Massimo Alioto
VLSI-SoC3
2016 Multi-function logic synthesis of silicon and beyond-silicon ultra-low power pass-gates circuits
abstract
Pass-gates logic is known to be intrinsically more energy efficient than static CMOS. This feature attracted the research interest over the years and many working implementations have been demonstrated. Recent works, in particular, have shown that pass-gates logic is well suited for ultra-low power adiabatic circuits mapped on emerging technologies. Despite the progress made, several design issues still prevent pass-gates logic circuits reaching large scale integration. In this work we deal with the lack of synthesis tools and methodologies. We propose a multi-function decomposition engine that yields (i) an efficient abstract circuit modeling through a more compact data-structure, the Multi-Function Pass Diagram (MFPD) and (ii) an effective multi-gate area/delay-driven low-power synthesis&optimization flow. Simulation results conducted on different technologies, i.e., silicon and graphene, demonstrate that logic circuits synthesized with the proposed tool are smaller in size and depth, hence less power consuming and faster than circuits obtained through conventional synthesis flows based on Binary Decision Diagrams.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
VLSI-SoC3
2015 One-pass logic synthesis for graphene-based Pass-XNOR logic circuits
abstract
Electrostatically controlled graphene P-N junctions are devices built on a single layer graphene sheet that can be turned-ON/OFF via external potential difference. Their electrical behavior resembles a CMOS transmission gate with an embedded XNOR Boolean functionality. Recent works presented an efficient design style, the Pass-XNOR logic (PXL), which allows the implementation of adiabatic logic circuits with ultra low-power features.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
DAC3
2015 A new distributed framework for integration of district energy data from heterogeneous devices
Francesco Gavino Brundu, Edoardo Patti, Andrea Acquaviva, Michelangelo Grosso, Gaetano Rasconà, Salvatore Rinaudo, Enrico Macii
DATE7
2015 A tool-chain to foster a new business model for photovoltaic systems integration exploiting an Energy Community approach
abstract
New approaches and business models for the development of renewable sources are needed as an alternative to feed-in tariffs. In this work, we present a tool-chain based on a distributed infrastructure for planning renewable energy systems deployment. This solution aims at fostering new services and business models by promoting energy community actions. Such tool-chain is able to: (i) evaluate the photovoltaic potential of the rooftops of a community; (ii) perform economic assessments of distributed photovoltaic system plants considering a “community based business model”. As case study, we considered a foothill community in north-west of Italy in which the tool-chain performed economic and energetic analyses. In order to integrate the proposed business model in the Italian regulatory framework, we analysed the Italian laws for electricity distribution and operation, highlighting the limitations in integrating such community approach.
Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Enrico Macii, Matteo Jarre, Michel Noussan
ETFA4
2015 Characterizing the Activity Factor in NBTI Aging Models for Embedded Cores
abstract
In deeply scaled CMOS technologies, device aging causes cores performance parameters to degrade over time. While accurate models to efficiently assess these degradation exist for devices and circuits, no reliable model for processor cores has gained strong acceptance in the literature. In this work, we propose a methodology for deriving an NBTI aging model for embedded cores. Based on an accurate characterization on the netlist of the core, we were able to (1) prove the independence of the aging on the workload (i.e., executed instructions), and (2) calculate an equivalent average constant aging factor that justifies the use of the baseline model template. We derived and assessed the proposed model by using a RISC-like processor core implemented in a 45nm process technology as a reference architecture, achieving a maximum error of 2.2% against simulated data on the core netlist.
Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2015 Exploiting the Expressive Power of Graphene Reconfigurable Gates via Post-Synthesis Optimization
abstract
As an answer to the new electronics market demands, semiconductor industry is looking for different materials, new process technologies and alternative design solutions that can support Silicon replacement in the VLSI domain. The recent introduction of graphene, together with the option of electrostatically controlling its doping profile, has shown a possible way to implement fast and power efficient Reconfigurable Gates (RGs). Also, and this is the most important feature considered in this work, those graphene RGs show higher expressive power, i.e., they implement more complex functions, like Majority, MUX, XOR, with less area w.r.t. CMOS counterparts. Unfortunately, state-of-the-art synthesis tools, which have been customized for standard NAND/NOR CMOS gates, do not exploit the aforementioned feature of graphene RGs.
Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino, Luca G. Amarù, Giovanni De Micheli, Pierre-Emmanuel Gaillardon
ACM Great Lakes Symposium on VLSI4
2015 Design and Characterization of Analog-to-Digital Converters using Graphene P-N Junctions
abstract
Electrostatically controlled graphene p-n junctions are devices built on single-layer graphene sheets whose in-to-out resistance can be dynamically tuned through external voltage potentials.
Roberto Giorgio Rizzo, Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI4
2015 An aging-aware battery charge scheme for mobile devices exploiting plug-in time patterns
abstract
The aging of a rechargeable battery is mainly due to stress during charge-discharge cycles. Although the discharge phase is difficult to control, the charging phase can be performed in a specific way in order to mitigate the aging of the battery during its usage. It therefore becomes important to select the correct charging algorithm. In the case of mobile systems, equipped mainly with lithiumion batteries, the standard widely adopted for charging a battery is the typical constant current/constant voltage (CC-CV) protocol usually based on a linearly regular charge process. In this work, we propose a charging protocol based on the standard CC-CV method in which the charge start time and the value of the charging current can be programmed in such a way that the aging of the battery is mitigated. To validate this charging scheme we use an aging model that includes the charge/discharge current among the major parameters, and an analytical macro-model for the CC-CV charge time analysis.
Alberto Bocca, Alessandro Sassone, Alberto Macii, Enrico Macii, Massimo Poncino
ICCD4
2015 An automated design flow for approximate circuits based on reduced precision redundancy
abstract
Reduced Precision Redundancy (RPR) is a popular Approximate Computing technique, in which a circuit operated in Voltage Over-Scaling (VOS) is paired to a reduced-bitwidth and faster replica so that VOS-induced timing errors are partially recovered by the replica, and their impact is mitigated. Previous works have provided various examples of effective implementations of RPR, which however suffer from three limitations: first, these circuits are designed using ad-hoc procedures, and no generalization is provided; second, error impact analysis is carried out statistically, thus neglecting issues like non-elementary data distribution and temporal correlation. Last, only dynamic power was considered in the optimization. In this work we propose a new generalized approach to RPR that allows to overcome all these limitations, leveraging the capabilities of state-of-the-art synthesis and simulation tools. By sacrificing theoretical provability in favor of an empirical input-based analysis, we build a design tool able to automatically add RPR to a preexisting gate-level netlist. Thanks to this method, we are able to confute some of the conclusions drawn in previous works, in particular those related to statistical assumptions on inputs; we show that a given inputs distribution may yield extremely different results depending on their temporal behavior.
Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino
ICCD3
2015 An equation-based battery cycle life model for various battery chemistries
abstract
The evaluation of the cycle life of batteries is an essential task in the assessment of the reliability and cost of battery-operated devices. Several compact cycle life models have been proposed in the literature, that exhibit a general trade-off between generality and accuracy. Some models are based on a compact equation derived from experimental data and try to extract a general relationship between cycle life and the relevant parameters (mostly the depth of discharge), but suffer from poor accuracy. At the other extreme, more accurate models, based on incorporating the aging effect into an equivalent circuit, tend to be focused on a specific device and are seldom applicable to another battery. In this work we propose an equation-based model that tries to overcome the accuracy limits of previous similar models. The model parameters are obtained by fitting the curve based on information reported in datasheets, and can be adapted (with different accuracy levels) to the amount of available information. We applied the model to various commercial batteries for which full information on their cycle life is available. Results show an average estimation error, in terms of the number of cycles, generally smaller than 10%, which is consistent with the typical tolerance provided in the datasheets, and much lower than previous equation-based models.
Alberto Bocca, Alessandro Sassone, Donghwa Shin, Alberto Macii, Enrico Macii, Massimo Poncino
VLSI-SoC5
2015 A Statistical Model-Based Cell-to-Cell Variability Management of Li-ion Battery Pack
abstract
The cell-to-cell variability of batteries is a well-known problem particularly when it comes to the assembly of large battery packs. Different battery cells exhibit substantial variability due to manufacturing tolerances, which should be assessed and managed carefully. Such variability has been approached mostly from the point of view of the chemical and physical phenomena, but these solutions are normally too complicated for the system-level design of electric applications. This paper proposes a combined cell-to-cell variability model of the capacity and internal resistance of a Li-ion battery that accounts for the variability effects in the cell manufacturing process. The proposed model allows to verify some known properties, such as the correlation between the capacity and internal resistance, to be verified qualitatively and the amount of variability and its impact on the design of battery packs to be assessed quantitatively. Using this model, the issue of how to consider the variability when constructing battery packs was also addressed. Modern battery packs normally incorporate some cell balancing circuitry, which is meant to balance cell voltages during charging at the expense of a bypassed (unstored) charge. For discharge, the cell-to-cell variability hides a part of the usable capacity of the battery pack. This paper proposes the use of variability information to assemble battery packs with minimal intracolumn variance of capacity. A weight-based variance minimization method, based on the correlation between cell capacity and weight is proposed to avoid resorting to direct battery capacity measurements, which is time-consuming and requires costly measurement equipment. The simulation result shows that the proposed weight-based approach allows an acceptable management of the cell-to-cell variability without the discharging experiment.
Donghwa Shin, Massimo Poncino, Enrico Macii, Naehyuck Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Statistical Battery Models and Variation-Aware Battery Management
abstract
Cell-to-cell variability of batteries is a well-known problem especially when it comes to assembling large battery packs. Different battery cells exhibit substantial variability among them due to manufacturing tolerances, which should be carefully assessed and managed. Although battery packs usually incorporate some cell balancing circuitry, it is supposed to balance cell voltages dynamically at the expense of bypassed (not stored) charge.
Donghwa Shin, Enrico Macii, Massimo Poncino
DAC2
2014 A cross-level verification methodology for digital IPs augmented with embedded timing monitors
abstract
Smart systems implement the leading technology advances in the context of embedded devices. Current design methodologies are not suitable to deal with tightly interacting subsystems of different technological domains, namely analog, digital, discrete and power devices, MEMS and power sources. The effects of interaction between components and with the environment must be modeled and simulated at system level to achieve high performance. Focusing on the digital domain, additional design constraints have to be considered as a result of the integration of multi-domain subsystems in a single device. The main digital design challenges, combined with those emerging from the heterogeneous nature of the whole system, directly impact on performance and on propagation delay of the digital component. This paper proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems, and a methodology to verify the added features at system-level. The design approach consists of augmenting the RTL model through the automatic insertion of delay sensors, which can detect and correct timing failures. The augmented model is abstracted to SystemC TLM and, then, mutants (i.e., code mutations for emulating timing failures) are automatically injected into the model. Experimental results demonstrate the applicability of the proposed design and verification methodology and the effectiveness of the simulation performance.
Valerio Guarnieri, Massimo Petricca, Alessandro Sassone, Sara Vinco, Nicola Bombieri, Franco Fummi, Enrico Macii, Massimo Poncino
DATE7
2014 Cache aging reduction with improved performance using dynamically re-sizable cache
abstract
Aging of transistors is a limiting factor for long term reliability of devices in sub-100nm technologies. It's a worst-case metric where the lifetime of a device is determined by the earliest failing component. Impact is more serious on memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. Previous works have shown that partitioning based strategies based on power management techniques can effectively control aging effects and can extend lifetime of the cache significantly. However, such a benefit comes as a tradeoff with performance which reduces proportionally as the time elapses. To address this problem and provide a single solution to concurrently improve aging, energy and performance of the cache, we propose an architectural solution based on the dynamically re-sizable cache and cache partitioning approaches. By this strategy, cache is dynamically re-sized and reconfigured whenever a cache block becomes unreliable. Coupling such aging mitigation technique along with dynamically re-sizable cache approach provides on average 30% lifetime improvement with less than 0.4x degradation in performance whereas, in previous solutions, performance degradation sometimes goes upto 10x.
Haroon Mahmood, Massimo Poncino, Enrico Macii
DATE3
2014 Thermal management of batteries using a hybrid supercapacitor architecture
abstract
Thermal analysis and management of batteries have been an important research issue for battery-operated systems such as electric vehicles and mobile devices. Nowadays, battery packs are designed considering heat dissipation, and external cooling devices such as a cooling fan are also widely used to enforce the reliability and extend the lifetime of a battery. This type of approaches that target the enhancement of the cooling efficiency via the reduction of the thermal resistance cannot achieve an immediate temperature drop to avoid a thermal emergency situation. Approaches based on removing the heat from the heat sources via idle period insertion (similar to what is done for silicon devices) would allow faster thermal response; however it is not obvious how to implement these schemes in the context of batteries. In this paper, we propose the use of a simple parallel battery-supercapacitor hybrid architecture with a dual-mode discharging strategy that can provide immediate temperature management, in which the supercapacitor is used as an energy buffer during the idle periods of the battery. Simulation results shows that the proposed method can keep the battery temperature within the safe range without external cooling devices while exploiting the advantage of the battery-supercapacitor parallel connection.
Donghwa Shin, Massimo Poncino, Enrico Macii
DATE3
2014 Pass-XNOR logic: A new logic style for P-N junction based graphene circuits
abstract
In this work we introduce a new logic style for p-n junctions based digital graphene circuits: the pass-XNOR logic style. The latter enables the realization of compact, energy efficient circuits that better exploit the characteristics of graphene. We first show how a single p-n junction can be conceived as a pass-XNOR gate, i.e., a transmission gate with embedded logic functionality, the XNOR Boolean operator. Secondly, we propose a smart integration strategy in which series/parallel connections of pass-XNOR gates allow to implement AND/OR logical conjunctions, and, therefore, all possible truth tables. Experimental results conducted on a set of representative logic functions show the superior of pass-XNOR logic circuits w.r.t. standard CMOS circuits and graphene circuits that use p-n junctions in a complementary-like structure.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE3
2014 Ultra Low-Power Computation via Graphene-Based Adiabatic Logic Gates
abstract
As an answer to the difficulties in improving the figures of merit of deeply-scaled CMOS devices, researchers have looked at alternative materials and technologies for implementing new devices that can overcome the limitations of CMOS. Graphene has emerged as one of the most promising candidates among these new materials, recent works have demonstrated the implementation of electrostatically-controlled p-n junctions that can serve as the basic primitive for a new class of compact, fast and energy-efficient graphene-based logic gates. In this work we revisit those gates from a different perspective, namely, as devices that can operate adiabatically, that is, that are able to reuse the dissipated energy. We show how to build the basic logic gates by appropriately interconnecting graphene based p-n junctions and characterize those adiabatic gates for power and performance. The comparison between these adiabatic gates and both adiabatic CMOS and their non-adiabatic graphene-based counterpart shows that the former can operate with 2X to 3X less average power, and about 4X better power-delay product.
Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino
DSD3
2014 Towards a Software Infrastructure for District Energy Management
abstract
Nowadays ICT is becoming a key factor to enhance the energy optimization in our cities. At district level, real-time information can be accessed to monitor and control the energy distribution network. Moreover, the fine grain monitoring and control done at building level can provide additional information to develop more efficient control policies for energy distribution in the district. In this paper we present a distributed software infrastructure for district energy management, which aims to provide a digital archive of the city in which energetic information is available. Such information is considered as the input for a decision system, which aims to increase the energy efficiency by promoting local balancing and shaving peak loads. As case study, we integrated in our proposed cloud the heating distribution network in Turin and we present exploitable options based on real-world environmental data to increase the energy efficiency and minimize the peak request.
Edoardo Patti, Andrea Acquaviva, Adriano Sciacovelli, Vittorio Verda, Dario Martellacci, Federico Boni Castagnetti, Enrico Macii
EUC7
2014 Modeling of the charging behavior of li-ion batteries based on manufacturer's data
abstract
The market of portable devices, wireless sensors, electric vehicles and storage systems has grown enormously in recent years. As a consequence, batteries and related technologies have become one of the major topics for researchers. Due to the large variety of applications in which batteries are involved, battery modeling is becoming an extremely important research topic. This relevance is witnessed by the number of papers addressing battery modeling.
Alessandro Sassone, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI5
2014 Automated generation of battery aging models from datasheets
abstract
The de-facto standard approach in battery modeling consists of the definition of a generic model template in terms of an equivalent electric circuit, which is then populated either using data obtained from direct measurements on actual devices or by some extrapolation of battery characteristics available from datasheets. These models typically describe only intra-cycle effects, that is, those manifesting within a single charge/discharge cycle of a battery. However, basic battery dynamics, during a single discharge, cannot provide a true estimate of the actual lifetime of the battery, e.g., how its usability decreases due to long-term and irreversible effects, such as the fading of capacity due to aging or to repeated cycling. While some solutions in the literature provide answers to this problem by proposing suitable models for these effects, they do not provide solutions for how to incorporate them into a generic model template. In this work we propose a method to include inter-cycle battery effects into a reference model template in an automated way, and using solely data reported by battery manufacturers. Flexibility and accuracy of the proposed strategy are demonstrated by modeling a commercial lithium iron phosphate battery, whose datasheet provides long-term capacity fading information.
Massimo Petricca, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino
ICCD5
2014 A compact macromodel for the charge phase of a battery with typical charging protocol
abstract
Availability of a simulation model of a battery is one of the most important requisites in the system-level design of battery-powered systems. The vast majority of the models describe the discharge behavior of the battery; so far, the estimation of charging time has been in fact only marginally studied because the charging phase is regarded as a relatively controlled process compared to discharge. In this paper, we present a compact macro-model for the estimation of charging time under the most widely used charge protocol, i.e., Constant Current-Constant Voltage (CC-CV). This model is derived under the consideration of the context of the existing models including the well-known Peukert's law and equivalent electric circuits. The estimation result with the proposed model based on the manufacturer's data of commercial Li-ion batteries shows fair accuracy, especially when compared to estimates on parameters extracted from discharge characteristics.
Donghwa Shin, Alessandro Sassone, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED5
2014 An open-source framework for formal specification and simulation of electrical energy systems
abstract
Electrical energy systems (EESs) are systems which consume, generate, distribute and store energy at various scales. This paper presents a modeling and simulation framework that uses principles borrowed from the system-level simulation of digital systems and extends them to the case of EESs. The framework relies on open-source standards such as SystemC (and its Analog and Mixed-Signal extensions) for simulation, and IP-XACT for interface definition.
Sara Vinco, Alessandro Sassone, Franco Fummi, Enrico Macii, Massimo Poncino
ISLPED4
2014 Modeling of Physical Defects in PN Junction Based Graphene Devices
Sandeep Miryala, Matheus Oleiro, Letícia Maria Veiras Bolzani, Andrea Calimera, Enrico Macii, Massimo Poncino
J. Electron. Test.5
2014 Dynamic Indexing: Leakage-Aging Co-Optimization for Caches
abstract
Traditional implementations of low-power states based on voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused by negative bias temperature instability (NBTI), which can be explained in terms of the intuitive correlation between the idleness and the reduced workload of a system. Such a joint benefit has been exploited only partially because of the different nature of energy and aging as cost functions: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually result in limited aging reductions. In this paper, we address this problem in the context of power-managed caches, which represent a critical target for NBTI-reduced aging: given their symmetric structure, SRAM structures are, in particular, sensitive to NBTI effects because they cannot take advantage of the value-dependent recovery typical of NBTI. We propose a strategy called dynamic indexing, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the various power managed units (e.g., lines). This distribution allows fully using the leakage optimization potential and extending the lifetime of a cache. We explore various alternatives, in particular different granularities of the power managed units as well as different reindexing functions. Experimental analysis shows that it is possible to simultaneously reduce leakage power and aging in caches, with minimal power consumption overhead.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Energy/Lifetime Cooptimization by Cache Partitioning With Graceful Performance Degradation
abstract
Aging of transistors can adversely impact the long-term reliability of devices in subnanometric technologies. Without any countermeasure, the first component that becomes unreliable will determine the life span of an entire device. The effect is more susceptible in memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. In this paper, we propose a reliability management technique based on the idea of cache partitioning, which deals with cell failures by gracefully degrading its performance. By this partitioning-based strategy, various subblocks will become unreliable at different times, and the cache will keep functioning with reduced efficiency. A coarse-grain implementation of this approach, with the use of a smart aging-driven partitioning algorithm, provides a lifetime extension of more than 2× . On the other hand, a fine-grain strategy with a single cache line as a unit of power management, stretch the lifetime to its maximum limits with an addition of small hardware overhead.
Haroon Mahmood, Mirko Loghi, Massimo Poncino, Enrico Macii
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Energy-optimal SRAM supply voltage scheduling under lifetime and error constraints
abstract
This work addresses the energy efficiency of the memory architecture in safety-critical systems that have to guarantee a given level of service and a minimum lifetime. We specifically target SRAM structures in which decreased reliability manifests itself in terms of the aging induced by NBTI (Negative Bias Temperature Instability), and in which the level of service is represented by the bit-error rate (BER).
Andrea Calimera, Enrico Macii, Massimo Poncino
DAC2
2013 HW-SW integration for energy-efficient/variability-aware computing
abstract
Recent trends in embedded system architectures brought a rapid shift towards multicore, heterogeneous and reconfigurable platforms. This imposes a large effort for programmers to develop their applications to efficiently exploit the underlying architecture. In addition, process variability issues lead to performance and power uncertainties, impacting expected quality of service and energy efficiency of the running software. In particular, variability may lead to sub-optimal runtime task allocation.
Gasser Ayad, Andrea Acquaviva, Enrico Macii, Brahim Sahbi, Romain Lemaire
DATE3
2013 A verilog-a model for reconfigurable logic gates based on graphene pn-junctions
abstract
Single layer sheets of graphene show special electrical properties that can enable the next generation of smart ICs. Recent works have proven the availability of an electrostatically controlled pn-junction upon which it is possible to design multi-function reconfigurable logic devices that naturally behave as multiplexers. In this work we introduce a stable large-signal Verilog-A model that mimics the behavior of the aforementioned devices. The proposed model, validated through the SPICE characterization of a MUX-based standard cell library we designed as benchmark, represents a first step towards the integration of Electronic Design Automation tools that can support the design of all-graphene ICs.
Sandeep Miryala, Mehrdad Montazeri, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE4
2013 SMAC: Smart Systems Co-design
abstract
In this paper we present the concepts and the organization of the FP7 Project SMAC (Smart systems Co-design), an Integrated Project (IP) of the 7th ICT Call under the Objective 3.2 "Smart components and Smart Systems integration". We describe in particular the project objectives and its organization, and how it addresses the challenges of the integration of heterogeneous and conflicting domains that emerge in the design of smart systems. The main outcome of the SMAC project is the development of flexible software platform (the SMAC platform) for smart subsystems/components design include methodologies and EDA tools enabling multi-disciplinary and multi-scale modeling and design, simulation of multi-domain systems, subsystems and components at all levels of abstraction, system integration and exploration for optimization of functional and non-functional metrics.
Nicola Bombieri, Giuliana Drogoudis, Giuliana Gangemi, Renaud Gillon, Enrico Macii, Massimo Poncino, Salvatore Rinaudo, Francesco Stefanni, Dimitrios Trachanis, Mark van Helvoort
DSD5
2013 Delay model for reconfigurable logic gates based on graphene PN-junctions
abstract
In this paper we address the problem of modeling the timing behavior of a new class of reconfigurable logic gates based on electrostatically controlled graphene pn-junctions. These gates naturally behave as a 2-to-1 multiplexer in which the polarity of the input select line can be dynamically reconfigured. Interconnection of multiple gates and proper assignments of the inputs signals allow to implement all the basic Boolean logic functions, and, at a larger scale, any digital circuit.
Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2013 An automated framework for generating variable-accuracy battery models from datasheet information
abstract
Models based on an electrical circuit equivalent have become the most popular choice for modeling the behavior of batteries, thanks to their ease of co-simulation with other parts of a digital system. Such circuit models are actually model templates: the specific values of their electrical elements must be derived by the analysis of the specific battery devices to be modeled. This process requires either to measure the battery characteristics or to derive them from the datasheet. In the latter case, however, very often not all information are available and the model fitting becomes then unfeasible. In this paper we present a methodology for deriving, in a semi-automatic way, circuit equivalent battery models solely from data available in a battery datasheet. In order to account for the different amount of information available, we introduce the concept of “level” of a model, so that models with different accuracy can be derived depending on the available data. The methodology requires only minimal intervention by the designer and it automatically generates MATLAB models once the required data for the corresponding model level are transcribed from the datasheet. Simulation results show that our methodology allows to accurately reconstruct the information reported in the datasheet as well as to derive missing ones.
Massimo Petricca, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED5
2013 A statistical model of cell-to-cell variation in Li-ion batteries for system-level design
abstract
Due to manufacturing tolerances, different battery cells exhibit substantial variability among them, which should be carefully assessed and managed, especially when assembling large battery packs. Cell-to-cell variability has been mostly approached from the point of view of the chemical and physical phenomena, but these studies did not provide a practical solution for the system-level design. In this work, we propose a combined cell-to-cell variation model of the capacity and of the internal resistance of a battery cell that accounts for variability effects in the cell manufacturing process. The model is derived from analytical models for a specific type of Li-ion cell provided in the literature, from which we identify what model parameters can be regarded as true random variables. This allows transforming capacity and internal resistance into the functions of random variables, which can be incorporated into an equivalent circuit model that is suitable for the system-level statistical simulations. The proposed model allows us to qualitatively verify some known properties such as the correlation between capacity and internal resistance, and quantitatively assess the amount of variability and its impact on the design of battery packs.
Donghwa Shin, Massimo Poncino, Enrico Macii, Naehyuck Chang
ISLPED3
2013 Integration of Literature with Heterogeneous Information for Genes Correlation Scoring
abstract
Determining the correlation between biomedical terms is a powerful instrument to help scientist research activity, both to understand experimental results and to design new ones. In particular, a great potential comes from the integration of the many heterogeneous information sources currently available on the Web. In this article we focus on the correlation between genes and biological processes. In this context, we present a methodology for integrating information from biomedical literature with other heterogeneous types of structured information. In particular, the information sources integrated in this work are PubMed abstracts, pathway databases, and NCI thesaurus definitions. The integration is performed at the semantic analysis level using a customized approach we developed to modulate the impact of the different sources on the correlation score. We report the results of a study concerning the impact of the information integration on the correlation score and of the user-level parameters we introduced to modulate the impact of pathway data or NCI definitions with respect to biomedical literature information, depending on the context of the search. To evaluate the methodology, we performed correlation measures on six biological processes and nine genes by comparing the results with and without the integration of pathways and NCI definitions.
Francesco Abate, Andrea Acquaviva, Elisa Ficarra, Enrico Macii
ACM J. Emerg. Technol. Comput. Syst.4
2013 Layout-Driven Post-Placement Techniques for Temperature Reduction and Thermal Gradient Minimization
abstract
With the continuing scaling of CMOS technology, on-chip temperature and thermal-induced variations have become a major design concern. To effectively limit the high temperature in a chip equipped with a cost-effective cooling system, thermal specific approaches, besides low power techniques, are necessary at the chip design level. The high temperature in hotspots and large thermal gradients are caused by the high local power density and the nonuniform power dissipation across the chip. With the objective of reducing power density in hotspots, we propose two placement techniques that spread cells in hotspots over a larger area. Increasing the area occupied by the hotspot directly reduces its power density, leading to a reduction in peak temperature and thermal gradient. To minimize the introduced overhead in delay and dynamic power, we maintain the relative positions of the coupling cells in the new layout. We compare the proposed methods in terms of temperature reduction, timing, and area overhead to the baseline method, which enlarges the circuit area uniformly. The experimental results showed that our methods achieve a larger reduction in both peak temperature and thermal gradient than the baseline method. The baseline method, although reducing peak temperature in most cases, has little impact on thermal gradient.
Wei Liu 0016, Andrea Calimera, Alberto Macii, Enrico Macii, Alberto Nannarelli, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2013 Gelsius: A Literature-Based Workflow for Determining Quantitative Associations between Genes and Biological Processes
abstract
An effective knowledge extraction and quantification methodology from biomedical literature would allow the researcher to organize and analyze the results of high-throughput experiments on microarrays and next-generation sequencing technologies. Despite the large amount of raw information available on the web, a tool able to extract a measure of the correlation between a list of genes and biological processes is not yet available. In this paper, we present Gelsius, a workflow that incorporates biomedical literature to quantify the correlation between genes and terms describing biological processes. To achieve this target, we build different modules focusing on query expansion and document cononicalization. In this way, we reached to improve the measurement of correlation, performed using a latent semantic analysis approach. To the best of our knowledge, this is the first complete tool able to extract a measure of genes-biological processes correlation from literature. We demonstrate the effectiveness of the proposed workflow on six biological processes and a set of genes, by showing that correlation results for known relationships are in accordance with definitions of gene functions provided by NCI Thesaurus. On the other side, the tool is able to propose new candidate relationships for later experimental validation. The tool is available at >http://bioeda1.polito.it:8080/medSearchServlet/.
Francesco Abate, Andrea Acquaviva, Elisa Ficarra, Roberto Piva, Enrico Macii
IEEE ACM Trans. Comput. Biol. Bioinform.5
2012 Application-specific memory partitioning for joint energy and lifetime optimization
abstract
Power management of caches based on turning idle cache lines into a low-energy state is also beneficial for the aging effects caused by Negative Bias Temperature Instability (NBTI), provided that idleness is correctly exploited; unlike energy, aging, being a measure of delay, is in fact a worst-case metric.
Haroon Mahmood, Massimo Poncino, Mirko Loghi, Enrico Macii
DATE4
2012 IR-drop analysis of graphene-based power distribution networks
abstract
Electromigration (EM) has been indicated as the killer effect for copper interconnects. ITRS projections show that for future technologies (22nm and beyond) the on-chip current demand will exceed the physical limit copper metal wires can tolerate. This represents a serious limitation for the design of power distribution networks of next generation ICs. New carbon nanomaterials, governed by ballistic transport, have shown higher immunity to EM, thereby representing potential candidate to replace copper. In this paper we make use of compact conductance models to benchmark Graphene Nanoribbons (GNRs) against copper. The two materials have been used to route a state-of-the-art multi-level power-grid architecture obtained through an industrial 45nm physical design flow. Although the adopted design style is optimized for metal grids, results obtained using our simulation framework show that GNRs, if properly sized, can outperform copper, thus allowing the design of reliable circuits with reduced IR-drop penalties.
Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE3
2012 Middleware services for network interoperability in smart energy efficient buildings
abstract
One of the major challenges in today's economy concerns the reduction in energy usage and CO2footprint in existing Public buildings and Spaces without significant construction works, by an intelligent ICT-based service monitoring and managing the energy consumption. In particular, interoperability between heterogeneous devices and networks, both existing and to be deployed is a key features to create efficient services and holistic energy control policies. In this paper we describe an innovative software infrastructure to provide a web-service based, hardware independent access to the heterogeneous networks of wireless sensor nodes, such as smart plugs for measuring energy motes for temperature, relative humidity and light monitoring. The proposed infrastructure allows easy extension to other networks, thus representing a contribute to the opening of a market for ICT-based customized solutions integrating numerous products from different vendors and offering services from design of integrated systems to the operation and maintenance phases.
Edoardo Patti, Andrea Acquaviva, Francesco Abate, Anna Osello, A. Cocuccio, Marco Jahn, Marc Jentsch, Enrico Macii
DATE8
2012 Investigating the effects of Inverted Temperature Dependence (ITD) on clock distribution networks
abstract
The aggressive scaling of CMOS technology toward nanometer lengths contributed to the surfacing of many effects that were not appreciable at the micrometer regime. Among them, Inverted Temperature Dependence (ITD) is certainly the most unusual. It manifests itself as a speed up of CMOS gates when the temperature increases, resulting in a reversal of the worst-case condition, i.e., CMOS gates show the largest delay at low temperatures. On the other hand, for metal interconnects an high temperature still holds as worst case condition. The two contrasting behaviors may invalidate the results obtained through standard design flow which do not consider temperature as an explicit variable in their optimizations. In this paper we focus on the impact of ITD on clock distribution networks (CDN), whose function is vital to guarantee the synchronization among physically spaced sequential components of digital circuits. Using our simulation framework, we characterized the thermal behavior of a clock tree mapped onto an industrial 65nm CMOS technology and obtained using a standard synthesis tool. Results demonstrate the presence of ITD at low operating voltages and open new potential research scenarios into the EDA field.
Alessandro Sassone, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino, Richard Goldman, Vazgen Melikyan, Eduard Babayan, Salvatore Rinaudo
DATE4
2012 NBTI effects on tree-like clock distribution networks
abstract
Negative Bias Temperature Instability (NBTI) is considered one of the most critical device reliability concerns in nanometer CMOS technologies, because it causes devices to exhibit a temporal drift of performance over time.
Wei Liu 0016, Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI5
2012 Applying textural features to the classification of HEp-2 cell patterns in IIF images
Santa Di Cataldo, Andrea Bottino, Elisa Ficarra, Enrico Macii
ICPR4
2012 Energy-optimal caches with guaranteed lifetime
abstract
This work addresses the aging of the memory sub-system due to NBTI (Negative Bias Temperature Instability) in systems that have to provide a guaranteed level of service, and specifically, a guaranteed lifetime.
Mirko Loghi, Haroon Mahmood, Andrea Calimera, Massimo Poncino, Enrico Macii
ISLPED5
2012 Aging-aware caches with graceful degradation of performance
abstract
Aging of transistors can substantially shorten the lifetime of devices in sub-nanometric technologies. Without any countermeasure, the first component which becomes unreliable will determine the life span of an entire device. This problem is even more relevant for memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. Traditional implementation of power management by turning idle cache lines into a low-energy state can also mitigate the aging effects caused by Negative Bias Temperature Instability (NBTI) provided that idleness is correctly exploited. In this work, we propose a cache structure which deals with cell failures by gracefully degrading its performance. By this partitioning-based strategy, various sub-blocks will become unreliable at different times, and the cache will keep functioning with reduced efficiency. Coupling such aging mitigation with the resulting energy reduction techniques we can obtain up to 2.5x lifetime extension and 40% energy savings with respect to a power managed cache.
Haroon Mahmood, Massimo Poncino, Mirko Loghi, Enrico Macii
VLSI-SoC4
2012 Bellerophontes: an RNA-Seq data analysis framework for chimeric transcripts discovery based on accurate fusion model
abstract
MOTIVATION: Next-generation sequencing technology allows the detection of genomic structural variations, novel genes and transcript isoforms from the analysis of high-throughput data. In this work, we propose a new framework for the detection of fusion transcripts through short paired-end reads which integrates splicing-driven alignment and abundance estimation analysis, producing a more accurate set of reads supporting the junction discovery and taking into account also not annotated transcripts. Bellerophontes performs a selection of putative junctions on the basis of a match to an accurate gene fusion model. RESULTS: We report the fusion genes discovered by the proposed framework on experimentally validated biological samples of chronic myelogenous leukemia (CML) and on public NCBI datasets, for which Bellerophontes is able to detect the exact junction sequence. With respect to state-of-art approaches, Bellerophontes detects the same experimentally validated fusions, however, it is more selective on the total number of detected fusions and provides a more accurate set of spanning reads supporting the junctions. We finally report the fusions involving non-annotated transcripts found in CML samples. AVAILABILITY AND IMPLEMENTATION: Bellerophontes JAVA/Perl/Bash software implementation is free and available at http://eda.polito.it/bellerophontes/.
Francesco Abate, Andrea Acquaviva, Giulia Paciello, Carmelo Foti, Elisa Ficarra, Alberto Ferrarini, Massimo Delledonne, Ilaria Iacobucci, Simona Soverini, Giovanni Martinelli, Enrico Macii
Bioinform.11
2011 Motion Artifact Correction in ASL images: An Improved Automated Procedure
abstract
Arterial Spin Labelling (ASL) is a perfusion MRI technique with tremendous applications in the study of biological markers and prognostic factors of brain tumors and in the assessment of neural diseases, moreover, it is completely non-invasive as it uses the magnetically inverted blood of the patient as an endogenous tracer. Unfortunately this powerful method is only viable in very limited conditions due to its extreme sensitivity to artifacts originated by head motion, that are not effectively addressed by the current software solutions. This paper presents a motion correction procedure that addresses this issue and provides improved solutions to enhance ASL images of the brain in presence of severe head motion. Experimental results run on a motion-affected pCASL dataset show the concept and demonstrate the superiority of our proposed procedure compared to standard 3D registration.
Santa Di Cataldo, Elisa Ficarra, Andrea Acquaviva, Enrico Macii
BIBM4
2011 Partitioned cache architectures for reduced NBTI-induced aging
abstract
Conventional power management knobs such as voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused Negative Bias Temperature Instability (NBTI). Such a benefit can be especially exploited in SRAM memories, which are particularly sensitive to NBTI effects: given their symmetric structure, they cannot in fact take advantage of value-dependent recovery. We propose an architectural solutions that is based on the idea of partitioning a memory into multiple banks of identical size. While this organization has been widely used for reducing both dynamic and static power, its exploitation for aging benefits requires proper management of the existing idleness of the various banks. This can be achieved by means of a sort of time-varying addressing scheme in which addresses are mapped to different banks over time in such a way that the idleness is uniformly distributed over all the banks. Experimental analysis shows that it is possible to simultaneously reducing leakage power and aging in caches, with minimal overhead and without modifying the internal structure of the SRAM arrays.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
DATE3
2011 Buffering of frequent accesses for reduced cache aging
abstract
Previous works have shown that typical power management knobs such as voltage scaling or power gating can also be exploited to reduce aging phenomena caused by Negative Bias Temperature Instability (NBTI). We propose a scheme for power-managed caches that allows to significantly improving the aging of the cache thanks to the use of a small buffer that stores a copy of the lines that are most critical for aging, that is, the ones with the least opportunity of being power-managed; by using the buffer instead of the cache when accessing these critical lines, the original cache is preserved and its lifetime is significantly prolonged. As a side effect, this scheme improves total power since the less energy-hungry buffer is accessed most of the time. Experimental analysis shows this scheme allows to achieve significant (>3x on average) lifetime extensions for the cache, with a concurrent energy saving between 18 and 24%, depending on cache size.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2011 miREE: miRNA Recognition Elements Ensemble
abstract
BACKGROUND: Computational methods for microRNA target prediction are a fundamental step to understand the miRNA role in gene regulation, a key process in molecular biology. In this paper we present miREE, a novel microRNA target prediction tool. miREE is an ensemble of two parts entailing complementary but integrated roles in the prediction. The Ab-Initio module leverages upon a genetic algorithmic approach to generate a set of candidate sites on the basis of their microRNA-mRNA duplex stability properties. Then, a Support Vector Machine (SVM) learning module evaluates the impact of microRNA recognition elements on the target gene. As a result the prediction takes into account information regarding both miRNA-target structural stability and accessibility. RESULTS: The proposed method significantly improves the state-of-the-art prediction tools in terms of accuracy with a better balance between specificity and sensitivity, as demonstrated by the experiments conducted on several large datasets across different species. miREE achieves this result by tackling two of the main challenges of current prediction tools: (1) The reduced number of false positives for the Ab-Initio part thanks to the integration of a machine learning module (2) the specificity of the machine learning part, obtained through an innovative technique for rich and representative negative records generation. The validation was conducted on experimental datasets where the miRNA:mRNA interactions had been obtained through (1) direct validation where even the binding site is provided, or through (2) indirect validation, based on gene expression variations obtained from high-throughput experiments where the specific interaction is not validated in detail and consequently the specific binding site is not provided. CONCLUSIONS: The coupling of two parts: a sensitive Ab-Initio module and a selective machine learning part capable of recognizing the false positives, leads to an improved balance between sensitivity and specificity. miREE obtains a reasonable trade-off between filtering false positives and identifying targets. miREE tool is available online at http://didattica-online.polito.it/eda/miREE/
Paula Helena Reyes-Herrera, Elisa Ficarra, Andrea Acquaviva, Enrico Macii
BMC Bioinform.4
2011 Fast Computation of Discharge Current Upper Bounds for Clustered Power Gating
abstract
The capability of accurately estimating an upper bound of the maximum current drawn by a digital macroblock from the ground or power supply line constitutes a major asset of automatic power-gating flows. In fact, the maximum current information is essential to properly size the sleep transistor in such a way that speed degradation and signal integrity violations are avoided. Loose upper bounds can be determined with a reasonable computational cost, but they lead to oversized sleep transistors. On the other hand, exact computation of the maximum drawn current is an NP-hard problem, even when conservative simplifying assumptions are made on gate-level current profiles. In this paper, we present a scalable algorithm for tightening upper bound computation, with a controlled and tunable computational cost. The algorithm exploits state-of-the-art commercial timing analysis engines, and it is tightly integrated into an industrial power-gating flow for leakage power reduction. The results we have obtained on large circuits demonstrate the scalability and effectiveness of our estimation approach.
Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Row-Based Power-Gating: A Novel Sleep Transistor Insertion Methodology for Leakage Power Optimization in Nanometer CMOS Circuits
abstract
Leakage power has become a serious concern in nanometer CMOS technologies, and power-gating has shown to offer a viable solution to the problem with a small penalty in performance. This paper focuses on leakage power reduction through automatic insertion of sleep transistors for power-gating. In particular, we propose a novel, layout-aware methodology that facilitates sleep transistor insertion and virtual-ground routing on row-based layouts. We also introduce a clustering algorithm that is able to handle simultaneously timing and area constraints, and we extend it to the case of multi-Vtsleep transistors to increase leakage savings. The results we have obtained on a set of benchmark circuits show that the leakage savings we can achieve are, by far, superior to those obtained using existing power-gating solutions and with much tighter timing and area constraints.
Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.4
2010 An Automated Tool for Scoring Biomedical Terms Correlation Based on Semantic Analysis
abstract
The considerable improvement in the biotechnogical field and the adoption of screen techniques such as high throughput arrays produced a spread of biological and genetical data and scientific papers, mostly diffused on the Web. Even if the huge amount of available information represents a major step forward for the biomedical research field, the main effort for a scientist is to evaluate the correlation among the concepts both in qualitative and in quantitative terms. In the presented work, exploiting the richness of the UMLS Metathesaurus in combination with the novelty of the literature provided by PubMed, an automatic flow aimed at the scoring the semantic correlation among biomedical terms (e. g. biomolecules and biological processes) is proposed. The experiments show that obtained correlations are fully coherent with the information coming from the biological literature. The accuracy of the results remarks the importance of combining text mining techniques with complete and well structured thesaurus, such as UMLS.
Francesco Abate, Elisa Ficarra, Andrea Acquaviva, Enrico Macii
CISIS4
2010 MicroRNA Target Prediction and Exploration through Candidate Binding Sites Generation
abstract
Gene regulation is one of the most important processes in the molecular biology, in the last years the microRNA molecule, one of the non-coding RNAs involved in the process, has been the focus of attention for several studies. The computational research on this area has gained a notable importance, considering the low amount of experimental information available and the lack of understanding of the microRNA binding mechanism. This article deals with the microRNA-target prediction and presents an innovative method for it. First it generates a set of promising binding sites for a given microRNA using a Genetic Algorithm, at the same time a set of target genes is selected based on the biological process under study. Secondly the set of promising binding sites is mapped into the selected set of target genes, in order to provide real binding sites and finally the resulting targets are filtered according to a biological or structural property. The objectives are to provide a flexible method that is capable of incorporating easily new knowledge, is independent of availability of the experimental information and is able to give hints on the research towards new characteristics among the microRNA binding sites such as motifs. The results present some of this novel properties and present a comparison with the most frequently used methods in the field.
Paula Helena Reyes-Herrera, Andrea Acquaviva, Elisa Ficarra, Enrico Macii
CISIS4
2010 Panel: First commandment at least, do nothing well!
Marco Casale-Rossi, Giovanni De Micheli, Antun Domic, Enrico Macii, Piero Perlo, Andreas Wild, Roberto Zafalon
DATE4
2010 Post-placement temperature reduction techniques
abstract
With technology scaled to deep submicron era, temperature and temperature gradient have emerged as important design criteria. We propose two post-placement techniques to reduce peak temperature by intelligently allocating whitespace in the hotspots. Both methods are fully compliant with commercial technologies, and can be easily integrated with state-of-the-art thermal-aware design flow. Experiments in a set of tests on circuits implemented in STM 65nm technologies show that our methods achieve better peak temperature reduction than directly increasing circuit's area.
Wei Liu 0016, Alberto Nannarelli, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE4
2010 An integrated thermal estimation framework for industrial embedded platforms
abstract
Next generation industrial embedded platforms require the development of complex power and thermal management solutions. Indeed, an increasingly fine and intrusive thermal control is required because of temperature impact on leakage and reliability. To be effective, the implementation of these policies involves decisions that must be taken during various phases along the design process, to enable the development of architectural level countermeasures and the required hardware knobs, such as power modes, power supply regulation granularity and the number of on-chip temperature sensors. As a consequence, a framework allowing thermal estimation exploiting design-time information is desirable.
Andrea Acquaviva, Andrea Calimera, Alberto Macii, Massimo Poncino, Enrico Macii, Matteo Giaconia, Claudio Parrella
ACM Great Lakes Symposium on VLSI5
2010 Aging effects of leakage optimizations for caches
abstract
Besides static power consumption, sub-90nm devices have to account for NBTI effects, which are one of the major concerns about system reliability.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2010 Thermal-aware floorplanning exploration for 3D multi-core architectures
abstract
Thermal effects are becoming increasingly important in today's sub-micron technologies. Thermal issues affect the performance, the reliability and the cooling costs of integrated systems. High peak temperatures are of major concern in modern 3D designs, where the stacking of multiple layers leads to higher power densities. Therefore, the integration of the thermal-aware design during the initial phases of the design can reduce the cost and the time-to-market of the resulting product. An efficient floorplanning in terms of thermal effects will reduce the appearance of critical hotspots and will spread heat across the chip area.
David Cuesta, José Luis Ayala, J. Ignacio Hidalgo, Massimo Poncino, Andrea Acquaviva, Enrico Macii
ACM Great Lakes Symposium on VLSI6
2010 Analysis of NBTI-induced SNM degradation in power-gated SRAM cells
abstract
Temporal reliability degradation mechanism, and NBTI in particular, are especially critical for SRAM cells. In fact, unlike logic gates, which under some conditions can be forced into an NBTI-immune state, SRAM cells are always subject to aging, whatever value they are storing. In this work, we first quantify the aging, in terms of degradation of the signal-to-noise margin (SNM), of an SRAM cell as a function of the value stored in the cell, on a 45nm industrial technology. Then, we show how it is possible, by applying power gating to the memory cell, to further reduce the SNM degradation. Finally, we study the joint effect of power gating and bit control techniques.
Andrea Calimera, Enrico Macii, Massimo Poncino
ISCAS2
2010 Dynamic indexing: concurrent leakage and aging optimization for caches
abstract
Previous works have shown that the traditional implementations of power management (i.e., using power gating or voltage scaling) can also mitigate the aging effect induced by Negative Bias Temperature Instability (NBTI), due to the partial recovery that occurs during the idle intervals used by power management. However, such a potential has been exploited only partially because of the different nature of energy and aging: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually turn into limited aging reductions. We address this problem in the context of caches, for which idleness is related to their access pattern. We propose a dynamic indexing scheme, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the cache lines. In this way it is possible to fully use the leakage optimization potential and to extend the lifetime of a cache. Experimental analysis shows that it is possible to obtain caches that are effectively aging-free, without any penalty in leakage energy reduction.
Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino
ISLPED3
2010 Power-aware partitioning of data converters
abstract
Serial data streaming, one of the most important functions in modern communication systems, is becoming more and more power consuming as bit-rate is increasing without standstill. In this work, we propose a novel technique for partitioning conventional N-bit registers in standard data converters, in order to reduce their switching activity, and therefore power consumption. The architecture here presented have a very low area overhead with respect to the standard ones for serializers and, furthermore, it allows different (i.e., custom) configurations for the partitioning. The proposed method even allows to extract idleness conditions of register banks in order to apply the well-known clock-gating technique to the circuit and thus furtherly reducing the total power consumption. This method has been applied to different data converters (i.e., serializers) in a base-band radio within an ultra low-power industrial design and the results highlight the effectiveness of the proposed technique.
Alberto Bonanno, Alberto Bocca, Alberto Macii, Enrico Macii
VLSI-SoC4
2010 Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Subbanking
abstract
Partitioning a memory into multiple blocks that can be independently accessed is a widely used technique to reduce its dynamic power. For embedded systems, its benefits can be even pushed further by properly matching the partition to the memory access patterns. When leakage energy comes into play, however, idle memory blocks must be put into a proper low-leakage sleep state to actually save energy when not accessed. In this case, the matching becomes an instance of the power management problem, because moving to and from this sleep state requires additional energy. In this work, we propose an effective solution to the problem of the leakage-aware partitioning of a memory into disjoint subblocks; in particular, we target scratchpad memories, which are commonly used in some embedded systems as a replacement for caches. We show that, although the solution space is extremely large (for a N--block partition, all the combinations of N-1 address boundaries) and nonconvex, it is possible to prove a nontrivial property that considerably reduces the number of partition boundaries to be enumerated, therefore, making exhaustive exploration feasible. We are thus able to provide an optimal solution to the leakage-aware partitioning problem. Experiments on a different sets of embedded applications have shown that total energy savings larger than 60 percent on average can be obtained, with a marginal overhead in execution time, thanks to an effective implementation of the low-leakage sleep state.
Mirko Loghi, Olga Golubeva, Enrico Macii, Massimo Poncino
IEEE Trans. Computers3
2010 NBTI-Aware Clustered Power Gating
abstract
The emergence of Negative Bias Temperature Instability (NBTI) as the most relevant source of reliability in sub-90nm technologies has led to a new facet of the traditional trade-off between power and reliability. NBTI effects in fact manifest themselves as an increase of the propagation delay of the devices over time, which adds up to the delay penalty incurred by most low-power design solutions. This implies that, given a desired lifetime of a circuit (i.e., a given performance target at some point in time), a power-managed component will fail earlier than a nonpower-managed one. In this work, we show how it is possible to partially overcome this conflict, by leveraging the benefits in terms of aging provided by power-gating (i.e., by using switches that disconnect a logic block from the ground). Thanks to some electrical properties, it is possible to nullify aging effects during standby periods. Based on this important property, we propose a methodology for a NBTI-aware power gating that allows synthesizing low-leakage circuits with maximum lifetime.
Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Trans. Design Autom. Electr. Syst.2
2010 Temperature-Insensitive Dual- Vth Synthesis for Nanometer CMOS Technologies Under Inverse Temperature Dependence
abstract
With the scaling of CMOS technologies, the gap between nominal supply voltage and threshold voltage has decreased significantly. This trend is further amplified in low-power nanometer libraries, which feature cells with identical size and functionality, but different threshold voltages. As a consequence, different cells may have different delay behaviors as the temperature varies within a circuit. For instance, cells with low-threshold devices may experience an increase in delay when temperature increases, whereas cells using high-threshold devices may experience the opposite behavior. The latter effect, also known as inverse temperature dependence (ITD), poses new challenges to circuit designers. Besides making timing analysis more difficult, ITD has important and unforeseeable consequences for power-aware logic synthesis. This paper describes the impact that ITD may have on the design of nanometer circuits. We also provide a threshold voltage assignment algorithm for dual threshold voltage synthesis, which guarantees temperature-insensitive operation of the circuits, together with a significant reduction of both leakage and total power consumption. Experiments performed on a set of standard benchmarks show timing compliance at any operating temperature, and an average leakage reduction around 28% compared to circuits synthesized with a standard synthesis flow that does not take ITD into account. We also apply our proposed synthesis algorithm to a realistic case study consisting of a 32-bit, IEEE-754 floating point unit.
Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Enabling concurrent clock and power gating in an industrial design flow
abstract
Clock-gating and power-gating have proven to be very effective solutions for reducing dynamic and static power, respectively. The two techniques may be coupled in such a way that the clock-gating information can be used to drive the control signal of the power-gating circuitry, thus providing additional leakage minimization conditions w.r.t. those manually inserted by the designer. This conceptual integration, however, poses several challenges when moved to industrial design flows. Although both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. This paper presents a layout-oriented synthesis flow which integrates the two techniques and that relies on leading-edge, commercial EDA tools. Starting from a gated-clock netlist, we partition the circuit in a number of clusters that are implicitly determined by the groups of cells that are clock-gated by the same register. Using a row-based granularity, we achieve runtime leakage reduction by inserting dedicated sleep transistors for each cluster. The entire flow has been benchmarked on a industrial design mapped onto a commercial, 65 nm CMOS technology library.
Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino
DATE4
2009 Physically clustered forward body biasing for variability compensation in nanometer CMOS design
abstract
Nanometer CMOS scaling has resulted in greatly increased circuit variability, with extremely adverse consequences on design predictability and yield. A number of recent works have focused on adaptive post-fabrication tuning approaches to mitigate this problem. Adaptive Body Bias (ABB) is one of the most successful tuning ldquoknobsrdquo in use today in high-performance custom design. Through forward body bias (FBB), the threshold voltage of the CMOS devices can be reduced after fabrication to bring the slow dies back to within the range of acceptable specs. FBB is usually applied with a very coarse core-level granularity at the price of a significantly increased leakage power. In this paper, we propose a novel, physically clustered FBB scheme on row-based standard-cell layout style that enables selective forward body biasing of only of the rows that contain most timing critical gates, thereby reducing leakage power overhead. We propose exact and heuristic algorithms to partition the design and allocate optimal body bias voltages to achieve minimum leakage power overhead. This style is fully compatible with state-of-the-art commercial physical design flows and imposes minimal area blowup. Benchmark results show large leakage power savings with a maximum savings of 30% in case of 5% compensation and 47.6% in case of 10% compensation with respect to block-level FBB and minimal implementation area overhead.
Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Giovanni De Micheli, Enrico Macii
DATE5
2009 NBTI-aware sleep transistor design for reliable power-gating
abstract
Negative Bias Temperature Instability (NBTI) has been regarded as most important source of reliability of CMOS devices, and specifically pMOS transistors.
Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2009 Using soft-edge flip-flops to compensate NBTI-induced delay degradation
abstract
We present a low-overhead solution to tackle the delay increase caused by Negative Bias Temperature Instability (NBTI), which has emerged as the most critical reliability issue in sub-90nm technology nodes.
Karthik Duraisami, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2009 Placement-aware Clustering for Integrated Clock and Power Gating
abstract
Clock-gating and power-gating are the most widely used solutions for reducing dynamic and static power. They can be potentially integrated so that clock-gating conditions can be used to control the power-gating circuitry thus also reducing static power. This integration becomes however difficult when applied in an industrial design flow. Even if both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available.
Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino
ISCAS4
2009 NBTI-aware power gating for concurrent leakage and aging optimization
abstract
Power and reliability are known to be intrinsically conflicting metrics: traditional solutions to improve reliability such as redundancy, increase of voltage levels, and up-sizing of critical devices do contrast with traditional low-power solutions, which rely on small devices and scaled supply voltages. The emergence of Negative Bias Temperature Instability (NBTI) as the most relevant source of unreliability in sub-90nm technologies has even exacerbated this incompatibility of the two metrics: NBTI manifests itself as an increase of the propagation delay over time, which adds up to the delay penalty introduced by most low-power design solutions. In this work, we show how the most widely adopted leakage reduction solution, that is, power-gating, can overcome this conflict, and how it can be used to naturally reduce the effects of NBTI on delay. Based on this important property, we present a methodology for NBTI-aware power gating that allows synthesizing low-leakage circuits with maximum lifetime.
Andrea Calimera, Enrico Macii, Massimo Poncino
ISLPED2
2009 Editorial
Enrico Macii
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Segmentation of nuclei in cancer tissue images: Contrasting active contours with morphology-based approach
abstract
In this paper we present a fully automated morphology-based technique for segmentation of nuclei in cancer tissue images and we compare it with a common technique for biomedical image processing, namely active contours. We discuss the limitations of active contours in the processing of immunohistochemical images characterized by heterogeneously stained nuclear region and noise caused by the presence of multiple tissue layers in the sample. We describe the integration of the proposed approach in a fully automated protein activity quantification tool. Finally, we demonstrate and motivate through extensive experiments that our fully automated morphology-based approach provides better accuracy compared to various active contours implementations.
Santa Di Cataldo, Elisa Ficarra, Andrea Acquaviva, Enrico Macii
BIBE4
2008 Optimal MTCMOS Reactivation Under Power Supply Noise and Performance Constraints
abstract
Sleep transistor insertion is one of today's most promising and widely adopted solutions for controlling stand-by leakage power in nanometer circuits. Although single-cycle power mode transition reduces wake-up latency, it originates large discharge current spikes, thereby causing IR-drop and inductive ground bounce for the surrounding circuit blocks. We propose a new reactivation solution which helps in controlling power supply fluctuations and in achieving minimum reactivation times. Our structure limits the turn-on current below a given threshold through sequential activation of the sleep transistors, which are connected in parallel and are sized using a novel optimal sizing algorithm. The proposed methodology is validated using HSPICE simulations of several benchmark circuits, which have been synthesized onto a commercial 65 nm CMOS technology library.
Andrea Calimera, Luca Benini, Enrico Macii
DATE3
2008 A Scalable Algorithmic Framework for Row-Based Power-Gating
abstract
Leakage power is a serious concern in nanometer CMOS technologies. In this paper we focus on leakage reduction through automatic insertion of sleep transistors for power gating in standard cell based designs. In particular, we propose clustering algorithms for row- based power-gating methodology which is based on using rows of the layout as the granularity for clustering. Our clustering methodology does timing and area constraint driven power-gating in contrast to only timing driven power-gating as proposed in the previous works. We present two distinct clustering algorithms with different accuracy-efficiency trade-off. An optimal one, which exploits a 0-1 or binary integer programming approach, and a heuristic one, which resorts to an implicit enumeration of the layout rows. Results show that, for all the benchmarks, the leakage power savings, as compared to previous techniques, are more than 75% when we have the same timing constraints but half sleep transistor area and at least 60% when area constraint is set at one fourth. We also show that we can perform clustering with no speed degradation and achieve maximum leakage power savings up-to 83%.
Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
DATE5
2008 Process Variation Tolerant Pipeline Design Through a Placement-Aware Multiple Voltage Island Design Style
abstract
A common technique to compensate process variation induced performance deviations during post-silicon testing consists of the dynamic adaptation of processor voltage. This however comes at a significant power cost. We envision multi supply voltage design (MSV) as a promising technique to mitigate such power overhead. Voltage islands are widely recognized as the state-of-the-art in MSV design. In this paper, we develop a novel design methodology that leverages voltage islands to compensate process variations through a commercial synthesis flow. Possible violation scenarios of performance requirements in fabricated chips are pre-characterized at design time through statistical static timing analysis. Then, during post-silicon testing the supply voltage of a proper number of voltage islands is raised depending on the actual violation scenario, thus bringing performance back within nominal values. Voltage islands are generated by exploiting cell proximity for minimal perturbation of performance pre-optimized placements.
Bonesi Stefano, Davide Bertozzi, Luca Benini, Enrico Macii
DATE4
2008 Integrating Clock Gating and Power Gating for Combined Dynamic and Leakage Power Optimization in Digital CMOS Circuits
abstract
Clock Gating and Power Gating are two of the most effective techniques that are applied today for reducing dynamic and leakage power, respectively, in digital CMOS circuits. The combined use of the two solutions, however, poses some challenges in terms of practical integration of the required control logic and the power/timing overhead associated to it. This paper presents an analysis methodology and a prototype CAD tool that support the designer in understanding when the joint application of Clock Gating and Power Gating may result in significant power savings.
Enrico Macii, Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Massimo Poncino
DSD1
2008 Temperature-insensitive synthesis using multi-vt libraries
abstract
Temperature fluctuations can alter the delay in MOS circuits. However, increases in temperature do not always lead to a corresponding increase in circuit delay, specifically when operating at low supply voltages. Instead a temperature inversion effect can be observed on the delay of MOS devices under certain conditions, where the delay actually decreases as temperature increases. Given these non-monotonic effects, guaranteeing timing correctness can no longer be achieved simply by characterizing the design under worst case (i.e., high temperature) conditions. In this paper, we present a synthesis methodology in which multi-Vth design is used to generate temperature-insensitive circuits, while minimizing leakage power dissipation as a side-effect. Our experiments with ISCAS benchmark circuits demonstrate the promise of this approach and show that significant reduction in static power is also possible.
Andrea Calimera, Enrico Macii, Massimo Poncino, R. Iris Bahar
ACM Great Lakes Symposium on VLSI2
2008 Energy efficiency bounds of pulse-encoded buses
abstract
Pulse-encoded buses, (i.e., in which a transition is encoded as a pulse) have recently emerged as an effective solution to solve crosstalk issues in global interconnects, since they suppress transitions in opposite directions by construction. As a side effect, this also reduces energy, since coupling capacitances in deep-submicron technologies are larger than ground capacitances. Furthermore, a single pulse consumes less energy than a conventional transition, because the limited length does not fully implies extra transitions since all input transitions are encoded with a pulse.
Karthik Duraisami, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2008 Optimal sleep transistor synthesis under timing and area constraints
abstract
Leakage power reduction in nano-CMOS designs has gained tremendous interest both in academia and industry. Many techniques have been proposed in the literature for leakage power reduction and one of the prominent techniques for leakage power reduction is the use of sleep transistors as power-gating elements to cut-off sub-threshold leakage current in circuits when they are in stand-by mode. Although sleep transistor insertion is very effective in cutting-off leakage, it also incurs timing, area and routing overhead. Since most of the sleep transistor insertion methodologies do post layout insertion, care should be taken such that there is minimal perturbation of the original layout. Over design of sleep transistors cells and sub-optimal sleep transistor placement must be avoided to achieve final design closure. Since the sleep transistor area plays an important and prominent role in this aspect, it necessitates for optimal sleep transistor sizing and synthesis technique under area constraints. In this paper, we first provide a methodology for optimal sleep transistor synthesis under given area constraints. We then apply our technique to the general timing and area constraint driven row-based power-gating methodology proposed in [13] and show how optimal low leakage designs with constraints on timing and area can be designed
Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI5
2008 On quantifying the figures of merit of power-gating for leakage power minimization in nanometer CMOS circuits
abstract
Power-gating has proved to be one of the most effective solutions for reducing stand-by leakage power in nanometer-scale CMOS circuits, and different strategies and algorithms for its application have been proposed recently. Unfortunately, power- gating comes with its own set of costs: Performance degradation, area increase, dynamic power increase and routing congestion. When a decision to power-gate a design has to be taken, pros and cons of power-gating have to be properly weighted to achieve optimal results. In this paper, we define "Figures of Merit" (FoMs) for power-gating, which can be used by designers to better understand the benefits and costs of power-gating, thereby allowing them to achieve optimal results. We then quantify the FoMs by applying a state-of-the-art, industry-strength power- gating flow on a set of designs implemented onto an industrial 65 nm CMOS process, and provide insightful discussion on how optimum power-gating can be achieved.
Ashoka Visweswara Sathanur, Andrea Calimera, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ISCAS6
2008 Reducing leakage power by accounting for temperature inversion dependence in dual-Vt synthesized circuits
abstract
The effects of temperature on delay depend on several parameters, such as cell size, load, supply voltage, and threshold voltage. In particular, variations in Vth can yield a temperature inversion effect causing a decreases of cell delay as temperature increases. This phenomenon, besides affecting timing analysis of a design, has important and unforeseeable consequences on power optimization techniques. In this paper, we focus on the impact of such effects on multi-Vt design; in particular, we show how traditional dual-Vt optimization may yield timing errors in circuits by ignoring temperature effects. Moreover, we present a temperature-aware dual-Vt optimization technique that reduces leakage power and can guarantee that the circuit is timing feasible at the boundary temperatures provided by the technology library. Our experiments show an average 27% leakage reduction with respect to a non temperature-aware design flow.
Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino
ISLPED3
2008 Multiple power-gating domain (multi-VGND) architecture for improved leakage power reduction
abstract
Row-based power-gating has recently emerged as a meet-in-the-middle sleep transistor insertion paradigm between cell-level and block-level granularity, in which each layout row defines the unit of gating, and different rows can be clustered and share the same sleep transistor. Previous works, however, assume the availability of a single virtual ground voltage, thus making the decision of whether to gate or not a given cluster a binary choice: a cluster is either gated or not.
Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED4
2008 Implementation of a thermal management unit for canceling temperature-dependent clock skew variations
Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Alberto Macii, Enrico Macii, Massimo Poncino
Integr.6
2008 Editorial
Enrico Macii
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Dynamic Thermal Clock Skew Compensation Using Tunable Delay Buffers
abstract
The thermal gradients existing in high-performance circuits may significantly affect their timing behavior, in particular, by increasing the skew of the clock net and/or altering hold/setup constraints, possibly causing the circuit to operate incorrectly. The knowledge of the spatial distribution of temperature can be used to properly design a clock network that is able to compensate such thermal non-uniformities. However, redesign of the clock network is effective only if temperature distribution is stationary, i.e., does not change over time. In this paper, we specifically address the problem of dynamically modifying the clock tree in such a way that it can compensate for temporal variations of temperature. This is achieved by exploiting the buffers that are inserted during the clock network generation, by transforming them into tunable delay elements. Temperature-induced delay variations are then compensated by applying the proper tuning to the tunable buffers, which is computed offline and stored in a tuning table inserted in the design. We propose an algorithm to minimize the number of inserted tunable buffers, as well as their tunable range (which directly relates to complexity). Results show that clock skew is kept within original bounds with worst-case power and area penalty of 3.5% and 5.5% respectively.
Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.7
2007 Gene-Markers Representation for Microarray Data Integration
abstract
When analyzing the relationship between genes under different scenarios, the integration of different microarray experiments becomes a relevant task. This paper presents a framework to address some intrinsic problems of integration, due for instance to scaling issues, error bias, different experimental conditions or technology and protocols. Our approach projects original microarray data in a common transformed space to create a common representation of different microarray datasets. This approach allows us to integrate data from various microarray platforms or microarrays based on different experimental conditions. We validate our framework with experiments on real microarray datasets. The results suggest that our approach can be a profitably exploited for microarray data integration and further gene expression analysis applications.
Elena Baralis, Elisa Ficarra, Alessandro Fiori, Enrico Macii
BIBE4
2007 Selection of Tumor Areas and Segmentation of Nuclear Membranes in Tissue Confocal Images: A Fully Automated Approach
abstract
An accurate and standardized technique for tumor tissue segmentation is a critical step for monitoring and quantifying the activity of specific families of pro- teins involved in multi-factorial genetic pathologies. However, fully automated tissue and cell segmentation in clinical images presents many challenges related to the characteristics of the images that make traditional approaches substantially ineffective or incomplete. In this paper we present a fully-automated algorithm that is able to perform accurate and fast segmentation of tissue images. Experimental results on several real-life datasets demonstrate the high level of accuracy achievable thanks to our approach.
Santa Di Cataldo, Elisa Ficarra, Enrico Macii
BIBM3
2007 Early Power-Aware Design & Validation: Myth or Reality?
Gila Kamhi, Stephen Bailey Mentor, Wolfgang Nebel, Y. C. Wong, Juergen Karmann, Enrico Macii, Stephen V. Kosonocky, Steve Curtis
DAC7
2007 Architectural leakage-aware management of partitioned scratchpad memories
Olga Golubeva, Mirko Loghi, Massimo Poncino, Enrico Macii
DATE4
2007 Interactive presentation: Efficient computation of discharge current upper bounds for clustered sleep transistor sizing
Ashoka Visweswara Sathanur, Andrea Calimera, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
DATE5
2007 Design of a family of sleep transistor cells for a clustered power-gating flow in 65nm technology
abstract
Clustered sleep transistor insertion is an effective leakage power reduction technique that is well-suited for integration in an automated design flow and offers a flexible tradeoff between area, delay overhead and turn-on transition time. In this work, we focus on the design of a family of sleep transistor cells, fully compatible with the physical design rules of a commercial 65nm CMOS library. We describe circuit-level and layout optimizations, as well as the cell characterization procedure required to support automated sleep transistor cell selection and instantiation in a clustered power-gating insertion flow.
Andrea Calimera, Antonio Pullini, Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI6
2007 Design Exploration of a Thermal Management Unit for Dynamic Control of Temperature-Induced Clock Skew
abstract
Power densities and temperatures in today's high performance circuits have reached alarmingly high levels due to increased scaling in feature sizes. Subsequently, the various techniques used to keep them under control have also created "zones" of varying temperatures, thus contributing to temperature gradients inside the chip. These gradients have detrimental effects on the delay of wires, as resistance in metals increases with temperature. Clock nets are extremely susceptible to this effect, since they run through the entire chip. Different techniques have been proposed to counter the impact of temperature on clock speed; they range from re-designing the clock network assuming a stationary profile to more adaptive solutions that allow to dynamically compensate the clock skew through replacement of the original buffers with a specially designed counterpart, called tunable delay buffers (TDBs). Dynamic skew management based on TDBs calls for the presence on the chip of a thermal management unit (TMU), whose purpose is that of periodically choosing the actual delay that each TDB must provide in order to achieve skew optimization. Preliminary implementations of such a unit for basic assumptions on the distribution of sensors and their accuracy have indicated negligible impact on the original design. This work aims at exploring in detail several issues related to TMU design, pivoting on the fact that sensor distribution and its accuracy could in fact impact the design in a significant way depending on the design. We provide the results of a careful exploration we have performed on a meaningful case study, quantifying values for area and power consumption.
Karthik Duraisami, Prassanna Sithambaram, Ashoka Visweswara Sathanur, Alberto Macii, Enrico Macii, Massimo Poncino
ISCAS5
2007 Locality-driven architectural cache sub-banking for leakage energy reduction
abstract
In most processors, caches account for the largest fraction of onchip transistors, thus being a primary candidate for tackling the leakage problem. Existing architectural solutions usually rely on customized cache structures, which are needed to implement some kind of power management policy. Memory arrays, however, are carefully developed and finely tuned by foundries, and their internal structure is typically non accessible to system designers.
Olga Golubeva, Mirko Loghi, Enrico Macii, Massimo Poncino
ISLPED3
2007 Power-optimal RTL arithmetic unit soft-macro selection strategy for leakage-sensitive technologies
abstract
With the advent of nanoscale technologies, developing power efficient ASICs increasingly requires consideration of static power. An effective approach to make RTL synthesis algorithms and tools leakage-aware consists of the smart inference of RTL macros based on design constraints and optimization directives. This involves exploring the new trade-offs spanned by the design of RTL functional units, as an effect of the features of nanoscale technologies and ofthe power optimizations performed by commercial synthesis tools. This work explores these new trade-offs and proves that making RTL macro selection strategies aware of them results in power savings as high as 43%.
Simone Medardoni, Davide Bertozzi, Enrico Macii
ISLPED3
2007 Timing-driven row-based power gating
abstract
In this paper we focus on leakage reduction through automatic insertion of sleep transistors using a row-based granularity. In particular, we tackle here the two main issues involved in this methodology: (i) Clustering and (ii) the interfacing of power-gated and non power-gated regions within the same block. The clustering algorithm automatically selects an optimal subset of rows that can be power-gated with a tightly controlled delay overhead. We then address the issue of interfacing different gated regions and propose a novel technique to address this issue with minimal area and power penalty. Our approach is compatible with state-of-the art logic and physical synthesis flows and it does not significantly impact design closure. We achieve leakage power reductions as high as 89% for a set of standard benchmarks, with minimum timing and area overhead.
Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED5
2007 Editorial
Enrico Macii
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 In Memoriam: Margarida F. Jacome
abstract
Recounts the career of Dr. Margarida F. Jacome, a Professor with the Department of Electrical and Computer Engineering, University of Texas (UT), Austin. Also includes recollections from her colleagues and students.
Gustavo de Veciana, Marcello Lajolo, Enrico Macii, Sachin S. Sapatnekar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 Computer-Aided Evaluation of Protein Expression in Pathological Tissue Images
abstract
This work presents the first fully-automated computer-aided analysis approach to the quantification of the expression of receptors for the non-small cell lung carcinoma. This immunohistochemical analysis is usually performed by pathologists via visual inspection of tissue samples images. Our techniques streamlines this error-prone and time-consuming process, thereby facilitating analysis and diagnosis. Experimental results on several real-life datasets demonstrate the high quantitative precision of our approach
Elisa Ficarra, Enrico Macii, Giovanni De Micheli, Luca Benini
CBMS2
2006 Enabling fine-grain leakage management by voltage anchor insertion
abstract
Functional unit shutdown based on MTCMOS devices is effective for leakage reduction in aggressively scaled technologies. However, the applicability of MTCMOS-based shutdown in a synthesis-based design flow poses the challenge of interfacing logic blocks in shutdown mode with active units: The outputs of inactive gates can float at intermediate voltages, causing very large short-circuit currents in the active gates they drive. In this paper, we propose two novel low-overhead elementary cells that fully address this issue. These cells can be added to any synthesis library, and they can be inserted into a netlist at the boundary between shutdown and active regions. Our results show that: (i) our cells solve the interfacing problem with minimum overhead; and (ii) a non-intrusive design flow enhancement is sufficient to automatically insert interface cells in post-synthesis netlists
Pietro Babighian, Luca Benini, Alberto Macii, Enrico Macii
DATE4
2006 Thermal resilient bounded-skew clock tree optimization methodology
abstract
The existence of non-uniform thermal gradients on the substrate in high performance IC's can significantly impact the performance of global on-chip interconnects. This issue is further exacerbated by the aggressive scaling and other factors such as dynamic power management schemes and non-uniform gate level switching activity. In high-performance systems, one of the most important problems is clock skew minimization since it has a direct impact on the maximum operating frequency of the system. Since clocks are routed across the entire chip, the presence of thermal gradients can significantly alter their characteristics because wire resistance increases linearly as the temperature increases. This often results in failure to meet original timing constraints thereby rendering the original topology unusable. Therefore it is necessary to perform a temperature aware re-embedding of the original topology to meet timing under these temperature effects. This work primarily explores these issues by proposing two algorithms that re-structure an existing clock tree topology to compensate for such temperature effects and as a result also meet timing constraints
Ashutosh Chakraborty, Prassanna Sithambaram, Karthik Duraisami, Alberto Macii, Enrico Macii, Massimo Poncino
DATE5
2006 Low-power design tools: are EDA vendors taking this matter seriously?
abstract
While transistors per square millimeter and on-chip clock keep scaling smoothly according to Moore’s Law, Vdd does not, nor does Vth. This leads to a dramatic increase in chip power density, and to a significant shift in the balance between dynamic and leakage power. In spite of the recent effort made by EDA vendors in delivering novel solutions that help mitigating the effects on power consumption of technology scaling, the question of whether EDA industry is taking the low-power matter seriously still remains. This session will provide an answer to this intriguing question, by first offering a short review of the state of-the-art in design technologies for dynamic and leakage power minimisation. The session will then continue with a public “trial”, in which OEMs, IDMs, IP and fabless semiconductor vendors will play the role of the public prosecutor, against defendant EDA industry. The court’s ruling will tell us about the future targets the EDA vendors will pursue in low-power design technologies.
Enrico Macii, Massoud Pedram, Dirk Friebel, Robert C. Aitken, Antun Domic, Roberto Zafalon
DATE1
2006 STV-Cache: a leakage energy-efficient architecture for data caches
abstract
We propose a low-leakage cache architecture based on the observation of the spatio-temporal properties of data caches. In particular, we exploit the fact that during the program lifetime a few data values tend to exhibit both spatial and temporal locality in cache, i.e., values that are simultaneously stored by several lines at the same time. Leakage energy can be reduced by turning off those lines and storing these values in a smaller, separate memory. In this work we introduce an architecture that implements such a scheme, as well as an algorithm to detect these special values. We show that by using as few as four values we can achieve 18.45% leakage energy savings, with an additional 13.85% reduction of dynamic energy as a consequence of a reduced average cache access cost.
Kimish Patel, Luca Benini, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2006 Implications of ultra low-voltage devices on design techniques for controlling leakage in NanoCMOS circuits
abstract
Enabled by technology scaling, ultra low-voltage devices have now found wide application in modern VLSI circuits. While low-voltage implies reduced dynamic power, it also signifies increased leakage power, as lower supply voltages are usually paired with lower threshold voltages in order to preserve circuit speed. This originates an increase in sub-threshold leakage currents that constitute, today, one of the most serious bottlenecks to further technology and supply voltage scaling. The need of controlling leakage power in nanometric devices is imposing a significant shift in the way integrated circuits are designed and manufactured. The behavior of devices with nanometric feature sizes is much more sensitive to parameters such as the operating temperature of the circuit, which in the past were neglected. In this paper we quantitatively analyze the leakage control capabilities of some well-established circuit-level design techniques, and assess how the effectiveness of such techniques scales with respect to decreased supply voltages (as induced by technology scaling) and temperature variations, thus providing an interesting insight on how leakage control solutions that are in use today is applicable in future designs
Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Alberto Macii, Enrico Macii, Massimo Poncino
ISCAS6
2006 Low-energy pixel approximation for DVI-based LCD interfaces
abstract
Several options are available for the approximating of color images in hardware, both in terms of the type of transformation (e.g., quantization, dithering) as well as in terms of where the approximation takes place (e.g., graphics controller, frame buffer, LCD controller). In this work, we propose a color approximation approach, orthogonal to other color simplification schemes, which is done during the digital transmission of the color data to the LCD. The proposed technique targets a specific digital standard (namely, DVI) and its serial protocols TMDS, to provide a 45-66% savings in the number of bit transitions on the LCD bus (which translate to a corresponding saving of energy), depending on the allowed degradation of image quality
A. Nurrachmat, Enrico Macii, Massimo Poncino
ISCAS2
2006 Mining Gene Sets for Measuring Similarities
abstract
In recent years, the development of high throughput devices for the massive parallel analyses of genomic data has lead to the generation of large amount of new biological evidences and has triggered the proliferation of data mining algorithms for the extraction of meaningful information. Microarrays for gene expression analyses are part of this revolution and provide important insight in molecular biology often in the form of coherent sets of genes representing previously uncharacterized processes. Large amount of data are continuously produced in this form, and computational approaches can significantly improve the efficient use of these results, since comparison among numbers of genes sets can give new meaningful information at no cost from the experimental biology point of view. To address this opportunity we designed and implemented FIT, a scalable, unsupervised algorithm that quantitatively compares different populations of gene sets using two distinct measures of similarity between any two gene sets. These measures are then used to obtain a summary statistic that describes the tightness of fit between sets belonging to two distinct populations of gene sets. We present the results of FIT on two data sets for the study of Lymphoma and Acute Lymphoblastic Leukemia. In both cases FIT was able to recapitulate the previous analyses on these datasets, to extend the results and to extract information likely to offer potential insights into the underlying biology.
Christine Nardini, Daniele Masotti, Sungroh Yoon, Enrico Macii, Michael D. Kuo, Giovanni De Micheli, Luca Benini
ISCC4
2006 Dynamic thermal clock skew compensation using tunable delay buffers
abstract
The thermal gradients existing in high-performance circuits may significantly affect their timing behavior, in particular by increasing the skew of the clock net and/or altering hold/setup constraints, possibly causing the circuit to operate incorrectly. The knowledge of the spatial distribution of temperature can be used to properly design a clock network that is able to compensate such thermal non-uniformities. However, re-design of the clock network is effective only if temperature distribution is stationary, i.e., does not change over time. In this work, we specifically address the problem of dynamically modifying the clock tree in such a way that it can compensate for temporal variations of temperature. This is achieved by exploiting the buffers that are inserted during the clock network generation, by transforming them into tunable delay elements. Temperature-induced delay variations are then compensated by applying the proper tuning to the tunable buffers, which is computed off-line and stored in a tuning table inserted in the design. We propose an algorithm to minimize the number of inserted tunable buffers, as well as their tunable range (which directly relates to complexity). Results show that clock skew is kept within original bounds with minimum area and power penalty. The maximum increase in power is 23.2% with most benchmarks exhibiting less than 5% increase in power.
Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED7
2006 Reducing Conflict Misses by Application-Specific Reconfigurable Indexing
abstract
The predictability of memory access patterns in embedded systems can be successfully exploited to devise effective application-specific cache optimizations. In this paper, an improved indexing scheme for direct-mapped caches, which drastically reduces the number of conflict misses by using application-specific information, is proposed. The indexing scheme is based on the selection of a subset of the address bits. With respect to similar approaches, the solution has two main strengths. First, owing to an analytical model for the conflict-miss conditions of a given trace, it provides a symbolic algorithm to compute the optimum solution (i.e., the subset of address bits to be used as cache index that minimize the number of conflict misses). Second, owing to a reconfigurable bit selector that can be programmed at run time, it allows the optimal cache indexing to fit to a given application. Results show an average reduction of conflict misses of 24%, measured over a set of standard benchmarks, and for different cache configurations
Kimish Patel, Luca Benini, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Low-overhead state-retaining elements for low-leakage MTCMOS design
abstract
Multi-threshold CMOS (MTCMOS) has shown to be a very effective technique for reducing sub-threshold leakage currents in DSM CMOS designs. Application of the MTC-MOS paradigm to sequential circuits requires the availability of data-retaining elements for storing circuit state during stand-by mode. In this paper we propose two novel circuit schemes for sequential elements featuring low leakage currents in stand-by mode and high-speed/low-dynamic power in active mode. We present post-layout simulation results obtained after parasitic extraction for delay and power of circuits built in 130nm CMOS technology. Our experiments demonstrate several advantages of the proposed schemes over the best previously published solutions.
Pietro Babighian, Luca Benini, Alberto Macii, Enrico Macii
ACM Great Lakes Symposium on VLSI4
2005 Zero clustering: an approach to extend zero compression to instruction caches
abstract
We propose an energy-efficient architecture for instruction caches that relies on dynamic zero compression (DZC), that is, the possibility of reading and writing a single bit for every zero-valued byte [5]. We enhance the basic DZC by using a simple bit permutation to increase the number of zero-valued bytes, so that the corresponding overhead is negligible. The derivation of an effective permutation relies on a heuristic zero clustering algorithm that is based on the knowledge of the memory reference access trace, thus making this solution suitable for application-specific embedded systems. The architecture proposed in this work makes possible the application of zero compression to instruction caches; experiments showed an increase of zero clusters of more than 70% on average, which translates into a 10% improvement in dynamic energy savings with respect to DZC.
Kimish Patel, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI2
2005 Exploring the impact of architectural parameters on energy efficiency of application-specific block-enabled SRAMs
abstract
Application-Specific Block-Enabled (ASBE) SRAMs represent a viable solution for reducing energy consumption in embedded memories. The basic idea behind ASBE architectures is that of partitioning the memory array into a number of non-uniformly sized blocks, such that memory access cost is reduced. The number and sizes of the partitions yielding a minimum power implementation of the SRAM macro is determined by the partitioning algorithm based on the memory access profile obtained as a result of the application (or application mix) executed by the processor. Given the complexity of the design space we are dealing with, there are several degrees of freedom that the partitioning engine may exploit to come up with the most energy-efficient memory architecture. In this paper, we investigate how the quality of the partitioned memory depends on the architectural parameters that define the memory structure (e.g., min and max number of lines per partition, min and max number of words per line, granularity of the partitions); such parameters, in turn, are constrained by the technology and process of choice. We believe that the results presented in this work will provide very useful guidelines for a succesfull adoption of the ASBE approach in practice, as this design paradigm is gaining a lot of attention for the new generations of embedded systems.
Prassanna Sithambaram, Alberto Macii, Enrico Macii
ACM Great Lakes Symposium on VLSI3
2005 Energy-Efficient Color Approximation for Digital LCD Interfaces
abstract
The limited resolution capabilities of color displays, coupled with the limited perceptual resolution of the human eye have been exploited to reduce the actual number of colors that are simultaneously displayed. In this work, we propose a color approximation approach, orthogonal to traditional color simplification schemes done either in the frame buffer or in the LCD controller; that targets the reduction of the energy required for the transmission of color data on digital LCD interfaces. Our approach leverages the serial nature of the transmission so as to approximate RGB pixel values with suitable codes that minimize the number of transitions on the LCD bus, for a given tolerated image quality level. Application of this scheme on a set of images shows energy savings from 60% to 75%, depending on the image quality.
Andi Nourrachmat, Sabino Salerno, Enrico Macii, Massimo Poncino
ICCD3
2005 Frame Buffer Energy Optimization by Pixel Prediction
abstract
We propose a technique to reduce the energy consumption of the frame buffer memory, based on the spatial locality of images and display frames. Our scheme reduces energy by selectively avoiding reads from the frame buffer when identical adjacent pixels are detected. This is made possible by using an auxiliary memory that stores the locality information. The proposed architecture allows to dynamically update the locality information, and, unlike previous approaches, it works virtually independent of the size and position of the updates of the display frames. Experimental results evaluated on a set of typical graphical applications show a reduction of about 40% of frame buffer reads.
Kimish Patel, Enrico Macii, Massimo Poncino
ICCD2
2005 A scalable algorithm for RTL insertion of gated clocks based on ODCs computation
abstract
We propose a new algorithm for automatic clock-gating insertion applicable at the register transfer level (RTL). The basic rationale of our approach is to eliminate redundant computations performed by temporally unobservable blocks through aggressive exploitation of observability don't care (ODC) conditions. ODCs are efficiently detected from an RTL description by focusing only on data-path modules with easily detectable input unobservability conditions. ODCs are then propagated in the form of logic expressions toward the registers by backward traversal and levelization of the design. Finally, the logic expressions are mapped onto hardware to provide control signals to the clock-gating logic at a reduced cost in area and speed. The technique is characterized by fast processing time, high scalability to large designs, and tight user control on clock-gating overhead. Our approach is compatible with standard industrial design flows, and reduces power consumption significantly with a small overhead in delay and area. Experimental results obtained on a set of industrial RTL designs containing several tens of thousands of gates show average power reductions of around 42%. On the same examples, the application of traditional clock-gating leads to average savings reductions close to 29%.
Pietro Babighian, Luca Benini, Enrico Macii
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Automated DNA fragments recognition and sizing through AFM image processing
abstract
This paper presents an automated algorithm to determine DNA fragment size from atomic force microscope images and to extract the molecular profiles. The sizing of DNA fragments is a widely used procedure for investigating the physical properties of individual or protein-bound DNA molecules. Several atomic force microscope (AFM) real and computer-generated images were tested for different pixel and fragment sizes and for different background noises. The automated approach minimizes processing time with respect to manual and semi-automated DNA sizing. Moreover, the DNA molecule profile recognition can be used to perform further structural analysis. For computer-generated images, the root mean square error incurred by the automated algorithm in the length estimation is 0.6% for a 7.8 nm image pixel size and 0.34% for a 3.9 nm image pixel size. For AFM real images we obtain a distribution of lengths with a standard deviation of 2.3% of mean and a measured average length very close to the real one, with an error around 0.33%.
Elisa Ficarra, Luca Benini, Enrico Macii, Giampaolo Zuccheri
IEEE Trans. Inf. Technol. Biomed.3
2004 Techniques for Enhancing Computation of DNA Curvature Molecules
abstract
An automated algorithm is presented to determine the DNA molecule intrinsic curvature profiles and the molecular spatial orientations in atomic force microscope images. The curvature is composed by static and dynamic contributions. The former is the intrinsic curvature, a function of the DNA nucleotide sequence, while the latter is due to thermal fluctuations. This algorithm allows to reconstruct the intrinsic curvature profile excluding the thermal contribution. The automated algorithm reconstructs the intrinsic curvature profile with a mean square error of 3.812/spl middot/ 10/sup -4/ rads over a profile with a central peak value of 0.196 rads, and 6.1/spl middot/10/sup -3/ rads over a curvature profile with two symmetric peaks of about 0.08 rads. Moreover, it correctly detects the location of the peaks in the molecules with a deviation of about 1% of molecule length.
Daniele Masotti, Elisa Ficarra, Enrico Macii, Luca Benini
BIBE3
2004 A Scalable ODC-Based Algorithm for RTL Insertion of Gated Clocks
abstract
This paper describes a new automatic clock-gating extraction working at the RT-level. The key features of our approach are: (i) seamless merging with existing industrial design flows and commercial tools; (ii) high scalability to deal with large circuits; (iii) improved quality of results with respect to available commercial tools; (iv) smaller and well-controlled overhead in speed and area. Experimental results, on a set of industrial RTL designs, demonstrate the viability and practical impact of our approach.
Pietro Babighian, Luca Benini, Enrico Macii
DATE3
2004 Sizing and Characterization of Leakage-Control Cells for Layout-Aware Distributed Power-Gating
abstract
This paper proposes a methodology for sleep transistor sizing for usage in a novel, single-threshold leakage cut-off approach, where power gating cells are distributed row-by-row in a fully placed circuit. Sizing equations are obtained by performing SPICE simulations for a 130nm technology. Furthermore, the layout of a test case is considered and power and delay values are extracted in order to demonstrate the practical impact of our solution.
Pietro Babighian, Luca Benini, Enrico Macii
DATE3
2004 Block-Enabled Memory Macros: Design Space Exploration and Application-Specific Tuning
abstract
In this paper, we propose a combined solution that allows us to customize the architecture of internally partitioned SRAM macros according to the given application be executed. Energy savings with respect to monolithic memory configurations are above 40%, without access time violation.
Luca Benini, Alessandro Ivaldi, Alberto Macii, Enrico Macii
DATE4
2004 Synthesis of Partitioned Shared Memory Architectures for Energy-Efficient Multi-Processor SoC
abstract
Accesses to the shared memory in multi-processor systems-on-chip represent a significant performance bottleneck. Multi-port memories are a common solution to this problem, because they allow parallel accesses. However, they are not an energy-efficient solution. We propose an energy-efficient shared-memory architecture that can be used as a substitute for multi-port memories, which is based on an application-driven partitioning of the shared address space into a multi-bank architecture. Experiments on a set of parallel benchmarks show energy savings of about 56% with respect to a dual-port memory architecture, at a very limited performance penalty.
Kimish Patel, Enrico Macii, Massimo Poncino
DATE2
2004 Energy-efficient bus encoding for LCD displays
abstract
This paper presents a low-power bus encoding technique suitable for the digital interface to a Liquid Crystal Display (LCD). In particular, we focus on interfaces that are compliant to the Digital Visual Interface (DVI) standard, in which the three color channels are serially transmitted to achieve high bandwidth.The proposed technique exploits the well-know inter-pixel correlation that exists in typical images by serially transmitting an encoded representation of the difference between adjacent pixels. The encoding is based on the principle of clustering the 1's in the code towards either ends of the pixel data, in such a way that serial transmission of a code yields at most 1 transition per pixel.The application of the encoding to a series of standard images resulted in energy savings of around 60% on average, with respect to a plain transmission of 8-bit pixel data.
Alberto Bocca, Sabino Salerno, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3
2004 Reducing cache misses by application-specific re-configurable indexing
abstract
The predictability of memory access patterns in embedded systems can be successfully exploited to devise effective application-specific cache optimizations. In this work, we propose an improved indexing scheme for direct-mapped caches, which drastically reduces the number of conflict misses by using application-specific information; the scheme is based on the selection of a subset of the address bits. With respect to similar approaches, our solution has two main strengths. First, it models the misses analytically by building a miss equation, and exploits a symbolic algorithm to compute the exact optimum solution (i.e., the subset of address bits to be used as cache index that minimizes conflict misses). Second, we designed a re-configurable bit selector, which can be programmed at run-time to fit the optimal cache indexing to a given application. Results show an average reduction of conflict misses of 24%, measured over a set of standard benchmarks, and for different cache configurations.
Kimish Patel, Enrico Macii, Luca Benini, Massimo Poncino
ICCAD2
2004 Post-layout leakage power minimization based on distributed sleep transistor insertion
abstract
This paper introduces a new approach to sub-threshold leakage power reduction in CMOS circuits. Our technique is based on automatic insertion of sleep transistors for cutting sub-threshold current when CMOS gates are in stand-by mode. Area and speed overhead caused by sleep transistor insertion are tightly controlled thanks to: (i) a post-layout incremental modification step that inserts sleep transistors in an existing row-based layout; (ii) an innovative algorithm that selects the subset of cells that can be gated for maximal leakage power reduction, while meeting user-provided constraints on area and delay increase. The presented technique is highly effective and fully compatible with industrial back-end flows, as demonstrated by post-layout analysis on several benchmarks placed and routed with state-of-the art commercial tools for physical design.
Pietro Babighian, Luca Benini, Alberto Macii, Enrico Macii
ISLPED4
2004 Limited intra-word transition codes: an energy-efficient bus encoding for LCD display interfaces
abstract
We propose a class of low-power codes, called Limited Intra-Word Transition (LIWT) codes, suitable for the digital interface to Liquid Crystal Displays (LCD).The proposed technique exploits the existing inter-pixel correlation of typical images, by transmitting an encoded representation of the difference between adjacent pixels. Since all standard LCD transmission protocols are serial, the LIWT specifically targets the minimization of intra-word transitions.The application of the encoding to a series of standard images resulted in transitions savings over 60% on average with respect to two standard TMDS and LVDS protocols.
Sabino Salerno, Alberto Bocca, Enrico Macii, Massimo Poncino
ISLPED3
2004 Power-aware clock tree planning
abstract
Modern processors and SoCs require the adoption of power-oriented design styles, due to the implications that power consumption may have on reliability, cost and manufacturability of integrated circuits featuring nanometric technologies. And the power problem is further exacerbated by the increasing demand of devices for mobile, battery-operated systems, for which reduced power dissipation is mandatory. A large fraction of the power consumed by a synchronous circuit is due to the clock distribution network. This is for two reasons: First, the clock nets are long and heavily loaded. Second, they are subject to a high switching activity.The problem of automatically synthesizing a power efficient clock tree has been addressed recently in a few research contributions. In this paper, we introduce a methodology in which low-power clock trees are obtained through aggressive exploitation of the clock-gating technology. Distinguishing features of the methodology are: (i) The capability of calculating powerful clock-gating conditions that go beyond the simple topological search of the RTL source code. (ii) The capability of determining the clock tree logical structure starting from an RTL description. (iii) The capability of including in the cost function that drives the generation of the clock tree structure both functional (i.e., clock activation conditions) and physical (i.e., floorplanning) information. (iv) The capability of generating a clock tree structure that can be synthesized and routed using standard, commercially-available back-end tools.We illustrate the methodology for power-aware RTL clock tree planning, we provide details on the fundamental algorithms that support it and information on how such a methodology can be integrated into an industrial design flow. The results achieved on several benchmarks, as well as on a real design case demonstrate the feasibility and the potential of the proposed approach.
Monica Donno, Enrico Macii, Luca Mazzoni
ISPD2
2004 Memory energy minimization by data compression: algorithms, architectures and implementation
abstract
Storing data in compressed form is becoming common practice in high-performance systems, where memory bandwidth constitutes a serious bottleneck to program execution speed. In this paper, we suggest hardware-assisted data compression as a tool for reducing energy consumption of processor-based systems. We propose a novel and efficient architecture for on-the-fly data compression and decompression whose field of operation is the cache-to-memory path. Uncompressed cache lines are compressed before they are written back to main memory, and decompressed when cache refills take place. We explore two classes of table-based compression schemes. The first, based on offline data profiling, is particularly suitable to embedded systems, where predictability of the data set is usually higher than in general-purpose systems. The second solution we introduce is adaptive, that is, it takes decisions on whether data words should be compressed according to the data statistics of the program being executed. We describe in details the architecture of the compression/decompression unit and we provide an insight about its implementation as a hardware (HW) block. We present experimental results concerning memory traffic and energy consumption in the cache-to-memory path of a core-based system running standard benchmark programs. The obtained energy savings range from 8%-39% when profile-driven compression is adopted, and from 7%-26% when the adaptive scheme is used. Performance improvements are also achieved as a by-product, showing the practical applicability of the proposed approach.
Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii
IEEE Trans. Very Large Scale Integr. Syst.4
2003 Energy-aware design techniques for differential power analysis protection
abstract
Differential power analysis is a very effective cryptanalysis technique that extracts information on secret keys by monitoring instantaneous power consumption of cryptoprocessors. To protect against differential power analysis, power supply noise is added in cryptographic computations, at the price of an increase in power consumption. We present a novel technique, based on well-known power-reducing transformations coupled with randomized clock gating, that introduces a significant amount of scrambling in the power profile without increasing (and, in some cases, by even reducing) circuit power consumption.
Luca Benini, Alberto Macii, Enrico Macii, Elvira Omerbegovic, Fabrizio Pro, Massimo Poncino
DAC3
2003 Clock-tree power optimization based on RTL clock-gating
abstract
As power consumption of the clock tree in modern VLSI designs tends to dominate, measures must be taken to keep it under control. This paper introduces an approach for reducing clock power based on clock gating. We present a methodology that, starting from an RTL description, automatically generates a set of constraints for driving the construction of the clock tree by the clock synthesis tool. The methodology has been fully integrated into an industry-strength design flow, based on Synopsys DesignCompiler (front-end) and Cadence Silicon Ensemble (back end). The power savings achieved on some industrial examples show that, when the size of the circuits is significant, savings on the power consumption of the clock tree are up to 75% larger than those achieved by applying traditional clock gating at the clock inputs of the RTL modules of the designs.
Monica Donno, Alessandro Ivaldi, Luca Benini, Enrico Macii
DAC4
2003 A New Algorithm for Energy-Driven Data Compression in VLIW Embedded Processors
Alberto Macii, Enrico Macii, Fabrizio Crudo, Roberto Zafalon
DATE2
2003 Improving the Efficiency of Memory Partitioning by Address Clustering
Alberto Macii, Enrico Macii, Massimo Poncino
DATE2
2003 A novel architecture for power maskable arithmetic units
abstract
Power maskable units have been proposed as a viable solution for preventing side-channel attacks to cryptoprocessors. This paper presents a novel architecture for the implementation of a class of such kinds of units, namely arithmetic components, which find wide usage in cryptographic applications and which are not suitable to traditional masking techniques. Results of extensive exploration and architectural trade-off analysis show the viability of the proposed solution.
Luca Benini, Alberto Macii, Enrico Macii, Elvira Omerbegovic, Massimo Poncino, Fabrizio Pro
ACM Great Lakes Symposium on VLSI3
2003 Combining wire swapping and spacing for low-power deep-submicron buses
abstract
We propose an approach for reducing the energy consumption of address buses that targets both the switching and the crosstalk components of power dissipation.The method is based on the combined application of two techniques. First, selective wire swapping is applied in such a way that bus wires with high coupling activity are kept far away from each other. Then, the slack available in the floorplanning for the routing of the bus wire is exploited to realize a bus with non-uniform inter-wire spacing. Both swapping and placement are driven by the switching data obtained from the analysis of typical address bus traces, and can be successfully applied to any address bus.Results on a set of profiled address streams show the effectiveness of the proposed approach.
Enrico Macii, Massimo Poncino, Sabino Salerno
ACM Great Lakes Symposium on VLSI1
2003 Energy-efficient data scrambling on memory-processor interfaces
abstract
Crypto-processors are prone to security attacks based on the observation of their power consumption profile. We propose new techniques for increasing the non-determinism of such profile, which rely on the idea of introducing randomness in the bus data transfers. This is achieved by combining data scrambling with energy-efficient bus encoding, thus providing high information protection at no energy cost.Results on a set of bus traces originated by real-life applications demonstrate the applicability of the proposed solution.
Luca Benini, Angelo Galati, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED4
2003 Discharge Current Steering for Battery Lifetime Optimization
abstract
Portable and wearable computers can be powered by different combinations of two or more battery packs to give the user the possibility of choosing an optimal compromise between lifetime and weight/size. Recent work on battery-driven power management has demonstrated that sequential discharge is suboptimal in multibattery systems and lifetime can be maximized by distributing (steering) the current load on the available batteries, thereby discharging them in a partially concurrent fashion. Based on these observations, we formulate multibattery lifetime maximization as a continuous, constrained optimization problem, which can be efficiently solved by nonlinear optimizers. We show that significant lifetime extensions can be obtained with respect to standard sequential discharge (up to 160 percent), as well to previously proposed battery scheduling algorithms (up to 12 percent).
Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii, Massimo Poncino
IEEE Trans. Computers4
2003 Scheduling battery usage in mobile systems
abstract
The use of multibattery power supplies is becoming common practice in electronic appliances of the latest generations. Economical and manufacturing constraints are at the basis of this choice. Unfortunately, a partitioned battery subsystem is not able to deliver the same amount of charge as a monolithic battery with the same total capacity. In this paper, we define the concept of battery scheduling, we investigate several policies for solving the problem of optimal charge delivery, and we study the relationship of such policies with different configurations of the battery subsystem. Experimental results, obtained for different kinds of current workloads, demonstrate that the choice of the proper scheduling can make system lifetime as close as 1% of the theoretical upper bound, that is, a monolithic power supply of equal capacity.
Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
IEEE Trans. Very Large Scale Integr. Syst.3
2002 Hardware-Assisted Data Compression for Energy Minimization in Systems with Embedded Processors
abstract
In this paper, we suggest hardware-assisted data compression as a tool for reducing energy consumption of core-based embedded systems. We propose a novel and efficient architecture for on-the-fly data compression and decompression whose field of operation is the cache-to-memory path. Uncompressed cache lines are compressed before they are written back to main memory, and decompressed when cache refills take place. We explore two classes of compression methods, profile-driven and differential, since they are characterized by compact HW implementations, and we compare their performance to those provided by some state-of-the-art compression methods (e.g., we have considered a few variants of the Lempel-Ziv encoder). We present experimental results about memory traffic and energy consumption in the cache-to-memory path of a core-based system running standard benchmark programs. The achieved average energy savings range from 4.2% to 35.2%, depending on the selected compression algorithm.
Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii
DATE4
2002 Wire Placement for Crosstalk Energy Minimization in Address Buses
abstract
We propose a novel approach to bus energy minimization that targets crosstalk effects. Unlike previous approaches, we try to reduce energy through capacitance optimization, by adopting nonuniform spacing between wires. This allows reduction of power and at the same time takes into account signal integrity. Therefore, performance is not degraded. Results show that the method saves up to 30% of total bus energy at no cost in performance or complexity of the design (no encoding-decoding circuitry is needed), and limited cost in area.
Luca Macchiarulo, Enrico Macii, Massimo Poncino
DATE2
2002 Enhanced clustered voltage scaling for low power
abstract
This paper presents a voltage scaling approach that is based on an enhanced variant of clustered voltage scaling originally proposed by Usami and Horowitz ([1]) The results show that subtituting the original depth first strategy with a breadth first one results in improved speed and quality of results. Data are validated through power and timing analysis performed with a commercial tool.
Monica Donno, Luca Macchiarulo, Alberto Macii, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI4
2002 Discharge current steering for battery lifetime optimization
abstract
Recent work on battery-driven power management has demonstrated that sequential discharge is suboptimal in multi-battery systems, and lifetime can be maximized by distributing (steering) the current load on the available batteries, thereby discharging them in a partially concurrent fashion. Based on these observations, we formulate multi-battery lifetime maximization as a continuous, constrained optimization problem, which can be efficiently solved by non-linear optimizers. We show that great lifetime extensions can be obtained with respect to standard sequential discharge, as well to previously proposed battery allocation schemes.
Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED3
2002 Minimizing memory access energy in embedded systems by selective instruction compression
abstract
We propose a technique for reducing the energy spent in the memory-processor interface of an embedded system during the execution of firmware code. The method is based on the idea of compressing the most commonly executed instructions so as to reduce the energy dissipated during memory access. Instruction decompression is performed on-the-fly by a hardware block located between processor and memory: No changes to the processor architecture are required. Hence, our technique is well suited for systems employing IP cores whose internal architecture cannot be modified. We describe a number of decompression schemes and architectures that effectively trade off hardware complexity and static code size increase for memory energy and bandwidth reduction, as proved by the experimental data we have collected by executing several test programs on different design templates.
Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.3
2002 Guest editorial: low-power electronics and design
abstract
status: Published
Enrico Macii, Ingrid Verbauwhede
IEEE Trans. Very Large Scale Integr. Syst.1
2001 From Architecture to Layout: Partitioned Memory Synthesis for Embedded Systems-on-Chip
abstract
We propose an integrated front-end/back-end flow for the automatic generation of a multi-bank memory architecture for embedded systems. The flow is based on an algorithm for the automatic partitioning of on-chip SRAM. Starting from the dynamic execution profile of an embedded application running on a given processor core, we synthesize a multi-banked SRAM architecture optimally fitted to the execution profile.
Luca Benini, Luca Macchiarulo, Alberto Macii, Enrico Macii, Massimo Poncino
DAC4
2001 Extending lifetime of portable systems by battery scheduling
abstract
Multi-battery power supplies are becoming popular in electronic appliances of the latest generations, due to economical and manufacturing constraints. Unfortunately, a partitioned battery subsystem is not able to deliver the same amount of charge as a monolithic battery with the same total capacity. In this paper, we define the concept of battery scheduling, we investigate policies for solving the problem of optimal charge delivery, and we study the relationship of such policies with different configurations of the battery subsystem. Results, obtained for different workloads, demonstrate that the choice of the proper scheduling can make, in the best cease, system lifetime as close as 1% of that guaranteed by a monolithic battery of equal capacity.
Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
DATE4
2001 On-the-fly layout generation for PTL macrocells
abstract
Pass transistor logic (PTL) has been recently proposed as an alternative to standard MOS for aggressive circuit design. Even though PTL has been successful in a few hand-crafted designs, its acceptance into mainstream digital design critically depends on the availability of tools for logic and physical synthesis and optimization. The automatic synthesis of pass transistor circuits starting from BDDs has been intensively studied in the past with promising results, but back-end tools for PTL cell generation are still missing. We describe an automatic layout generator that has been designed for seamless integration in a library-free PTL design flow. The generator exploits the distinctive characteristics of pass transistor networks produced by synthesis to achieve quality of results comparable with state-of-the art commercial cell generation tools in a function of the execution time.
Luca Macchiarulo, Luca Benini, Enrico Macii
DATE3
2001 Low-energy for deep-submicron address buses
abstract
Article Low-energy for deep-submicron address buses Share on Authors: Luca Macchiarulo Politecnico di Torino, Torino, Italy Politecnico di Torino, Torino, ItalyView Profile , Enrico Macii Politecnico di Torino, Torino, Italy Politecnico di Torino, Torino, ItalyView Profile , Massimo Poncino Politecnico di Torino, Torino, Italy Politecnico di Torino, Torino, ItalyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 176–181https://doi.org/10.1145/383082.383127Online:06 August 2001Publication History 22citation275DownloadsMetricsTotal Citations22Total Downloads275Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Luca Macchiarulo, Enrico Macii, Massimo Poncino
ISLPED2
2001 Synthesis of power-managed sequential components based oncomputational kernel extraction
abstract
This paper introduces a power optimization paradigm for sequential components based on the concept of computational kernel, a highly simplified logic block whose behavior mimics the steady-state behavior of the original specification. We present a flexible framework that supports a number of algorithmic options for carrying out kernel extraction. We first describe an exact symbolic procedure that is applicable to components for which only a functional specification (i.e., the state transition graph) is available. Due to its computational complexity, this procedure is mainly of theoretical interest and it is not usable for large circuits. We then propose two approximate algorithms that can be adopted in practical situations. The first one is simulation-based and it is suitable to cases where input data streams representing typical operation of the component are available. The second approach performs kernel extraction by iteratively refining a structural representation of the component obtained through synthesis. The impact of the power optimization paradigm based on kernel extraction is demonstrated by the results of extensive experimentation carried out on a number of benchmarks of different characteristics and nature.
Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2001 Discrete-time battery models for system-level low-power design
abstract
For portable applications, long battery lifetime is the ultimate design goal. Therefore, the availability of battery and voltage converter models providing accurate estimates of battery lifetime is key for system-level low-power design frameworks. In this paper, we introduce a discrete-time model for the complete power supply subsystem that closely approximates the behavior of its circuit-level continuous-time counterpart. The model is abstract and efficient enough to enable event-driven simulation of digital systems described at a very high level of abstraction and that includes, among their components, also the power supply. The model gives the designer the possibility of estimating battery lifetime during system-level design exploration, as shown by the results we have collected on meaningful case studies. In addition, it is flexible and it can thus be employed for different battery chemistries.
Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
IEEE Trans. Very Large Scale Integr. Syst.4
2001 Parameterized RTL power models for soft macros
abstract
We propose a new power macromodel for usage in the context of register-transfer level (RTL) power estimation. The model is suitable for reconfigurable, synthesizable, soft macros because it is parameterized with respect to the input data size (i.e., bit width) and can also be automatically scaled with respect to different technology libraries and/or synthesis options. The power model is precharacterized once and for all for each soft macro and then adapted to each specific instance by means of a single additional experiment to be performed by the end user. No intellectual-property disclosure is required for model scaling. The proposed model is derived from empirical analysis of the sensitivity of power consumption on input statistics, input data size, and technology. The experiments prove that with limited approximation, it is possible to decouple the effects on power of these three factors. The proposed solution is innovative since no previous macromodel supports automatic technology scaling and yields average estimation errors around 10%.
Alessandro Bogliolo, Roberto Corgnati, Enrico Macii, Massimo Poncino
IEEE Trans. Very Large Scale Integr. Syst.3
2001 Estimation of lower and upper bounds on the power consumption from scheduled data flow graphs
abstract
In this paper, we present an approach for the calculation of lower and upper bounds on the power consumption of data path resources like functional units, registers, I/O ports, and busses from scheduled data flow graphs executing a specified input data stream. The low power allocation and binding problem is formulated. First, it is shown that this problem without constraining the number of resources can be relaxed to the bipartite weighted matching problem which is solvable in O(n)/sup 3/. n is the number of arithmetic operations, variables, I/O-access or bus-access operations which have to be bound to data path resources. In a second step we demonstrate that the relaxation can be efficiently extended by including Lagrange multipliers in the problem formulation to handle a resource constraint. The estimated bounds take into account the effects of resource sharing. The technique can be used, for example, to prune the design space in high-level synthesis for low power before the allocation and binding of the resources. The application of the technique on benchmarks with real application input data shows the tightness of the bounds.
Lars Kruse, Eike Schmidt, Gerd von Cölln, Ansgar Stammermann, Arne Schulz, Enrico Macii, Wolfgang Nebel
IEEE Trans. Very Large Scale Integr. Syst.6
2001 Stream synthesis for efficient power simulation based on spectral transforms
abstract
One way of minimizing the time required to perform simulation-based power estimation is that of reducing the length of the input trace to be fed to the simulator. Obviously, the use of a reduced stream may introduce some errors in the estimation results. The generation (or synthesis) of the short input sequence to be used for power simulation should then be carried out in such a way that the resulting error is minimized. Existing techniques exploit the knowledge of some statistical and correlation characteristics concerning the original input trace to generate a reduced stream that closely matches such characteristics. In this paper, we introduce a new stream synthesis method. Its distinguishing feature is the use of spectral analysis based on the discrete Fourier transform to determine a reduced sequence of vectors that enables us to shorten the overall power simulation time at a very limited penalty in accuracy. The effectiveness and the robustness, in terms of estimation accuracy, of the proposed synthesis procedure are demonstrated by the experimental results we have obtained on standard combinational benchmarks for a variety of input streams with different statistical and correlation properties. Data for sequential circuits are also reported.
Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
IEEE Trans. Very Large Scale Integr. Syst.2
2000 Synthesis of application-specific memories for power optimization in embedded systems
abstract
This paper presents a novel approach to memory power optimization for embedded systems based on the exploitation of data locality. Locations with highest access frequency are mapped onto a small, low-power application-specific memory which is placed close the processor. Although, in principle, a cache may be used to implement such a memory, more efficient solutions may be adopted. We propose an architecture that outperforms (power-wise) different types of cache memories at no penalty in performance. Power savings (averaged over a number of embedded applications running on ARM processors) range from 12% to 68%.
Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
DAC3
2000 A Discrete-Time Battery Model for High-Level Power Estimation
abstract
In this paper, we introduce a discrete-time model for the complete power supply sub-system that closely approximates the behavior of its circuit-level (i.e., HSpice), continuous-time counterpart. The model is abstract and efficient enough to enable event-driven simulation of digital systems described at a very high level of abstraction and that include, among their components, also the power supply. Therefore, it can be successfully used for the purpose of battery life-time estimation during design optimization, as shown by the results we have collected on a meaningful case study. Experiments prove also that the accuracy of our model is very close to that provided by the corresponding Spice-level model.
Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
DATE4
2000 Regression-based RTL power models for controllers
abstract
Power consumption of the controller of an RTL design, though typically smaller than that of the datapath, cannot be neglected. Controller power modeling is a more challenging task than that of datapath modules, because the controller is usually specified in an abstract fashion that has no structural relationship with its final implementation. This paper presents an extensive experimental study on pre-state assignment power models. Starting from a large set of controllers, we study the relation between power and static parameters, such as number of inputs and outputs, as well as dynamic parameters, such as input and output switching activity. We then quantify the loss of accuracy caused by the high level of abstraction of the models, and we show how previously published results, obtained under favorable experimental conditions, have underestimated model inaccuracy. Finally, we perform extensive exploration on a large number of alternative macro-model structures, and we select an optimal model equation.
Luca Benini, Alessandro Bogliolo, Enrico Macii, Massimo Poncino, Mihai Surmei
ACM Great Lakes Symposium on VLSI3
2000 Supporting system-level power exploration for DSP applications
abstract
System-level power exploration requires tools for estimation of the overall power consumed by a system, as well as a detailed breakdown of the consumption of its main functional blocks. We focus on power estimation for data-dominated systems specified as synchronous data-flows and implemented on a single-processor architecture. Our estimator is integrated within the Ptolemy design environment, and provides information to system designers on the power dissipated by every task in a given specification. Power estimation is based on instruction-level power models. We demonstrate the applicability of our tool on a few design examples and target architectures.
Luca Benini, Marco Ferrero, Alberto Macii, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI4
2000 A multilevel engine for fast power simulation of realistic inputstreams
abstract
Power estimation for validation and sign-off is a critical step in the design process. In this phase, accuracy is a key requirement, but there are hard constraints on the time that can be dedicated to power estimation. Moreover, it is important to estimate the power dissipated by the system while running typical applications, i.e., extremely long streams of validation patterns provided by the designer. The power dissipated by digital systems under realistic input stimuli is not accurately described by a single average value, but by a waveform that shows how power consumption varies over time as the system responds to the inputs. In this paper, we face the problem of obtaining accurate power waveforms for combinational and sequential circuits under typical usage patterns. We propose a multilevel simulation engine that achieves high accuracy in estimating the time-domain power waveform, as well as the average power with high computational efficiency.
Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Riccardo Scarsi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 Symbolic optimization of interacting controllers based onredundancy identification and removal
abstract
This paper presents a binary decision diagram (BDD)-based algorithm for the optimization of the driven machine, M/sub 2/, of a finite-state machine (FSM) network with cascade connection, M/sub 1//spl rarr/M/sub 2/. The technique we propose relies on redundant faults identification and removal. A fault, f, located into machine M/sub 2/, is redundant with respect to the overall network if the driving machine M/sub 1/ is not able to generate any test sequence for such a fault. When the state transition graph (STG) specifications of the network components are available, the standard way for checking the redundancy condition for the considered fault requires one to first construct the product machine M/sub 2//spl times/M/sub 2//sup F/, where M/sub 2//sup F/ is the faulty FSM, then to connect it to the driving machine, and finally to perform reachability analysis on the composed machine M/sub 1//spl rarr/M/sub 2//spl times/M/sub 2//sup F/. Clearly, the size of such machine limits the applicability of the approach above to systems whose components have a few tens of states at most, even when symbolic traversal algorithms are used. Since we are interested in dealing with networks of larger FSM's (i.e., machines whose STGs can not be represented explicitly), we propose to use the product automaton P'=A/sub 1//spl times/A/sub f/, where A/sub 1/' is the finite automaton (FA) accepting all the output sequences of M/sub 1/, and A/sub f/ is the FA accepting all the test sequences for fault f, instead of machine M/sub 1//spl rarr/M/sub 2//spl times/M/sub 2//sup F/. This simplifies sensibly the task of the reachability analysis program, since A/sub f/ has considerably less states and less edges than the product machine M/sub 2//spl times/M/sub 2//sup F/ and, thus, the size of the BDD representation of its transition relation is much more easily manageable. In addition, differently from other approaches, automaton A/sub 1/' is not required to be deterministic and state minimal. This allows us to avoid the application of determinization and state minimization procedures whose complexity is exponential. We present experimental results For examples (i.e., network of interacting controllers) on which existing optimization methods are not applicable, due to the size of the component FSM's. We also provide a comparison to the data produced by state-of-the-art FSM network optimizers on small benchmarks in order to show the effectiveness of our approach.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 Power optimization of technology-dependent circuits based on symbolic computation of logic implications
abstract
This paper presents a novel approach to the problem of optimizing combinational circuits for low power. The method is inspired by the fact that power analysis performed on a technology mapped network gives more realistic estimates than it would at the technology-independent level. After each node's switching activity in the circuit is determined, high-power nodes are eliminated through redundancy addition and removal. To do so, the nodes are sorted according to their switching activity, they are considered one at a time, and learning is used to identify direct and indirect logic implications inside the network. These logic implications are exploited to add gates and connections to the circuit; this may help in eliminating high-power dissipating nodes, thus reducing the total switching activity and power dissipation of the entire circuit. The process is iterative; each iteration starts with a different target node. The end result is a circuit with a decreased switching power. Besides the general optimization algorithm, we propose a new BDD-based method for computing satisfiability and observability implications in a logic network; futhermore, we present heuristic techniques to add and remove redundancy at the technology-dependent level, that is, restructure the logic in selected places without destroying the topology of the mapped circuit. Experimental results show the effectiveness of the proposed technique. On average, power is reduced by 34%, and up to a 64% reduction of power is possible, with a negligible increase in the circuit delay.
R. Iris Bahar, Ernest T. Lampe, Enrico Macii
ACM Trans. Design Autom. Electr. Syst.3
2000 Glitch power minimization by selective gate freezing
abstract
This paper presents a technique for glitch power minimization in combinational circuits. The total number of glitches is reduced by replacing some existing gates with functionally equivalent ones (called F-Gates) that can be "frozen" by asserting a control signal. A frozen gate cannot propagate glitches to its output. Algorithms for gate selection and clustering that maximize the percentage of filtered glitches and reduce the overhead for generating the control signals are introduced. A power-efficient CMOS implementation of F-Gates is also described. An important feature of the proposed method is that it can be applied in place directly to layout-level descriptions; therefore, it guarantees very predictable results and minimizes the impact of the transformation on circuit size and speed.
Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
IEEE Trans. Very Large Scale Integr. Syst.4
1999 Kernel-Based Power Optimization of RTL Components: Exact and Approximate Extraction Algorithms
abstract
Article Free Access Share on Kernel-based power optimization of RTL components: exact and approximate extraction algorithms Authors: L. Benini Università di Bologna, Bologna, Italy 40136 Università di Bologna, Bologna, Italy 40136View Profile , G. De Micheli Stanford University, Stanford, CA Stanford University, Stanford, CAView Profile , E. Macii Politecnico di Torino, Torino, Italay 10129 Politecnico di Torino, Torino, Italay 10129View Profile , G. Odasso Politecnico di Torino, Torino, Italy 10129 Politecnico di Torino, Torino, Italy 10129View Profile , M. Poncino Politecnico di Torino, Torino, Italy 10129 Politecnico di Torino, Torino, Italy 10129View Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 247–252https://doi.org/10.1145/309847.309922Published:01 June 1999Publication History 0citation244DownloadsMetricsTotal Citations0Total Downloads244Last 12 Months6Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Luca Benini, Giovanni De Micheli, Enrico Macii, Giuseppe Odasso, Massimo Poncino
DAC3
1999 Synthesis of Low-Overhead Interfaces for Power-Efficient Communication over Wide Buses
abstract
In this paper we present algorithms for the synthesis of encoding and decoding interface l o gic that minimizes the average number of transitions on heavily-loaded global bus lines.The approach automatically constructs low-transition activity codes and hardware implementation of encoders and decoders, given information on word-level statistics.We present an accurate method that is applicable to low-width buses, as well as approximate methods that scale well with bus width.Furthermore, we introduce an adaptive architecture that automatically adjusts encoding to reduce t r ansition activity on buses whose word-level statistics are not known a-priori.Experimental results demonstrate that our approach well outperforms low-power encoding schemes presented in the past.
Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
DAC3
1999 Glitch Power Minimization by Gate Freezing
abstract
This paper presents a technique for glitch power minimization in combinational circuits. The total number of glitches is reduced by replacing some existing gates with functionally equivalent ones (called F-gates) that can be "frozen" by asserting a control signal. A frozen gate cannot propagate glitches to its output. An important feature of the proposed method is that it can be applied in-place directly to layout-level descriptions; therefore, it guarantees very predictable results and minimizes the impact of the transformation on circuit size and speed.
Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
DATE4
1999 Clustered Table-Based Macromodels for RTL Power Estimation
abstract
Macromodeling is considered the most effective approach to RTL power estimation. Among the macromodels presented in the literature, table-based ones have overcome some of the limitations of conventional, equation-based solutions. In this paper we propose some enhancements to the basic implementation of table-based macromodels that improve the estimation accuracy while preserving the intrinsic robustness.
Roberto Corgnati, Enrico Macii, Massimo Poncino
Great Lakes Symposium on VLSI2
1999 Regression-Based Macromodeling for Delay Estimation of Behavioral Components
abstract
This paper presents a methodology for delay estimation of hardware components described at the behavioral-level. The basis of the proposed technique is a well-known theoretical result that relates the entropy of a logic function to the delay of a multi-level implementation of the same function. We propose an improved model for delay estimation, and we prove its validity by means of experiments performed on a set of standard benchmarks.
Alberto Macii, Enrico Macii, Giuseppe Odasso, Massimo Poncino, Riccardo Scarsi
Great Lakes Symposium on VLSI2
1999 Parameterized RTL power models for combinational soft macros
abstract
We propose a new RTL power macromodel that is suitable for re-configurable, synthesizable soft-macros. The model is parameterized with respect to the input data size (i.e., bit-width), and can be automatically scaled with respect to different technology libraries and/or synthesis options. Scalability is obtained through a single additional characterization run, and does not require the disclosure of any intellectual property. The model is derived from empirical analysis of the sensitivity of power on input statistics, input data size and technology. The experiments prove that, with limited approximation, it is possible to de-couple the effects on power of these three factors. The proposed solution is innovative, since no previous macromodel supports automatic technology scaling, and yields estimation errors within 15%.
Alessandro Bogliolo, Roberto Corgnati, Enrico Macii, Massimo Poncino
ICCAD3
1999 Selective instruction compression for memory energy reduction in embedded systems
abstract
We propose a technique for reducing the energy required by firmware code to ezecute on embedded systema. The method ia based on the idea of compressing the moat commonly ezecuted instructions 80 a ~ to reduce the energy dissipated in memory acceeees. Instruction decompression is performed on the fly by a hardware module located between processor and memory: No changea to the processor architecture ore required. Hence, our technique is well-suited for systema employing IP cone whose internal architecture cannot be modified. We describe a number of decompnaaion achemea and architec-tura that effectively trade off hardware complezity for memory energy and bandwidth nduction, aa proved by experimental data collected by executing aeveml sample programs. 1
Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino
ISLPED3
1999 Automatic Synthesis of Large Telescopic Units Based on Near-Minimum Timed Supersetting
abstract
In high-performance systems, variable-latency units are often employed to improve the average throughput when the worst-case delay exceeds the cycle time. Traditionally, units of this type have been hand-designed. In this paper, we propose a technique for the automatic synthesis of variable-latency units that is applicable to large data-path modules. We define and study an optimization problem, timed supersetting, whose solution is at the kernel of the procedure for automatic generation of variable-latency units. We contribute a new algorithm for solving timed supersetting in the most difficult case, that is, when the timing behavior of the circuit is expressed through an accurate delay model. The proposed solution overcomes the computational limitations of previous approaches and its robustness is experimentally demonstrated by obtaining high-throughput, variable-latency implementations for all the largest circuits in the Iscas '85 and Iscas '89 benchmark suites, as well as for some realistic, high-performance arithmetic units.
Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino
IEEE Trans. Computers4
1999 Symbolic synthesis of clock-gating logic for power optimization of synchronous controllers
abstract
Recent results have shown that dynamic power management is effective in reducing the total power consumption of sequential circuits. In this paper, we propose a bottom-up approach for the automatic extraction and synthesis of dynamic power management circuitry starting from structural logic-level specifications. Our techniques leverage the compact BDD-based representation of Boolean and pseudo-Boolean functions to detect idle conditions where the clock can be stopped without compromising functional correctness. Moreover, symbolic techniques allow accurate probabilistic computations; in particular, they enable the use of non-equiprobable primary input distributions, a key step in the construction of models that match the behavior of real hardware devices with a high degree of fidelity. The results are encouraging, since power savings of up to 34% have been obtained on standard benchmark circuits.
Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Riccardo Scarsi
ACM Trans. Design Autom. Electr. Syst.3
1998 Computational Kernels and their Application to Sequential Power Optimization
abstract
We introduce a new sequential optimization paradigm based on the extraction of computational kernels, i.e., logic blocks whose behavior mimics the steady-state behavior of the original circuit. We present a procedure for the automatic extraction of such kernels directly from the gate-level description of the design. The advantage of this solution with respect to extraction algorithms based on STG analysis is that it can be applied to large circuits, since it does not require to manipulate the STG specification.
Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino
DAC4
1998 In-Place Power Optimization for LUT-Based FPGAs
abstract
This paper presents a new technique to perform power-oriented re-configuration of a system implemented using LUT FPGAs. The main features of our approach are: Accurate exploitation of degrees of freedom, concurrent optimization of multiple LUTs based on Boolean relations, and in-place re-programming without re-routing. Our tool optimizes the combinational component of the CLBs after layout, and does not require any re-wiring. Hence, delay and CLB usage are left unchanged, while power is minimized. As the algorithm operates locally on the various LUT clusters, it best performs on large examples as demonstrated by our experimental results: An average power reduction of 20.6% has been obtained on standard benchmarks.
Balakrishna Kumthekar, Luca Benini, Enrico Macii, Fabio Somenzi
DAC3
1998 Address Bus Encoding Techniques for System-Level Power Optimization
abstract
The power dissipated by system-level buses is the largest contribution to the global power of complex VLSI circuits. Therefore, the minimization of the switching activity at the I/O interfaces can provide significant savings on the overall power budget. This paper presents innovative encoding techniques suitable for minimizing the switching activity of system-level address buses. In particular, the schemes illustrated here target the reduction of the average number of bus line transitions per clock cycle. Experimental results, conducted on address streams generated by a real microprocessor, have demonstrated the effectiveness of the proposed methods.
Luca Benini, Giovanni De Micheli, Donatella Sciuto, Enrico Macii, Cristina Silvano
DATE4
1998 Power Estimation of Behavioral Descriptions
abstract
This paper presents a methodology for power estimation of designs described at the behavioral-level as the interconnection of functional modules. The input/output behavior of each module is implicitly stored using BDDs, and the power consumed by the network is estimated using a novel and accurate entropy-based approach. As a demonstration example, we have used the proposed power estimation technique to evaluate and compare the effects of some architectural transformations applied to a reference design specification on the power dissipation of the corresponding implementations.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino
DATE3
1998 Timed Supersetting and the Synthesis of Telescopic Units
abstract
In high-performance systems, variable-latency units are often employed to improve the average throughput when the worst-case delay exceeds the cycle time. Although such units have traditionally been hand-designed, recent results have shown that variable-latency units can be automatically generated. Unfortunately, the existing synthesis procedure has limited applicability due to its computational complexity. In this work, we define and study an optimization problem, timed supersetting, whose solution is at the kernel of the procedure for automatic generation of variable-latency units. We contribute a new algorithm for solving timed supersetting in the most difficult case, that is, when the timing behaviour of the circuits is expressed through an accurate delay model. The proposed solution overcomes the complexity limitation of previous approaches, and its robustness is experimentally demonstrated by obtaining high-throughput, variable-latency implementations for all the largest circuits in the Iscas'85 and Iscas'89 benchmark suites.
Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino
Great Lakes Symposium on VLSI4
1998 Reducing Power Consumption of Dedicated Processors Through Instruction Set Encoding
abstract
With the increased clock frequency of modern, high-performance processors (over 500 MHz, in some cases), limiting the power dissipation has become the most stringent design target. It is thus mandatory for processor engineers to resort to a large variety of optimization techniques to reduce the power requirements in the hot zones of the chip. In this paper, we focus on the power dissipated By the instruction fetch and decode logic, a portion of the processor architecture where a lot of capacitance switching normally takes place. We propose a methodology for determining an encoding of the instruction set that guarantees the minimization of the number of bit transitions occurring inside the registers of the pipeline stages involved in instruction fetching and decoding. The assignment of the binary patterns to the op-codes is driven by the statistics concerning instruction adjacency collected through instruction-level simulation of typical software applications; therefore, the technique is best exploited when applied to encode the instruction set of core processors and microcontrollers, since components of these types ore commonly used to execute fixed portions of machine code within embedded systems. We illustrate the effectiveness of the methodology through the experimental data we have obtained on an existing microprocessor.
Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino
Great Lakes Symposium on VLSI4
1998 Symbolic algorithms for layout-oriented synthesis of pass transistor logic circuits
abstract
This paper presents a nouel methodology for synthesizing PTL circuifg, whose disfincfiue feafures are the use of a symbolic algorifhm for the covem.ng of fhe initial network in ferms of PTL cells, and the eqloifation of layout-level ama and delay models dun.ng fhe selection of fhe be~t couem.ngsolufion.The results produced by the synthesis procedure on the full guite of fhe kcas'85 combinational circuifs are very encouraging.
Fabrizio Ferrandi, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi, Fabio Somenzi
ICCAD3
1998 Stream synthesis for efficient power simulation based on spectral transforms
abstract
In this paper, we present a power estimation technique for control-flow intensive designs that is tailored towards driving iterative high-level synthesis systems, where hundreds of architectural trade-offs are explored and compared. Our method is fast and relatively accurate. The algorithm utilizes the behavioral information to extract branch probabilities, and uses these in conjunction with switching activity and circuit capacitance information, to estimate the power consumption of a given architecture.
Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi
ISLPED2
1998 Telescopic units: a new paradigm for performance optimization of VLSI designs
abstract
This paper introduces a novel optimization paradigm for increasing the throughput of digital systems. The basic idea consists of transforming fixed-latency units into variable-latency ones that run with a faster clock cycle. The transformation is fully automatic and can be used in conjunction with traditional design techniques to improve the overall performance of speed-critical units. In addition, we introduce procedures for reducing the area overhead of the modified units, and we formulate an algorithm for automatically restructuring the controllers of the data paths in which variable-latency units have been introduced. Results, obtained on a large set of benchmark circuits, show an average throughput improvement exceeding 27%, at the price of a modest area increase (less than 8% on average).
Luca Benini, Enrico Macii, Massimo Poncino, Giovanni De Micheli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 High-level power modeling, estimation, and optimization
abstract
Silicon area, performance, and testability have been, so far, the major design constraints to be met during the development of digital very-large-scale-integration (VLSI) systems. In recent years, however, things have changed; increasingly, power has been given weight comparable to the other design parameters. This is primarily due to the remarkable success of personal computing devices and wireless communication systems, which demand high-speed computations with low power consumption. In addition, there exists a strong pressure for manufacturers of high-end products to keep power under control, due to the increased costs of packaging and cooling this type of device. Last, the need of ensuring high circuit reliability has turned out to be more stringent. The availability of tools for the automatic design of low-power VLSI systems has thus become necessary. More specifically, following a natural trend, the interests of the researchers have lately shifted to the investigation of power modeling, estimation, synthesis, and optimization techniques that account for power dissipation during the early stages of the design flow. This paper surveys representative contributions to this area that have appeared in the recent literature.
Enrico Macii, Massoud Pedram, Fabio Somenzi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1998 Power optimization of core-based systems by address bus encoding
abstract
This paper presents a solution to the problem of reducing the power dissipated by a digital system containing an intellectual proprietary core processor which repeatedly executes a special-purpose program. The proposed method relies on a novel, application-dependent low-power address bus encoding scheme. The analysis of the execution traces of a given program allows an accurate computation of the correlations that may exist between blocks of bits in consecutive patterns; this information can be successfully exploited to determine an encoding which sensibly reduces the bus transition activity. Experimental results, obtained on a set of special-purpose applications, are very satisfactory; reductions of the bus activity up to 64.8% (41.8% on average) have been achieved over the original address streams. In addition, data concerning the quality and the performance of the automatically synthesized encoding/decoding circuits, as well as the results obtained for a realistic core-based design, indicate the practical usefulness of the proposed power optimization strategy.
Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Stefano Quer
IEEE Trans. Very Large Scale Integr. Syst.3
1997 Telescopic Units: Increasing the Average Throughput of Pipelined Designs by Adaptive Latency Control
abstract
This paper presents a technique, alternative to performance-drivensynthesis, that allows to drastically increase the averagethroughput of combinational logic blocks by transforming fixed-latencyunits into variable-latency ones that run with a fasterclock cycle.The transformation is fully automatic and can beused in conjunction with traditional design techniques, such aspipelining, to improve the overall performance of speed-criticalsystems.Results, obtained on a large set of benchmark circuits,are very promising.
Luca Benini, Enrico Macii, Massimo Poncino
DAC2
1997 High-Level Power Modeling, Estimation, and Optimization
abstract
In the past, the major concern of the VLSI designers werearea, performance, cost, and reliability.In recent years,however, this has changed and, increasingly, power is beinggiven comparable weight to area and speed.This is mainlydue to the remarkable success of personal computing devicesand wireless communication systems, which demandhigh-speed computation and complex functionality with lowpower consumption.In addition, there exists a strong pressurefor manufacturers of high-end products to keep powerunder control.The main driving factors for lower powerdissipation in these products are the costs associated withpackaging and cooling, and circuit reliability.Tools for the automatic design of low-power VLSI systemshave thus become mandatory.More specifically, followinga natural trend, interests of researchers have latelyshifted to the investigation of high-level power modeling,estimation, synthesis, and optimization techniques that accountfor power dissipation as the primary cost factor.This paper provides a non-exhaustive survey of the mostsuccessful and innovative ideas in this area that have appearedin the literature in the last few years.
Enrico Macii, Massoud Pedram, Fabio Somenzi
DAC1
1997 Asymptotic Zero-Transition Activity Encoding for Address Busses in Low-Power Microprocessor-Based Systems
abstract
In microprocessor-based systems, large power savings can be achieved through reduction of the transition activity of the on- and off-chip buses. This is because the total capacitance being switched when a voltage change occurs on a bus line is usually sensibly larger than the capacitive load that must be charged/discharged when internal nodes toggle. In this paper, we propose an encoding scheme which is suitable for reducing the switching activity on the lines of an address bus. The technique relies on the observation that, in a remarkable number of cases, patterns traveling onto address buses are consecutive. Under this condition it may therefore be possible, for the devices located at the receiving end of the bus, to automatically calculate the address to be received at the next clock cycle; consequently, the transmission of the new pattern can be avoided, resulting in an overall switching activity decrease. We present analytical and experimental analyses showing the improved performance of our encoding scheme when compared to both binary and Gray addressing schemes, the latter being widely accepted as the most efficient method for address bus encoding. We also propose power and timing efficient implementations of the encoding and the decoding logic, and we discuss the applicability of the technique to real microprocessor-based designs.
Luca Benini, Giovanni De Micheli, Enrico Macii, Donatella Sciuto, Cristina Silvano
Great Lakes Symposium on VLSI3
1997 Accurate Entropy Calculation for Large Logic Circuits Based on Output Clustering
abstract
Entropy-based estimation is a promising approach to the problem of predicting the power dissipated by a digital system for which an architectural description is available. For achieving good performance of the power estimation tool, an accurate computation of the input and output entropies of the Boolean functions implemented by the circuit is essential. For small designs, the calculation can be carried out exactly, thanks to the compact representation and ease of manipulation of Boolean and pseudo-Boolean functions provided by BDD-like data structures. For large circuits, on the other hand, resorting to approximate computations is mandatory. Techniques to determine an upper bound on the exact entropy values have been developed in the recent past. Unfortunately, the results provided by such techniques are, in some ceases, not satisfactory; in other words, the assumptions made to simplify the calculation-total absence of correlation among the output signals of a circuit are in many cases too strong to guarantee a reasonable lightness of the approximate entropy values to the exact ones. In this paper, we propose a method to determine the entropy of large logic circuits with a level of accuracy which is far beyond the one provided by existing approaches. We partition the set of output signals according to the information about the functional correlations that may exist among such signals, and we compute the approximate entropy values after performing output clustering. Experimental results, obtained on a large collection of benchmarks, are very promising.
Antonio Lioy, Enrico Macii, Massimo Poncino, Massimo Rossello
Great Lakes Symposium on VLSI2
1997 Fast power estimation for deterministic input streams
abstract
The power dissipated by digital systems under realistic input stimuli is not accurately described by a single average value, but by a waveform that shows how power consumption varies over time as the system responds to the inputs. We face the problem of obtaining accurate power waveforms for combinational and sequential circuits under typical usage patterns. We propose a multi level simulation engine that achieves high accuracy in estimating the average power as well as the time domain power waveform with high computational efficiency.
Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Riccardo Scarsi
ICCAD3
1997 System-level power optimization of special purpose applications: the beach solution
abstract
Article Free Access Share on System-level power optimization of special purpose applications: the beach solution Authors: Luca Benini Stanford University, Computer Systems Laboratory, Stanford, CA Stanford University, Computer Systems Laboratory, Stanford, CAView Profile , Giovanni De Micheli Stanford University, Computer Systems Laboratory, Stanford, CA Stanford University, Computer Systems Laboratory, Stanford, CAView Profile , Enrico Macii Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129 Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129View Profile , Massimo Poncino Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129 Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129View Profile , Stefano Quer Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129 Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129View Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 24–29https://doi.org/10.1145/263272.263277Published:01 August 1997Publication History 33citation259DownloadsMetricsTotal Citations33Total Downloads259Last 12 Months17Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Stefano Quer
ISLPED3
1997 Algebraic Decision Diagrams and Their Applications
R. Iris Bahar, Erica A. Frohm, Charles M. Gaona, Gary D. Hachtel, Enrico Macii, Abelardo Pardo, Fabio Somenzi
Formal Methods Syst. Des.5
1997 Symbolic timing analysis and resynthesis for low power of combinational circuits containing false paths
abstract
This paper presents applications of algebraic decision diagrams (ADDs) to timing analysis and resynthesis for low power of combinational CMOS circuits. We first propose a symbolic algorithm to perform true delay calculation of a technology mapped network; the procedure we propose, implemented as an extension of the SIS synthesis system, is able to provide more accurate timing information than any other method presented so far; in particular, it is able to compute and store the arrival times of all the gates of the circuit for all possible input vectors, as opposed to the traditional methods which consider only the worst case primary inputs combination. Furthermore, the approach does not require any explicit false path elimination. We then extend our timing analysis tool to the symbolic calculation of required times and slacks, and we use this information to perform resynthesis for low power of the circuit by gate resizing. Our approach takes into account false paths naturally; in fact, it guarantees that resizing of the gates does not increase the true delay of the circuit, even in the presence of false paths. Our experiments have shown that many circuits, originally free of false paths, exhibit a large number of these false paths when optimized for area; therefore, the ability to deal with circuits containing false paths is of primary importance. We present experimental results for ADD-based and static timing analysis-based resynthesis, which clearly show that our tool is superior in the case of circuits containing false paths, but at the same time, it provides competitive results in the case of circuits which are free of false paths.
R. Iris Bahar, Gary D. Hachtel, Enrico Macii, Fabio Somenzi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1997 Formal verification of digital systems by automatic reduction of data paths
abstract
Verification of properties (tasks) on a system P containing data paths may require too many resources (memory space and/or computation time) because such systems have very large and deep state spaces. As pointed out by Kurshan, what is needed is a reduced system P' which behaves exactly as P with respect to the properties that must be proved, but more compact than P, so that the verification can be easily performed. The process of finding P' from P is called reduction. P is specified by a network of interacting finite-state machines for data paths and controllers, and tasks are specified by finite-state automate. The verification of a task T on P is performed by the language containment check L(P)/spl sube/L(T), where L(P) is the language generated by P and L(T) is the language accepted by T. It has been shown that, under appropriate conditions, the system P can be reduced to P' and the task T to T' such that L(P')/spl sube/L(T')/spl hArr/L(P)/spl sube/L(T). The direct language containment check L(P)/spl sube/L(T) is no longer needed; it is replaced by L(P')/spl sube/L(T'), which is less expensive. More specifically, for the purpose of simplifying the verification of some properties, the system implementation is abstracted locally with respect to the behavior under observation (i.e., bottom-up reduction), in the context of an integrated top-down design/verification technique. The tasks that one may want to verify can express both safety and fairness constraints. In this paper, we prove that the reduction of some data paths to four-state, nondeterministic finite-state machines, and the redundancy removal performed on the controllers is a homomorphic transformation, so that the simplified language containment check can automatically be applied without testing the validity of the homomorphism. This homomorphism correctness verification, required when a formal proof is not available, can be executed using a tool like Cospan, but it may not be completed when the state space to be traversed is too large and deep. The redundancy removal performed on the controllers is important because it eliminates the spurious behaviors introduced in the system by the nondeterminism of the reduced data paths. Redundancy, in fact, may induce a failure in the verification of L(P')/spl sube/L(T'), while L(P)/spl sube/L(T) actually holds. In order to show the effectiveness of the proposed methodology, we verify properties on an extended version of the Mead-Conway Traffic Light Controller, on a modified IRQ communication protocol, and on a relatively prime integers checker and generator.
Enrico Macii, Bernard Plessier, Fabio Somenzi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1996 Symbolic Optimization of FSM Networks Based on Sequential ATPG Techniques
abstract
This paper presents a novel optimization algorithm for FSM networks that relies on sequential test generation and redundancy removal. The implementation of the proposed approach, which is based on the exploitation of input don't care sequences through regular language intersection, is fully symbolic. Experimental results, obtained on a large set of standard benchmarks, improve over the ones of state-of-the-art methods.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto
DAC3
1996 Test Generation for Networks of Interacting FSMs Using Symbolic Techniques
abstract
This paper presents a new testing strategy for networks of interacting FSMs. The approach allows us to generate test patterns for faults in the network by separately handling the network's components. The proposed algorithms are fully symbolic; therefore, they allow the manipulation of large designs. Experimental results, though preliminary, are promising.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto
Great Lakes Symposium on VLSI3
1996 Exact Computation of the Entropy of a Logic Circuit
abstract
Computing the entropy of a digital circuit has proved to be very useful for several applications in the area of VLSI system design. Recently, a method for entropy calculation has been used in the context of power estimation for logic circuits described at the register-transfer level. The technique has shown to be reasonably effective concerning the trade-off between the accuracy of the estimates produced and the execution time. However, the assumptions required to make the computation feasible are such that the obtained results are approximate. In this paper, we propose a symbolic algorithm for the exact calculation of the entropy of a logic circuit which is able to handle reasonably large examples without introducing any approximation. We present experimental data on standard benchmark designs in order to show the effectiveness of the new method; in addition, we compare our results to the ones obtained with the approximate approach. As a result, we observe a marginal penalty in the performance of the symbolic procedure; on the other hand, accuracy in the calculation increases significantly.
Enrico Macii, Massimo Poncino
Great Lakes Symposium on VLSI1
1996 Enhancing FSM Traversal by Temporary Re-Encoding
abstract
Synthesis and optimization of large finite-state machines has improved dramatically over the last few years with the introduction and rapid improvement of symbolic-state manipulation techniques. The algorithms efficiently visit each reachable state in the machine while computing and storing information about these states. We propose a new technique for improving the efficacy of traversal algorithms: re-encoding the states of the machine to more efficiently represent state sets or state transitions, or to more efficiently compute the next set of states. Our technique can be embedded in existing traversal algorithms. Experiments reveal that re-encoding can indeed reduce the time and/or space required for traversal.
Gianpiero Cabodi, Luciano Lavagno, Enrico Macii, Massimo Poncino, Stefano Quer, Paolo Camurati, Ellen Sentovich
ICCD3
1996 Symbolic computation of logic implications for technology-dependent low-power synthesis
abstract
This paper presents a novel technique for re-synthesizing circuits for low-power dissipation. Power consumption is reduced through redundancy addition and removal by using learning to identify indirect logic implications within a circuit. Such implications are exploited by adding gates and connections to the circuit without altering its overall behavior and thereby enabling us to eliminate other, high power dissipating, nodes. We propose a new BDD-based method for computing indirect implications in a logic network; furthermore, we present heuristic techniques to perform redundancy addition and removal without destroying the topology of the mapped circuit. Experimental results show the effectiveness of the proposed technique in reducing power while keeping within delay and area constraints.
R. Iris Bahar, M. Burns, Gary D. Hachtel, Enrico Macii, H. Shin, Fabio Somenzi
ISLPED4
1996 Automatic state space decomposition for approximate FSM traversal based on circuit analysis
abstract
Exploiting circuit structure is a key issue in the implementation of algorithms for state space decomposition when the target is approximate FSM traversal. Given the gate-level description of a sequential circuit, the information about its structure can be captured by evaluating the affinity between pairs or groups of latches. Two main factors have to be considered in carrying out the structural analysis of a sequential circuit: latch connectivity and latch correlation. The first one takes into account the mutual dependency of each memory element on the others; the second one tells us how related are the functions realized by the logic feeding each latch. In this paper we estimate the affinity of two latches by combining these two factors, and we use this measure to formulate the state space decomposition problem as a graph partitioning problem. We propose an algorithm to automatically determine "good" partitions of the latch set which induce state space decomposition, and we present approximate FSM traversal and logic optimization results for the largest ISCAS'89 sequential benchmarks.
Gary D. Hachtel, Enrico Macii, Massimo Poncino, Fabio Somenzi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 Algorithms for approximate FSM traversal based on state space decomposition
abstract
This paper presents algorithms for approximate finite state machine traversal based on state space decomposition. The original finite state machine is partitioned in component submachines, and each of them is traversed separately; the result of the computation is an over-estimation of the set of reachable states of the original machine. Different traversal strategies, which reduce the effects of the degrees of freedom introduced by the decomposition, are discussed. Efficient partitioning is a key point for the performance of the traversal techniques; a method to heuristically find a good decomposition of the overall finite state machine, based on the exploration of its state variable dependency graph, is proposed. Applications of the approximate traversal methods to logic optimization of sequential circuits and behavioral verification of finite state machines are described; experimental results for such applications, together with data concerning pure traversal, are reported.
Gary D. Hachtel, Enrico Macii, Bernard Plessier, Fabio Somenzi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 Markovian analysis of large finite state machines
abstract
Regarding finite state machines as Markov chains facilitates the application of probabilistic methods to very large logic synthesis and formal verification problems. In this paper we present symbolic algorithms to compute the steady-state probabilities for very large finite state machines (up to 10/sup 27/ states). These algorithms, based on Algebraic Decision Diagrams (ADD's)-an extension of BDD's that allows arbitrary values to be associated with the terminal nodes of the diagrams-determine the steady-state probabilities by regarding finite state machines as homogeneous, discrete-parameter Markov chains with finite state spaces, and by solving the corresponding Chapman-Kolmogorov equations. We first consider finite state machines with state graphs composed of a single terminal strongly connected component; for this type of system we have implemented two solution techniques: One is based on the Gauss-Jacobi iteration, the other one is based on simple matrix multiplication. Then we extend our treatment to the most general case of systems which can be modelled as finite state machines with arbitrary transition structures; here our approach exploits structural information to decompose and simplify the state graph of the machine. We report experimental results obtained for problems on which traditional methods fail.
Gary D. Hachtel, Enrico Macii, Abelardo Pardo, Fabio Somenzi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1995 Computing the Maximum Power Cycles of a Sequential Circuit
abstract
This paper studies the problem of estimating worst case power dissipation in a sequential circuit.We approach this problem by nding the maximum average weight cycles in a weighted directed g r aph.In order to handle practical sized examples, we use symbolic methods, based o n A lgebraic Decision Diagrams (ADDs), for computing the maximum average length cycles as well as the number of gate transitions in the circuit, which is necessary to construct the weighted directed g r aph.
Srilatha Manne, Abelardo Pardo, R. Iris Bahar, Gary D. Hachtel, Fabio Somenzi, Enrico Macii, Massimo Poncino
DAC6
1995 Estimating worst-case power consumption of CMOS circuits modeled as symbolic neural networks
abstract
In this paper we propose a new approach to the problem of estimating worst-case power consumption of CMOS combinational circuits based on neural models. Given the gate level description of a circuit, we build the corresponding neural network, we store it, we calculate the energy dissipated by the network and, finally, we derive the power dissipated by the original circuit. All the operations above are executed in the symbolic domain; that is, Algebraic Decision Diagrams are used to represent and manipulate the graph specification of the neural network modeling the circuit. We present preliminary results to show the feasibility of the method.
Enrico Macii, Massimo Poncino
Great Lakes Symposium on VLSI1
1995 Using symbolic Rademacher-Walsh spectral transforms to evaluate the correlation between Boolean functions
abstract
The use of symbolic techniques to store integer-valued functions has been shown to be extremely effective in handling both transform matrices and spectral representations of large Boolean functions. In this paper we propose a novel application of symbolic Rademacher-Walsh spectral transforms to the evaluation of Boolean function correlation. In particular, we present an ADD-based algorithm to compute the agreement between two Boolean functions starting from their spectral representations. The method, operating in the transform domain, has appeared to be more advantageous than traditional approaches, using operations in the Boolean domain, concerning both memory occupation and execution time on some classes of functions.
Enrico Macii, Massimo Poncino
Great Lakes Symposium on VLSI1
1995 Using connectivity and spectral methods to characterize the structure of sequential logic circuits
Enrico Macii, Massimo Poncino
Microprocess. Microprogramming1
1994 Probabilistic Analysis of Large Finite State Machines
abstract
Regarding finite state machines as Markov chains facilitates the application of probabilistic methods to very large logic synthesis and formal verification problems. Recently, we have shown how symbolic algorithms based on Algebraic Decision Diagrams may be used to calculate the steadystate probabilities of finite state machines with more than 10 8 states. These algorithms treated machines with state graphs composed of a single terminal strongly connected component. In this paper we consider the most general case of systems which can be modeled as state machines with arbitrary transition structures. The proposed approach exploits structural information to decompose and simplify the state graph of the machine. 1 Introduction Finite state machines (FSMs), or their extensions, are often employed to model real digital systems for formal verification. As the complexity of those systems increases, probabilistic approaches to design and implementation verification become of interest; for...
Gary D. Hachtel, Enrico Macii, Abelardo Pardo, Fabio Somenzi
DAC2
1994 A symbolic method to reduce power consumption of circuits containing false paths
R. Iris Bahar, Gary D. Hachtel, Enrico Macii, Fabio Somenzi
ICCAD3
1994 A Structural Approach to State Space Decomposition for Approximate Reachability Analysis
abstract
Exploiting circuit structure is a key issue in the implementation of algorithms for state space decomposition when the target is approximate FSM traversal. Given the gate-level description of a sequential circuit, the information about its structure can be captured by evaluating the affinity between pairs or groups of latches. Two main factors have to be considered in carrying out the structural analysis of a sequential circuit: latch connectivity and latch correlation. We estimate the affinity of two latches by combining these two factors, and we use this measure to translate the state space decomposition problem into a graph partitioning problem. Traversal results obtained on the largest ISCAS'89 benchmarks show the effectiveness of the method.>
Gary D. Hachtel, Enrico Macii, Massimo Poncino, Fabio Somenzi
ICCD3
1994 A test generation program for sequential circuits
Enrico Macii, Angelo Raffaele Meo
J. Electron. Test.1
1993 Algorithms for Approximate FSM Traversal
abstract
Article Algorithms for approximate FSM traversal Share on Authors: Hyunwoo Cho View Profile , Gary D. Hachtel View Profile , Enrico Macii View Profile , Bernard Plessier View Profile , Fabio Somenzi View Profile Authors Info & Claims DAC '93: Proceedings of the 30th international Design Automation ConferenceJuly 1993 Pages 25–30https://doi.org/10.1145/157485.164555Online:01 July 1993Publication History 72citation370DownloadsMetricsTotal Citations72Total Downloads370Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Gary D. Hachtel, Enrico Macii, Bernard Plessier, Fabio Somenzi
DAC3
1993 Δ-trees of a graph: introduction and formal definition
abstract
The authors introduce and formally define the Delta -trees of a graph. They have used Delta -trees in developing formal techniques for hardware verification based on the exploration of the graph structure of interacting finite state machines and omega -regular automata.>
Jason K. Davis, Enrico Macii
Great Lakes Symposium on VLSI2
1993 Modeling stuck-open faults in CMOS iterative circuits
abstract
Considers testability criteria for stuck-open faults in one-dimensional, unilateral, iterative circuits of CMOS combinational cells, and gives necessary and sufficient conditions for the testability of such faults. These conditions are extended to include stuck-open C-testability of the circuit. The authors consider the problem of test patterns which may be invalidated by the presence of delays in the input changes and propose restrictions to be imposed during the test vector computation in order to guarantee the robustness of the test pattern.>
Enrico Macii
Great Lakes Symposium on VLSI1
1993 Algebraic decision diagrams and their applications
abstract
In this paper we present theory and experiments on the algebraic decision diagrams (ADDs). These diagrams extend BDD's by allowing values from an arbitrary finite domain to be associated with the terminal nodes. We present a treatment founded in Boolean algebras and discuss algorithms and results in applications like matrix multiplication and shortest path algorithms. Furthermore, we outline possible applications of ADD's to logic synthesis, formal verification, and testing of digital systems.
R. Iris Bahar, Erica A. Frohm, Charles M. Gaona, Gary D. Hachtel, Enrico Macii, Abelardo Pardo, Fabio Somenzi
ICCAD5
1992 Verification of systems containing counters
abstract
It is pointed out that systems containing counters have very large and deep state spaces, and the verification of properties on these systems can be very expensive in terms of memory space and computation time. A technique for automatically reducing the state space associated with the system on which some properties that can express both safeness and fairness constraints have to be proved is presented. In particular, a set of conditions upon which some counters can be reduced to three-state, nondeterministic machines is given. The controllers can be simplified by removing the redundancy induced by their interaction with the counter, so that the verification tasks can be more easily performed.>
Enrico Macii, Bernard Plessier, Fabio Somenzi
ICCAD1
1992 Techniques to increase sequential ATPG performance
abstract
An automatic test pattern generation (ATPG) system for sequential circuits is described. Techniques such as testability measures, 9-valued functions, incompatibility function and fault simulation have been added to the basic algorithm in order to increase the fault coverage and reduce the test generation time. The entire ATPG system has been benchmarked on the set of ISCAS89 circuits.>
Enrico Macii, Angelo Raffaele Meo
VTS1
1991 An algebraic approach to test generation for sequential circuits
abstract
The authors describe an algebraic algorithm for automatic test pattern generation for sequential circuits. Three innovative concepts have been introduced in order to reduce the computational time required for pattern generation. These are: firstly, circuit partitioning in fanout-free regions; then, computation of observability and excitability functions for state propagation and justification; and finally, assignment of an observability and an excitability order to each node of the decision tree, for fast test pattern detection of each fault.>
Antonio Lioy, Enrico Macii, Angelo Raffaele Meo, Matteo Sonza Reorda
Great Lakes Symposium on VLSI2