VLDB 2026 Research / reviewers in the wild / expert
Massimo Poncino
dblp:20/691
· DBLP profile ↗
251ranked-venue papers
1as first author
36since 2021 · last 2026
0000-0002-1369-9688ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 240 · 1 first-author · 28 since 2021Software engineering, systems software and programming languages · 51 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 10 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: CHESSY: Coupled Hybrid Emulation with SystemC-FPGA SynchronizationabstractThe growing complexity of cyber-physical systems (CPSs) calls for early prototyping tools that combine accuracy, speed, and usability. Virtual Platforms (VPs) provide fast functional simulation, but hybrid co-emulation solutions, in which key digital components are deployed on FPGA, become necessary when accurate timing modelling is required and RTL simulation is too costly. However, existing hybrid emulation tools are mostly proprietary, and rely on vendor-specific FPGA features. To address this gap, we introduce an open-source framework that connects SystemC-based VPs with FPGA emulation, enabling full-system co-emulation of digital and non-digital components. The FPGA accelerates the execution of main digital subsystems, while a wrapper coordinates timing and communication with the VP through JTAG, maintaining synchronization with simulated peripherals. Evaluations using a RISC-V SoC, with an example in the biosignals processing domain, show up to 2500× speedup compared to RTL simulation, while maintaining less than 2× total simulation time relative to pure FPGA emulation. Lorenzo Ruotolo, Giovanni Pollo, Mohamed Amine Hamdi, Matteo Risso, Yukai Chen, Enrico Macii, Massimo Poncino, Sara Vinco, Alessio Burrello, Daniele Jahier Pagliari |
DATE | 7 |
| 2026 | An Open Source Design Exploration Tool for Battery and Coolant ConfigurationabstractEnsuring both electrical performance and effective thermal management in large-scale battery packs is a critical challenge for next-generation electric mobility and energy storage systems. Current modeling approaches often rely on rigid configurations or computationally expensive CFD simulations, limiting their use in early design stages. This work introduces a modular, compositional framework that enables the dynamic construction of battery packs of arbitrary size, where each cell is modeled individually with coupled electrical and thermal dynamics. The framework integrates a configurable liquid cooling system supporting multiple layouts and coolant types, allowing rapid evaluation of thermal management strategies under diverse operating conditions. By combining scalability, flexibility, and high computational efficiency, the proposed approach accelerates design iterations, reduces prototyping costs, and supports the development of safer and more reliable battery systems for real-world applications. Francesco Tosoni 0002, Yukai Chen, Massimo Poncino, Franco Fummi, Sara Vinco |
DATE | 3 |
| 2026 | End-to-end Automated Deep Neural Network Optimization for PPG-based Blood Pressure Estimation on WearablesabstractPhotoplethysmography-based Blood Pressure (BP) estimation is a challenging task, particularly on resource-constrained wearable devices. However, fully on-board processing is desirable to ensure user data confidentiality. Recent Deep Neural Networks (DNNs) have achieved high BP estimation accuracy by reconstructing BP waveforms or directly regressing BP values, but their large memory, computation, and energy requirements hinder deployment on wearables. This work introduces a fully automated DNN design pipeline that combines hardware-aware Neural Architecture Search, pruning, and Mixed-Precision Search to generate accurate yet compact BP prediction models optimized for ultra-low-power multi-core Systems-on-Chip (SoCs). Starting from state-of-the-art baseline models on four public datasets, our optimized networks achieve up to 7.99% lower error with a 7.5 \(\times\) parameter reduction, or up to 83 \(\times\) fewer parameters with negligible accuracy loss. All models fit within 512 kB of memory on our target SoC (GreenWaves’ GAP8), requiring less than 55 kB and achieving an average inference latency of 142 ms and energy consumption of 7.25 mJ. Patient-specific fine-tuning further improves accuracy by up to 64%, enabling fully autonomous, low-cost BP monitoring on wearables. Francesco Carlucci, Giovanni Pollo, Xiaying Wang, Massimo Poncino, Enrico Macii, Luca Benini, Sara Vinco, Alessio Burrello, Daniele Jahier Pagliari |
ACM Trans. Comput. Heal. | 4 |
| 2025 | Energy-Aware Error Correction Method for Indoor Positioning and TrackingabstractIndoor positioning is crucial for the effective use of drones in smart environments, enabling precise navigation and control in complex indoor spaces where GPS signals are weak or unavailable and wireless communication-based systems must be used. In order to improve positioning accuracy, various distance measurement techniques and related error correction methods have been proposed in the literature. However, these methods are mostly focused on accuracy and often require a significant amount of computational resources, which is quite inefficient when deployed on battery-operated devices like small robots or drones because of their limited battery capacity. Moreover, conventional error correction methods are little effective for the tracking of moving objects. In this paper, we first analyze the trade-off between energy consumption and accuracy for the error correction and identify the most energy-efficient error correction method. Based on this analysis in the accuracy/energy space, we introduce a new energy-efficient error correction method that is especially targeted for tracking a moving object. We validated our solution by implementing an Ultra-Wideband based indoor positioning system and demonstrated that the proposed method improves positioning accuracy by 15% and reduces energy consumption by 33% compared to the state-of-the-art method. Donkyu Baek, Yukai Chen, Enrico Macii, Massimo Poncino |
DATE | 5 |
| 2025 | Coupling Neural Networks and Physics Equations For Li-Ion Battery State-of-Charge PredictionabstractEstimating the evolution of the battery's State of Charge (SoC) in response to its usage is critical for implementing effective power management policies and for ultimately improving the system's lifetime. Most existing estimation methods are either physics-based digital twins of the battery or data-driven models such as Neural Networks (NNs). In this work, we propose two new contributions in this domain. First, we introduce a novel NN architecture formed by two cascaded branches: one to predict the current SoC based on sensor readings, and one to estimate the SoC at a future time as a function of the load behavior. Second, we integrate battery dynamics equations into the training of our NN, merging the physics-based and data-driven approaches, to improve the models' generalization over variable prediction horizons. We validate our approach on two publicly accessible datasets, showing that our Physics-Informed Neural Networks (PINNs) outperform purely data-driven ones while also obtaining superior prediction accuracy with a smaller architecture with respect to the state-of-the-art. Giovanni Pollo, Alessio Burrello, Enrico Macii, Massimo Poncino, Sara Vinco, Daniele Jahier Pagliari |
DATE | 4 |
| 2025 | Building Damage Assessment in Conflict Zones: A Deep Learning Approach Using Geospatial Sub-Meter Resolution DataabstractVery High Resolution (VHR) geospatial image analysis is crucial for humanitarian assistance in both natural and anthropogenic crises, as it allows to rapidly identify the most critical areas that need support. Nonetheless, manually inspecting large areas is time-consuming and requires domain expertise. Thanks to their accuracy, generalization capabilities, and highly parallelizable workload, Deep Neural Networks (DNNs) provide an excellent way to automate this task. Nevertheless, there is a scarcity of VHR data pertaining to conflict situations, and consequently, of studies on the effectiveness of DNNs in those scenarios. Motivated by this, our work extensively studies the applicability of a collection of state-of-the-art Convolutional Neural Networks (CNNs) originally developed for natural disasters damage assessment in a war scenario. To this end, we build an annotated dataset with pre- and post-conflict images of the Ukrainian city of Mariupol. We then explore the transferability of the CNN models in both zero-shot and learning scenarios, demonstrating their potential and limitations. To the best of our knowledge, this is the first study to use sub-meter resolution imagery to assess building damage in combat zones. Matteo Risso, Alessia Goffi, Beatrice Alessandra Motetti, Alessio Burrello, Jean Baptiste Bove, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari, Giuseppe Maffeis |
IPAS | 7 |
| 2025 | MEbots: Integrating a RISC-V Virtual Platform with a Robotic Simulator for Energy-aware DesignabstractVirtual Platforms (VPs) enable early software validation of autonomous systems’ electronics, reducing costs and time-to-market. While many VPs support both functional and non-functional simulation (e.g., timing, power), they lack the capability of simulating the environment in which the system operates. In contrast, robotics simulators lack accurate timing and power features. This twofold shortcoming limits the effectiveness of the design flow, as the designer can not fully evaluate the features of the solution under development. This paper presents a novel, fully open-source framework bridging this gap by integrating a robotics simulator (Webots) with a VP for RISC-V-based systems (MESSY). The framework enables a holistic, mission-level, energy-aware co-simulation of electronics in their surrounding environment, streamlining the exploration of design configurations and advanced power management policies. Giovanni Pollo, Mohamed Amine Hamdi, Matteo Risso, Lorenzo Ruotolo, Pietro Furbatto, Matteo Isoldi, Yukai Chen, Alessio Burrello, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari, Sara Vinco |
ISLPED | 10 |
| 2025 | Foundation Models for Structural Health MonitoringabstractStructural Health Monitoring (SHM) is a critical task for ensuring the safety and reliability of civil infrastructures, typically realized on bridges and viaducts by means of vibration monitoring. In this paper, we propose for the first time the use of Transformer neural networks, with a Masked Auto-Encoder architecture, asFoundation Modelsfor SHM. We demonstrate the ability of these models to learn generalizable representations from multiple large datasets through self-supervised pre-training, which, coupled with task-specific fine-tuning, allows them to outperform state-of-the-art traditional methods on diverse tasks, including Anomaly Detection (AD) and Traffic Load Estimation (TLE). We then extensively explore model size versus accuracy trade-offs and experiment with Knowledge Distillation (KD) to improve the performance of smaller Transformers, enabling their embedding directly into the SHM edge nodes. We showcase the effectiveness of our foundation models using data from three operational viaducts. For AD, we achieve a near-perfect 99.9% accuracy with a monitoring time span of just 15 windows. In contrast, a state-of-the-art method based on Principal Component Analysis (PCA) obtains its first good result (95.03% accuracy), only considering 120 windows. On two different TLE tasks, our models obtain state-of-the-art performance on multiple evaluation metrics (R2score, MAE% and MSE%). On the first benchmark, we achieve an R2score of 0.97 and 0.90 for light and heavy vehicle traffic, respectively, while the best previous approach (a Random Forest) stops at 0.91 and 0.84. On the second one, we achieve an R2score of 0.54 versus the 0.51 of the best competitor method, a Long-Short Term Memory network. Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta, Yhorman Alexander Bedoya Velez, Andrea Acquaviva, Massimo Poncino, Enrico Macii, Luca Benini, Alessio Burrello |
IEEE Trans. Sustain. Comput. | 6 |
| 2024 | Model-Driven Feature Engineering for Data-Driven Battery SOH ModelabstractAccurate State of Health (SoH) estimation is indispensable for ensuring battery system safety, reliability, and run-time monitoring. However, as instantaneous runtime measurement of SoH remains impractical when not unfeasible, appropriate models are required for its estimation. Recently, various data-driven models have been proposed, which solve various weaknesses of traditional models. However, the accuracy of data-driven models heavily depends on the quality of the training datasets, which usually contain data that are easy to measure but that are only partially or weakly related to the physical/chemical mechanisms that determine battery aging. In this study, we propose a novel feature engineering approach, which involves augmenting the original dataset with purpose-designed features that better represent the aging phenomena. Our contribution does not consist of a new machine-learning model but rather in the addition of selected features to an existing model. This methodology consistently demonstrates enhanced accuracy across various machine-learning models and battery chemistries, yielding an approximate 25% SoH estimation accuracy improve-ment. Our work bridges a critical gap in battery research, offering a promising strategy to significantly enhance SoH estimation by optimizing feature selection. Khaled Alamin, Daniele Jahier Pagliari, Yukai Chen, Enrico Macii, Sara Vinco, Massimo Poncino |
DATE | 6 |
| 2024 | HW-SW Optimization of DNNs for Privacy-Preserving People Counting on Low-Resolution Infrared ArraysabstractLow-resolution infrared (IR) array sensors enable people counting applications such as monitoring the occupancy of spaces and people flows while preserving privacy and minimizing energy consumption. Deep Neural Networks (DNNs) have been shown to be well-suited to process these sensor data in an accurate and efficient manner. Nevertheless, the space of DNNs' archi-tectures is huge and its manual exploration is burdensome and often leads to sub-optimal solutions. To overcome this problem, in this work, we propose a highly automated full-stack optimization flow for DNNs that goes from neural architecture search, mixed-precision quantization, and post-processing, down to the realization of a new smart sensor prototype, including a Microcontroller with a customized instruction set. Integrating these cross-layer optimizations, we obtain a large set of Pareto-optimal solutions in the 3D-space of energy, memory, and accuracy. Deploying such solutions on our hardware platform, we improve the state-of-the-art achieving up to 4.2 x model size reduction, 23.8 x code size reduction, and 15.38 x energy reduction at iso-accuracy. Matteo Risso, Francesco Daghero, Alessio Burrello, Seyedmorteza Mollaei, Marco Castellano, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
DATE | 8 |
| 2024 | Dynamic Decision Tree Ensembles for Energy-Efficient Inference on IoT Edge NodesabstractWith the increasing popularity of Internet of Things (IoT) devices, there is a growing need for energy-efficient machine learning (ML) models that can run on constrained edge nodes. Decision tree ensembles, such as random forests (RFs) and gradient boosting (GBTs), are particularly suited for this task, given their relatively low complexity compared to other alternatives. However, their inference time and energy costs are still significant for edge hardware. Given that said costs grow linearly with the ensemble size, this article proposes the use of dynamic ensembles, that adjust the number of executed trees based both on a latency/energy target and on the complexity of the processed input, to tradeoff computational cost and accuracy. We focus on deploying these algorithms on multicore low-power IoT devices, designing a tool that automatically converts a Python ensemble into optimized C code, and exploring several optimizations that account for the available parallelism and memory hierarchy. We extensively benchmark both static and dynamic RFs and GBTs on three state-of-the-art IoT-relevant data sets, using an 8-core ultralow-power System-on-Chip (SoC), GAP8, as the target platform. Thanks to the proposed early stopping mechanisms, we achieve an energy reduction of up to 37.9% with respect to static GBTs (8.82 uJ versus 14.20 uJ per inference) and 41.7% with respect to static RFs (2.86 uJ versus 4.90 uJ per inference), without losing accuracy compared to the static model. Francesco Daghero, Alessio Burrello, Enrico Macii, Paolo Montuschi, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Internet Things J. | 5 |
| 2023 | Energy-efficient Wearable-to-Mobile Offload of ML Inference for PPG-based Heart-Rate EstimationabstractModern smartwatches often include photoplethysmographic (PPG) sensors to measure heartbeats or blood pressure through complex algorithms that fuse PPG data with other signals. In this work, we propose a collaborative inference approach that uses both a smartwatch and a connected smartphone to maximize the performance of heart rate (HR) tracking while also maximizing the smartwatch's battery life. In particular, we first analyze the trade-offs between running on-device HR tracking or offloading the work to the mobile. Then, thanks to an additional step to evaluate the difficulty of the upcoming HR prediction, we demonstrate that we can smartly manage the workload between smartwatch and smartphone, maintaining a low mean absolute error (MAE) while reducing energy consumption. We benchmark our approach on a custom smartwatch prototype, including the STM32WB55 MCU and Bluetooth Low-Energy (BLE) communication, and a Raspberry Pi3 as a proxy for the smartphone. With our Collaborative Heart Rate Inference System (CHRIS), we obtain a set of Pareto-optimal configurations demonstrating the same MAE as State-of-Art (SoA) algorithms while consuming less energy. For instance, we can achieve approximately the same MAE of TimePPG-Small [1] (5.54 BPM MAE vs. 5.60 BPM MAE) while reducing the energy by 2.03×, with a configuration that offloads 80% of the predictions to the phone. Furthermore, accepting a performance degradation to 7.16 BPM of MAE, we can achieve an energy consumption of 179 uJ per prediction, 3.03× less than running TimePPG-Small on the smartwatch, and 1.82× less than streaming all the input data to the phone. Alessio Burrello, Matteo Risso, Noemi Tomasello, Yukai Chen, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
DATE | 7 |
| 2023 | Model-Driven Dataset Generation for Data-Driven Battery SOH ModelsabstractEstimating the State of Health (SOH) of batteries is crucial for ensuring the reliable operation of battery systems. Since there is no practical way to instantaneously measure it at run time, a model is required for its estimation. Recently, several data-driven SOH models have been proposed, whose accuracy heavily relies on the quality of the datasets used for their training. Since these datasets are obtained from measurements, they are limited in the variety of the charge/discharge profiles. To address this scarcity issue, we propose generating datasets by simulating a traditional battery model (e.g., a circuit-equivalent one). The primary advantage of this approach is the ability to use a simulatable battery model to evaluate a potentially infinite number of workload profiles for training the data-driven model. Furthermore, this general concept can be applied using any simulatable battery model, providing a fine spectrum of accuracy/complexity tradeoffs. Our results indicate that using simulated data achieves reasonable accuracy in SOH estimation, with a 7.2 % error relative to the simulated model, in exchange for a 27X memory reduction and a$\approx 2000\mathrm{X}$speedup. Khaled Alamin, Francesco Daghero, Giovanni Pollo, Daniele Jahier Pagliari, Yukai Chen, Enrico Macii, Massimo Poncino, Sara Vinco |
ISLPED | 7 |
| 2023 | Precision-aware Latency and Energy Balancing on Multi-Accelerator Platforms for DNN InferenceabstractThe need to execute Deep Neural Networks (DNNs) at low latency and low power at the edge has spurred the development of new heterogeneous Systems-on-Chips (SoCs) encapsulating a diverse set of hardware accelerators. How to optimally map a DNN onto such multi-accelerator systems is an open problem. We propose ODiMO, a hardware-aware tool that performs a fine-grain mapping across different accelerators on-chip, splitting individual layers and executing them in parallel, to reduce inference energy consumption or latency, while taking into account each accelerator's quantization precision to maintain accuracy. Pareto-optimal networks in the accuracy vs. energy or latency space are pursued for three popular dataset/DNN pairs, and deployed on the DIANA heterogeneous ultra-low power edge AI SoC. We show that ODiMO reduces energy/latency by up to 33%/31% with limited accuracy drop (−0.53%/-0.32%) compared to manual heuristic mappings. Matteo Risso, Alessio Burrello, Giuseppe Maria Sarda, Luca Benini, Enrico Macii, Massimo Poncino, Marian Verhelst, Daniele Jahier Pagliari |
ISLPED | 6 |
| 2023 | Efficient Deep Learning Models for Privacy-Preserving People Counting on Low-Resolution Infrared ArraysabstractUltralow-resolution infrared (IR) array sensors offer a low cost, energy efficient, and privacy-preserving solution for people counting, with applications, such as occupancy monitoring and visitor flow analysis in private and public spaces. Previous work has shown that deep learning (DL) can yield superior performance on this task. However, the literature was missing an extensive comparative analysis of various efficient DL architectures for IR array-based people counting, that considers not only their accuracy but also the cost of deploying them on memory- and energy-constrained Internet of Things (IoT) edge nodes. Such analysis is key for system designers, since it helps them select the most appropriate DL model given the constraints of their target hardware. In this work, we address this need by comparing six different DL architectures on a novel data set composed of IR images collected from a commercial$8\times8$array, which we made openly available. With a wide architectural exploration of each model type, we obtain a rich set of Pareto-optimal solutions, spanning cross-validated balanced accuracy scores in the 55.70%–82.70% range. When deployed on a commercial microcontroller (MCU) by STMicroelectronics, the STM32L4A6ZG, these models occupy 0.41–9.28kB of memory, and require 1.10–7.74 ms per inference, while consuming 17.18–$120.43 \mu \text{J}$of energy. Our models are significantly more accurate than a previous deterministic method (up to +39.9%), while being up to$3.53\times $faster and more energy efficient. So, our work serves also as a demonstration that DL can not only achieve higher accuracy but also higher efficiency compared to classic algorithms for this type of task. Further, our models’ accuracy is comparable to state-of-the-art DL solutions on similar resolution sensors, despite a much lower complexity. All our models enable continuous, real-time inference on an MCU-based IoT node, with years of autonomous operation without battery recharging. Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Internet Things J. | 8 |
| 2023 | Lightweight Neural Architecture Search for Temporal Convolutional Networks at the EdgeabstractNeural Architecture Search (NAS) is quickly becoming the go-to approach to optimize the structure of Deep Learning (DL) models for complex tasks such as Image Classification or Object Detection. However, many other relevant applications of DL, especially at the edge, are based on time-series processing and require models with unique features, for which NAS is less explored. This work focuses in particular on Temporal Convolutional Networks (TCNs), a convolutional model for time-series processing that has recently emerged as a promising alternative to more complex recurrent architectures. We propose the first NAS tool that explicitly targets the optimization of the most peculiar architectural parameters of TCNs, namely dilation, receptive-field and number of features in each layer. The proposed approach searches for networks that offer good trade-offs between accuracy and number of parameters/operations, enabling an efficient deployment on embedded platforms. Moreover, its fundamental feature is that of being lightweight in terms of search complexity, making it usable even with limited hardware resources. We test the proposed NAS on four real-world, edge-relevant tasks, involving audio and bio-signals: (i) PPG-based Heart-Rate Monitoring, (ii) ECG-based Arrythmia Detection, (iii) sEMG-based Hand-Gesture Recognition, and (iv) Keyword Spotting.Results show that, starting from a single seed network, our method is capable of obtaining a rich collection of Pareto optimal architectures, among which we obtain models with the same accuracy as the seed, and 15.9-152× fewer parameters. Moreover, the NAS finds solutions that Pareto-dominate state-of-the-arthand-tuned models for 3 out of the 4 benchmarks, and are Pareto-optimal on the fourth (sEMG). Compared to three state-of-the-art NAS tools, ProxylessNAS, MorphNet and FBNetV2, our method explores a larger search space for TCNs (up to 1012×) and obtains superior solutions, while requiring low GPU memory and search time. We deploy our NAS outputs on two distinct edge devices, the multicore GreenWaves Technology GAP8 IoT processor and the single-core STMicroelectronics STM32H7 microcontroller. With respect to the state-of-the-art hand-tuned models, we reduce latency and energy of up to 5.5× and 3.8× on the two targets respectively, without any accuracy loss. Matteo Risso, Alessio Burrello, Francesco Conti 0001, Lorenzo Lamberti, Yukai Chen, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Trans. Computers | 8 |
| 2022 | Bioformers: Embedding Transformers for Ultra-Low Power sEMG-based Gesture RecognitionabstractHuman-machine interaction is gaining traction in rehabilitation tasks, such as controlling prosthetic hands or robotic arms. Gesture recognition exploiting surface electromyographic (sEMG) signals is one of the most promising approaches, given that sEMG signal acquisition is non-invasive and is directly related to muscle contraction. However, the analysis of these signals still presents many challenges since similar gestures result in similar muscle contractions. Thus the resulting signal shapes are almost identical, leading to low classification accuracy. To tackle this challenge, complex neural networks are employed, which require large memory footprints, consume relatively high energy and limit the maximum battery life of devices used for classification. This work addresses this problem with the introduction of the Bioformers. This new family of ultra-small attention-based architectures approaches state-of-the-art performance while reducing the number of parameters and operations of 4.9 ×. Additionally, by introducing a new inter-subjects pre-training, we improve the accuracy of our best Bioformer by 3.39 %, matching state-of-the-art accuracy without any additional inference cost. Deploying our best performing Bioformer on a Parallel, Ultra-Low Power (PULP) microcontroller unit (MCU), the GreenWaves GAP8, we achieve an inference latency and energy of 2.72 ms and 0.14 mJ, respectively, 8.0× lower than the previous state-of-the-art neural network, while occupying just 94.2 kB of memory. Alessio Burrello, Francesco Bianco Morghet, Moritz Scherer 0001, Simone Benatti, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
DATE | 7 |
| 2022 | Quality inspection of critical aircraft engine components: towards full automationabstractThe quality of products has become a key factor for success in the current manufacturing industry. This is especially true in the aviation field where components for aircraft engines called honeycombs are produced. Due to their small dimension and peculiar shape, such components typically undergo a severe and cumbersome visual inspection by specialized operators. However, this process is highly prone to human error, requires a lot of time and a high number of undetected defects have been reported. In order to reduce the whole inspection time, ensure higher quality and guarantee standardization and process control, this paper presents an innovative strategy for the fully-automated inspection of honeycomb engine parts. The proposed solution is a two-phase process fully controlled by a robot, leveraging a camera as well as a purposely designed optic fibers sensor, coupled with Artificial Intelligence (AI) algorithms for the detection of different types of defects. To assess the functionality and validity of the proposed solution, a fully functioning prototype is described and characterized. Davide Cannizzaro, Filomena Simone, Klaus Illgner-Fehns, Sara Mata, Ivan Mondino, Alberto Ghiazza, Massimo Poncino, Santa Di Cataldo |
ETFA | 7 |
| 2022 | C-NMT: A Collaborative Inference Framework for Neural Machine TranslationabstractCollaborative Inference (CI) optimizes the latency and energy consumption of deep learning inference through the inter-operation of edge and cloud devices. Albeit beneficial for other tasks, CI has never been applied to the sequence-to-sequence mapping problem at the heart of Neural Machine Translation (NMT). In this work, we address the specific issues of collaborative NMT, such as estimating the latency required to generate the (unknown) output sequence, and show how existing CI methods can be adapted to these applications. Our experiments show that CI can reduce the latency of NMT by up to 44% compared to a non-collaborative approach. Yukai Chen, Roberta Chiaro, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISCAS | 4 |
| 2022 | Privacy-preserving Social Distance Monitoring on Microcontrollers with Low-Resolution Infrared Sensors and CNNsabstractLow-resolution infrared (IR) array sensors offer a low-cost, low-power, and privacy-preserving alternative to optical cameras and smartphones/wearables for social distance monitoring in indoor spaces, permitting the recognition of basic shapes, without revealing the personal details of individuals. In this work, we demonstrate that an accurate detection of social distance violations can be achieved processing the raw output of a 8x8 IR array sensor with a small-sized Convolutional Neural Network (CNN). Furthermore, the CNN can be executed directly on a Microcontroller (MCU)-based sensor node.With results on a newly collected open dataset, we show that our best CNN achieves 86.3% balanced accuracy, significantly outperforming the 61% achieved by a state-of-the-art deterministic algorithm. Changing the architectural parameters of the CNN, we obtain a rich Pareto set of models, spanning 70.5-86.3% accuracy and 0.18-75k parameters. Deployed on a STM32L476RGMCU, these models have a latency of 0.73-5.33ms, with an energy consumption per inference of 9.38-68.57$\mu$J. Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISCAS | 8 |
| 2022 | Multi-Complexity-Loss DNAS for Energy-Efficient and Memory-Constrained Deep Neural NetworksabstractNeural Architecture Search (NAS) is increasingly popular to automatically explore the accuracy versus computational complexity trade-off of Deep Learning (DL) architectures. When targeting tiny edge devices, the main challenge for DL deployment is matching the tight memory constraints, hence most NAS algorithms consider model size as the complexity metric. Other methods reduce the energy or latency of DL models by trading off accuracy and number of inference operations. Energy and memory are rarely considered simultaneously, in particular by low-search-cost Differentiable NAS (DNAS) solutions. Matteo Risso, Alessio Burrello, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISLPED | 5 |
| 2022 | Embedding Temporal Convolutional Networks for Energy-efficient PPG-based Heart Rate MonitoringabstractPhotoplethysmography (PPG) sensors allow for non-invasive and comfortable heart rate (HR) monitoring, suitable for compact wrist-worn devices. Unfortunately, motion artifacts (MAs) severely impact the monitoring accuracy, causing high variability in the skin-to-sensor interface. Several data fusion techniques have been introduced to cope with this problem, based on combining PPG signals with inertial sensor data. Until now, both commercial and reasearch solutions are computationally efficient but not very robust, or strongly dependent on hand-tuned parameters, which leads to poor generalization performance. In this work, we tackle these limitations by proposing a computationally lightweight yet robust deep learning-based approach for PPG-based HR estimation. Specifically, we derive a diverse set of Temporal Convolutional Networks for HR estimation, leveraging Neural Architecture Search. Moreover, we also introduce ActPPG, an adaptive algorithm that selects among multiple HR estimators depending on the amount of MAs, to improve energy efficiency. We validate our approaches on two benchmark datasets, achieving as low as 3.84 beats per minute of Mean Absolute Error on PPG-Dalia, which outperforms the previous state of the art. Moreover, we deploy our models on a low-power commercial microcontroller (STM32L4), obtaining a rich set of Pareto optimal solutions in the complexity vs. accuracy space. Alessio Burrello, Daniele Jahier Pagliari, Pierangelo Maria Rapa, Matilde Semilia, Matteo Risso, Tommaso Polonelli, Massimo Poncino, Luca Benini, Simone Benatti |
ACM Trans. Comput. Heal. | 7 |
| 2022 | A Smart Meter Infrastructure for Smart Grid IoT ApplicationsabstractElectric infrastructures have been pushed forward to handle tasks they were not originally designed to perform. To improve reliability and efficiency, state-of-the-art power grids include improved security, reduced peak loads, increased integration of renewable sources, and lower operational costs. In this framework, “smart grids” are built around bidirectional communication technologies, where “smart meters” communicate with all other entities and collect data from the power grid, offering specific features to each actor playing in the energy marketplace. In this article, to overcome some of the challenges raised by smart grids and smart meters, we propose a distributed metering infrastructure, which provides bidirectional communication, self-configuration, and autoupdate capabilities. Our 3-phase smart meters follow the basics Internet of Things principles and have the ability to run, either onboard or distributed on the network, multiple algorithms for smart grid management. These algorithms can be freely added, updated, or removed on the fly, thanks to the autoupdate feature of the system. Moreover, to reduce costs and improve scalability, we prove that it is possible to implement our smart meters using only off-the-shelf and inexpensive hardware devices. A digital real-time simulator (i.e., Opal-RT) has been used to assess the capabilities of both the infrastructure and the meter. Our experimental analysis shows that the latency introduced by the data transmission over the Internet is compliant with the limits imposed by the IEC 61850 standard. As a consequence, our architecture does not affect the operational status of the smart grid, making it a viable solution to support the deployment of novel services. Matteo Orlando, Abouzar Estebsari, Enrico Pons, Marco Pau, Stefano Quer, Massimo Poncino, Lorenzo Bottaccioli, Edoardo Patti |
IEEE Internet Things J. | 6 |
| 2022 | Human Activity Recognition on Microcontrollers with Quantized and Adaptive Deep Neural NetworksabstractHuman Activity Recognition (HAR) based on inertial data is an increasingly diffused task on embedded devices, from smartphones to ultra low-power sensors. Due to the high computational complexity of deep learning models, most embedded HAR systems are based on simple and not-so-accurate classic machine learning algorithms. This work bridges the gap between on-device HAR and deep learning, proposing a set of efficient one-dimensional Convolutional Neural Networks (CNNs) that can be deployed on general purpose microcontrollers (MCUs). Our CNNs are obtained combining hyper-parameters optimization with sub-byte and mixed-precision quantization, to find good trade-offs between classification results and memory occupation. Moreover, we also leverage adaptive inference as an orthogonal optimization to tune the inference complexity at runtime based on the processed input, hence producing a more flexible HAR system. With experiments on four datasets, and targeting an ultra-low-power RISC-V MCU, we show that (i) we are able to obtain a rich set of Pareto-optimal CNNs for HAR, spanning more than 1 order of magnitude in terms of memory, latency, and energy consumption; (ii) thanks to adaptive inference, we can derive >20 runtime operating modes starting from a single CNN, differing by up to 10% in classification scores and by more than 3× in inference complexity, with a limited memory overhead; (iii) on three of the four benchmarks, we outperform all previous deep learning methods, while reducing the memory occupation by more than 100×. The few methods that obtain better performance (both shallow and deep) are not compatible with MCU deployment; (iv) all our CNNs are compatible with real-time on-device HAR, achieving an inference latency that ranges between 9 μs and 16 ms. Their memory occupation varies in 0.05–23.17 kB, and their energy consumption in 0.05 and 61.59 μJ, allowing years of continuous operation on a small battery supply. Francesco Daghero, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2021 | Ultra-compact binary neural networks for human activity recognition on RISC-V processorsabstractHuman Activity Recognition (HAR) is a relevant inference task in many mobile applications. State-of-the-art HAR at the edge is typically achieved with lightweight machine learning models such as decision trees and Random Forests (RFs), whereas deep learning is less common due to its high computational complexity. In this work, we propose a novel implementation of HAR based on deep neural networks, and precisely on Binary Neural Networks (BNNs), targeting low-power general purpose processors with a RISC-V instruction set. BNNs yield very small memory footprints and low inference complexity, thanks to the replacement of arithmetic operations with bit-wise ones. However, existing BNN implementations on general purpose processors impose constraints tailored to complex computer vision tasks, which result in over-parametrized models for simpler problems like HAR. Therefore, we also introduce a new BNN inference library, which targets ultra-compact models explicitly. With experiments on a single-core RISC-V processor, we show that BNNs trained on two HAR datasets obtain higher classification accuracy compared to a state-of-the-art baseline based on RFs. Furthermore, our BNN reaches the same accuracy of a RF with either less memory (up to 91%) or more energy-efficiency (up to 70%), depending on the complexity of the features extracted by the RF. Francesco Daghero, Daniele Jahier Pagliari, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino |
CF | 9 |
| 2021 | Design of District-level Photovoltaic Installations for Optimal Power Production and Economic BenefitabstractPhotoVoltaic (PV) installations are a widespread source of renewable energy, and are quite common urban buildings’ roofs. To soften both the initial investment and the recurrent maintenance costs, the current market trends delegate the construction of PV installations to Energy Aggregators, i.e., grouping of consumers and producers that act as a single entity to satisfy local energy demand and to sell the surplus energy to the grid. In this perspective, PV installations can be designed with a larger perspective, i.e., at district level, to maximize power production not of a single building but rather of a number of blocks of a city. This implies new challenges, including efficient data management (the covered area can be squared kilometers wide) and optimal PV installation (the number of PV modules can be in the order of hundreds or even thousands). This paper proposes a framework to combine detailed geographic and irradiance information to determine an optimal PV installation over a district, by maximizing both power production and economic convenience. Our simulation results run on a real-world district prove that the framework allows an advanced evaluation of costs and benefit, that can be used by Energy Aggregators to design a new PV installation, and demonstrate an improvement on power generation up to 20% w.r.t. standard installations. Matteo Orlando, Lorenzo Bottaccioli, Sara Vinco, Enrico Macii, Massimo Poncino, Edoardo Patti |
COMPSAC | 5 |
| 2021 | Pruning In Time (PIT): A Lightweight Network Architecture Optimizer for Temporal Convolutional NetworksabstractTemporal Convolutional Networks (TCNs) are promising Deep Learning models for time-series processing tasks. One key feature of TCNs is time-dilated convolution, whose optimization requires extensive experimentation. We propose an automatic dilation optimizer, which tackles the problem as a weight pruning on the time-axis, and learns dilation factors together with weights, in a single training. Our method reduces the model size and inference latency on a real SoC hardware target by up to 7.4× and 3×, respectively with no accuracy drop compared to a network without dilation. It also yields a rich set of Pareto-optimal TCNs starting from a single model, outperforming hand-designed solutions in both size and accuracy. Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari, Francesco Conti 0001, Lorenzo Lamberti, Enrico Macii, Luca Benini, Massimo Poncino |
DAC | 8 |
| 2021 | Digital Twin Extension with Extra-Functional PropertiesabstractDigital twins of production lines do not focus solely on the management of the production process, they can also monitor and optimize other extra-functional aspects such as energy consumption and communications. This paper proposes the extension of digital twin concept in such directions. First, we extend the digital twin with models of energy consumption, that allow the monitoring of production line components throughout production lifetime. Then, we propose a flow to design the communication network starting from information obtained from the digital twin concerning the production, usage and flowing of information through the plant. All these methodologies start from the production line specification, then they enrich it with data collected during operation, and finally information is used to perform design and optimization. Results have been shown on a real Industry 4.0 research facility. Khaled Alamin, Sara Vinco, Massimo Poncino, Nicola Dall'Ora, Enrico Fraccaroli, Davide Quaglia |
DATE | 3 |
| 2021 | Robust and Energy-Efficient PPG-Based Heart-Rate MonitoringabstractA wrist-worn PPG sensor coupled with a lightweight algorithm can run on a MCU to enable non-invasive and comfortable monitoring, but ensuring robust PPG-based heart-rate monitoring in the presence of motion artifacts is still an open challenge. Recent state-of-the-art algorithms combine PPG and inertial signals to mitigate the effect of motion artifacts. However, these approaches suffer from limited generality. Moreover, their deployment on MCU-based edge nodes has not been investigated. In this work, we tackle both the aforementioned problems by proposing the use of hardware-friendly Temporal Convolutional Networks (TCN) for PPG-based heart estimation. Starting from a single "seed" TCN, we leverage an automatic Neural Architecture Search (NAS) approach to derive a rich family of models. Among them, we obtain a TCN that outperforms the previous state-of-the- art on the largest PPG dataset available (PPGDalia), achieving a Mean Absolute Error (MAE) of just 3.84 Beats Per Minute (BPM). Furthermore, we tested also a set of smaller yet still accurate (MAE of 5.64 - 6.29 BPM) networks that can be deployed on a commercial MCU (STM32L4) which require as few as 5k parameters and reach a latency of 17.1 ms consuming just 0.21 mJ per inference. Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari, Simone Benatti, Enrico Macii, Luca Benini, Massimo Poncino |
ISCAS | 7 |
| 2021 | TCN Mapping Optimization for Ultra-Low Power Time-Series Edge InferenceabstractTemporal Convolutional Networks (TCNs) are emerging lightweight Deep Learning models for Time Series analysis. We introduce an automated exploration approach and a library of optimized kernels to map TCNs on Parallel Ultra-Low Power (PULP) microcontrollers. Our approach minimizes latency and energy by exploiting a layer tiling optimizer to jointly find the tiling dimensions and select among alternative implementations of the causal and dilated 1D-convolution operations at the core of TCNs. We benchmark our approach on a commercial PULP device, achieving up to $103 \times $ lower latency and $20.3 \times $ lower energy than the Cube-AI toolkit executed on the STM32L4 and from $2.9 \times $ to $26.6 \times $ lower energy compared to commercial closed-source and academic open-source approaches on the same hardware target. Alessio Burrello, Alberto Dequino, Daniele Jahier Pagliari, Francesco Conti 0001, Marcello Zanghieri, Enrico Macii, Luca Benini, Massimo Poncino |
ISLPED | 8 |
| 2021 | ACME: An Energy-Efficient Approximate Bus Encoding for I2CabstractIn ultra low power systems with many peripherals, off-chip serial interconnects contribute significantly to the total energy budget. Leveraging the error-resilience characteristics of many embedded applications, the approximate computing paradigm has been applied to serial bus encodings to reduce interconnect consumption. However, the power model considered in previous works was purely capacitive. Accordingly, the objective of these approximate encodings was simply to reduce the transition count. While this works well for most bus standards, one notable exception is represented by I2C, whose open-drain physical connection makes the static energy consumed by logic-0 values on the bus extremely relevant. In this work, we propose ACME, the first approximate serial bus encoding targeting specifically I2C connections. With a simple encoding/decoding scheme, ACME concurrently reduces both the static and dynamic energy on the bus by maximizing the number of logic-1 values in codewords, while simultaneously reducing transitions. Using an accurate bus model and realistic capacitance and resistance values selected according to the I2C standard, we show that our encoding outperforms state-of-the-art solutions and reduces the total energy consumption on the bus by 57% on average, with an error smaller than 0.1%. Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino |
ISLPED | 5 |
| 2021 | Adaptive Random Forests for Energy-Efficient Inference on MicrocontrollersabstractRandom Forests (RFs) are widely used Machine Learning models in low-power embedded devices, due to their hardware friendly operation and high accuracy on practically relevant tasks. The accuracy of a RF often increases with the number of internal weak learners (decision trees), but at the cost of a proportional increase in inference latency and energy consumption. Such costs can be mitigated considering that, in most applications, inputs are not all equally difficult to classify. Therefore, a large RF is often necessary only for (few) hard inputs, and wasteful for easier ones. In this work, we propose an early-stopping mechanism for RFs, which terminates the inference as soon as a high-enough classification confidence is reached, reducing the number of weak learners executed for easy inputs. The early-stopping confidence threshold can be controlled at runtime, in order to favor either energy saving or accuracy. We apply our method to three different embedded classification tasks, on a single-core RISC-V microcontroller, achieving an energy reduction from 38% to more than 90% with a drop of less than 0.5% in accuracy. We also show that our approach outperforms previous adaptive ML methods for RFs. Francesco Daghero, Alessio Burrello, Luca Benini, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
VLSI-SoC | 7 |
| 2021 | Enhancing manufacturing intelligence through an unsupervised data-driven methodology for cyclic industrial processes
Tania Cerquitelli, Francesco Ventura, Daniele Apiletti, Elena Baralis, Enrico Macii, Massimo Poncino |
Expert Syst. Appl. | 6 |
| 2021 | Manufacturing as a Data-Driven Practice: Methodologies, Technologies, and ToolsabstractIn recent years, the introduction and exploitation of innovative information technologies in industrial contexts have led to the continuous growth of digital shop floor environments. The new Industry 4.0 model allows smart factories to become very advanced IT industries, generating an ever-increasing amount of valuable data. As a consequence, the necessity of powerful and reliable software architectures is becoming prominent along with data-driven methodologies to extract useful and hidden knowledge supporting the decision-making process. This article discusses the latest software technologies needed to collect, manage, and elaborate all data generated through innovative Internet-of-Things (IoT) architectures deployed over the production line, with the aim of extracting useful knowledge for the orchestration of high-level control services that can generate added business value. This survey covers the entire data life cycle in manufacturing environments, discussing key functional and methodological aspects along with a rich and properly classified set of technologies and tools, useful to add intelligence to data-driven services. Therefore, it serves both as a first guided step toward the rich landscape of the literature for readers approaching this field and as a global yet detailed overview of the current state of the art in the Industry 4.0 domain for experts. As a case study, we discuss, in detail, the deployment of the proposed solutions for two research project demonstrators, showing their ability to mitigate manufacturing line interruptions and reduce the corresponding impacts and costs. Tania Cerquitelli, Daniele Jahier Pagliari, Andrea Calimera, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Massimo Poncino |
Proc. IEEE | 7 |
| 2021 | CRIME: Input-Dependent Collaborative Inference for Recurrent Neural NetworksabstractThe excellent accuracy of Recurrent Neural Networks (RNNs) for time-series and natural language processing comes at the cost of computational complexity. Therefore, the choice between edge and cloud computing for RNN inference, with the goal of minimizing response time or energy consumption, is not trivial. An edge approach must deal with the aforementioned complexity, while a cloud solution pays large time and energy costs for data transmission. Collaborative inference is a technique that tries to obtain the best of both worlds, by splitting the inference task among a network of collaborating devices. While already investigated for other types of neural networks, collaborative inference for RNNs poses completely new challenges, such as the strong influence of input length on processing time and energy, and is greatly unexplored. In this paper, we introduce a Collaborative RNN Inference Mapping Engine(CRIME), which automatically selects the best inference device for each input. CRIME is flexible with respect to the connection topology among collaborating devices, and adapts to changes in the connections statuses and in the devices loads. With experiments on several RNNs and datasets, we show that CRIME can reduce the execution time (or end-node energy) by more than 25% compared to any single-device approach. Daniele Jahier Pagliari, Roberta Chiaro, Enrico Macii, Massimo Poncino |
IEEE Trans. Computers | 4 |
| 2021 | A Microservices-Based Framework for Smart Design and Optimization of PV InstallationsabstractThe design of photovoltaic (PV) installations mostly relies on rule-of-thumb criteria and on gross estimates of the shading patterns, and the few optimized approaches are generally focused on the problem of identifying the most suitable surfaces (e.g., roofs) in a larger geographic area (e.g., city or district). This article proposes a framework to address the design and the optimization of PV installations through a set of microservices focusing on the different variables of the design: identification of the target surfaces, elaboration of weather data, modeling of the PV panel, and floorplanning of the panel on the surface. The microservices architecture ensures extensibility and generality, as the user may execute only a subset of the proposed services or provide novel algorithms to extend the existing ones. Additionally, the framework provides a set of built-in models that allow sensitivity to the distribution of shades and accurate modeling of the power production over time. We show the many benefits of the proposed framework on two different use cases. Sara Vinco, Daniele Jahier Pagliari, Lorenzo Bottaccioli, Edoardo Patti, Enrico Macii, Massimo Poncino |
IEEE Trans. Sustain. Comput. | 6 |
| 2020 | Optimal Configuration and Placement of PV Systems in Building Roofs with Cost AnalysisabstractFollowing the Smart Grid view, current energy generation systems based on fossil fuels will be replaced with renewable energy sources. Photovoltaic (PV) is currently considered the most promising technology, due to decreasing costs of the devices and to the limited invasiveness in existing infrastructures, that make PV installations quite common urban buildings' roofs. To maximise both power production and Return Of Investment (ROI) of PV installations, new techniques and methodologies should be applied to limit sources of inefficiencies, like shading and power losses due to an incorrect installation. In this paper, we propose a novel solution for an optimal configuration and placement of PV systems in buildings' roofs. Given a number of alternative configurations and a roof of interest, it combines detailed geographic and irradiance information to determine the optimal PV installation, by maximizing both power production and ROI. Our simulation results on two real-world roofs demonstrate an improvement on power generation up to 23% w.r.t. standard compact installations. These results also highlight that a cost analysis, often ignored by standard installation strategies, is nonetheless necessary to guarantee optimal results in terms of PV production and revenue. Matteo Orlando, Lorenzo Bottaccioli, Edoardo Patti, Enrico Macii, Sara Vinco, Massimo Poncino |
COMPSAC | 6 |
| 2020 | Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceabstractGiven the computational complexity of Recurrent Neural Networks (RNNs) inference, IoT and mobile devices typically offload this task to the cloud. However, the execution time and energy consumption of RNN inference strongly depends on the length of the processed input. Therefore, considering also communication costs, it may be more convenient to process short input sequences locally and only offload long ones to the cloud. In this paper, we propose a low-overhead runtime tool that performs this choice automatically. Results based on real edge and cloud devices show that our method is able to simultaneously reduce the total execution time and energy consumption of the system compared to solutions that run RNN inference fully locally or fully in the cloud. Daniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino |
DAC | 6 |
| 2020 | A Diode-Aware Model of PV Modules from Datasheet SpecificationsabstractSemi-empirical models of photovoltaic (PV) modules based only on datasheet information are popular in electrical energy systems (EES) simulation because they can be built without measurements and allow quick exploration of alternative devices. One key limitation of these models, however, is the fact that they cannot model the presence of bypass diodes, which are inserted across a set of series-connected cells in a PV module to mitigate the impact of partial shading; datasheet information refer in fact to the operations of the module under uniform irradiance. Neglecting the effect of bypass diodes may incur in significant underestimation of the extracted power.This paper proposes a semi-empirical model for a PV module, that, by taking into account the only available information about bypass diodes in a datasheet, i.e., its number, by a first downscaling the model to a single PV cell and a subsequent upscaling to the level of a substring and of a module, allows to take into accout the diode effect as much accurately as allowed by the datasheet information.Experimental results show that, in a typical PV array on a roof, using a diode-agnostic model can signifantly underestimate the output power production. Sara Vinco, Yukai Chen, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2020 | Logic Synthesis of Pass-Gate Logic Circuits With Emerging Ambipolar TechnologiesabstractEmerging devices and new ultrascaled silicon transistors have shown disruptive electrical and functional properties that might bring digital hardware to the next level. The key issue today concerns their integration. Even though the classical complementary logic style is the most intuitive option, other strategies such as pass-transistors that were discarded in the past because they did not fit silicon MOSFETs logic should be reconsidered. Obviously, the assessment of such alternatives requires customized CAD tools and optimization engines. The objective of this paper is to introduce a synthesis and optimization flow for pass-gate logic circuits mapped onto emerging ambipolar technologies. As main contributions we propose: 1) a novel EXNOR-based decomposition technique that fully exploits do not care conditions to generate compact logic function representations and 2) a dedicated one-pass synthesis flow where optimization and technology mapping are concurrently run on a common data structure, the reduced ordered pass-diagram. Experimental results demonstrate that the proposed flow outperforms existing synthesis tools by achieving more compact circuit representations with 8.5× less devices and about 8× shallower structures (on average), while still yielding lower CPU times. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Modeling and Simulation of Cyber-Physical Electrical Energy Systems With SystemC-AMSabstractModern cyber-physical electrical energy systems (CPEES) are characterized by wider adoption of sustainable energy sources and by an increased attention to optimization, with the goal of reducing pollution and wastes. This imposes a need for instruments supporting the design flow, to simulate and validate the behavior of system components and to apply additional optimization and exploration steps. Additionally, each system might be tested with a number of management policies, to evaluate their economic impact. It is thus evident that simulation is a key ingredient in the design flow of CPEES. This paper proposes a framework for CPEES modeling and simulation, that relies on the open-source standard SystemC-AMS. The paper formalizes the information and energy flow in a generic CPEES, by focusing on both AC and DC components, and by including support for mechanical and physical models that represent multiple energy sources and loads. Experimental results, applied to a complex CPEES case study, will prove the effectiveness of the proposed solution, in terms of accuracy, speed up w.r.t. the current state-of-the-art Matlab/Simulink, and support for the design flow. Yukai Chen, Sara Vinco, Daniele Jahier Pagliari, Paolo Montuschi, Enrico Macii, Massimo Poncino |
IEEE Trans. Sustain. Comput. | 6 |
| 2019 | Low-Overhead Power Trace Obfuscation for Smart Meter PrivacyabstractSmart meters communicate to the utility provider fine-grain information about a user's energy consumption, which could be used to infer the user's habits and pose thus a critical privacy risk. State-of-the-art solutions try to obfuscate the readings of a meter either by using a large re-chargeable battery to filter the trace or by adding random noise to alter it. Both solutions, however, have significant drawbacks: large batteries are prohibitively expensive, whereas digitally added noise implies that the user entrusts the utility provider to protect his/her privacy. Daniele Jahier Pagliari, Sara Vinco, Enrico Macii, Massimo Poncino |
DAC | 4 |
| 2019 | Irradiance-Driven Partial Reconfiguration of PV PanelsabstractAdaptive reconfiguration of a photo-voltaic (PV) panel by means of a switch network is a well-known approach to tackle shading issues dynamically and with a reasonable cost. Most of these approaches assume however that the entire panel is reconfigurable, resulting in high installation costs due to the large wiring overhead required by this solution. In this work we propose an architecture in which only a portion of the panel is made reconfigurable, while minimizing the loss in the extracted power with respect to a fully reconfigurable solution. The key feature of our approach is the use of environmental (irradiance and temperature) data to determine the reconfigurable subset at design time. Simulation results show that, by reconfiguring only about 50-70% of a panel, it is possible to achieve up to 45% power increase with respect to a static topology, while losing less than 5% power with respect to full reconfiguration. Daniele Jahier Pagliari, Sara Vinco, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2019 | Dynamic Beam Width Tuning for Energy-Efficient Recurrent Neural NetworksabstractRecurrent Neural Networks (RNNs) are state-of-the-art models for many machine learning tasks, such as language modeling and machine translation. Executing the inference phase of a RNN directly in edge nodes, rather than in the cloud, would provide benefits in terms of energy consumption, latency and network bandwidth, provided that models can be made efficient enough to run on energy-constrained embedded devices. Daniele Jahier Pagliari, Francesco Panini, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | Battery-Aware Electric Truck Delivery Route PlannerabstractFinding the energy-optimal route in the context of parcel delivery with electric vehicles (EVs) is more complicated than for conventional internal combustion engine (ICE) vehicles, where the energy cost of a path is mostly determined by the total traveled distance. In the case of EV delivery, the total energy consumption strongly depends on the order of delivery because the efficiency of the EV is affected by how the transported weight changes over time as it directly affects the battery efficiency. This makes impossible to find an optimal solution using traditional routing algorithms such as the traveling salesman problem (TSP) using a static quantity (e.g., distance) as a metric.In this paper, we propose a solution for the least-energy delivery problem using EVs; we implement an electric truck simulator and evaluate different static metrics to assess their quality on small size instances for which the optimal solution can be computed exhaustively. A greedy algorithm using the empirically best metric (namely, distance × residual weight) provides significant reductions (up to 33%) with respect to a common-sense heaviest first package delivery route determined using a metric suggested by the battery properties, and is sensibly faster than state-of-the-art TSP heuristic algorithms. Donkyu Baek, Yukai Chen, Enrico Macii, Massimo Poncino, Naehyuck Chang |
ISLPED | 4 |
| 2019 | CNN-Based Camera-less User Attention Detection for Smartphone Power ManagementabstractThe many sensors hosted by mobile electronic devices are commonly used to recognize user activities and context, in order to provide new functionalities, such as tracking physical activity and sleep cycles. Despite its potential, such context recognition is only employed for power management purposes in very specific scenarios (e.g. in-pocket detection). In this work we present a novel context recognition system able to reliably identify whether a mobile device is not being looked at, and to consequently trigger power management actions such as turning off the display and moving to suspended mode. Our method takes as input the readings from common low-power sensors present in virtually all mobile devices and classifies them using a Convolutional Neural Network. Most importantly, the power-hungry camera sub-system is not used, resulting in an extremely energy-efficient detection strategy. Results show that our system is able to identify scenarios in which a device is not being used with 95.6% accuracy, thus reducing the energy overheads by 91% compared to a standard timeout-based power management and by 58% compared to a system relying on the camera. Daniele Jahier Pagliari, Matteo Ansaldi, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2019 | Fine-Grain Back Biasing for the Design of Energy-Quality Scalable OperatorsabstractEnergy-quality scalable systems are a promising solution to cope with the small energy budgets and high processing demands of mobile and Internet of Things applications. These systems leverage the error resilience of applications to obtain high energy efficiency, at the expense of tolerable reductions in the output quality. Hardware datapath operators able to reconfigure their precision and power consumption at runtime are key components of such systems. However, most implementations of these operators require manual, architecture-specific modifications and tend to have large power overheads compared to standard designs, when working at maximum precision. One promising design-independent alternative is dynamic voltage and accuracy scaling, whose adoption, however, is hindered by incompatibilities with standard design flows. In this paper, we propose a new methodology for the design of energy-quality scalable operators; our solution leverages runtime tuning of transistors threshold voltages to obtain a fine-grain control of the speed and power consumption of standard-cells within an operator. Thanks to the additional flexibility provided by this fine-grain knob, our method overcomes the main limitations of previous solutions, at the cost of a small area overhead. We demonstrate our approach on a 28 nm FDSOI technology; by exploiting the strong effect of back-gate biasing on threshold voltage, we achieve a power consumption reduction of more than 40% compared to the state-of-the-art, for the same precision. Daniele Jahier Pagliari, Yves Durand, David Coriat, Edith Beigné, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | SystemC-AMS Thermal Modeling for the Co-simulation of Functional and Extra-Functional PropertiesabstractTemperature is a critical property of smart systems, due to its impact on reliability and to its inter-dependence with power consumption. Unfortunately, the current design flows evaluate thermal evolution ex-post on offline power traces. This does not allow to consider temperature as a dimension in the design loop, and it misses all the complex inter-dependencies with design choices and power evolution. In this article, by adopting the functional language SystemC-AMS (Analog Mixed Signal), we propose a method to enable thermal/power/functional co-simulation. The system thermal model is built by using state-of-the-art circuit equivalent models, by exploiting the support for electrical linear networks intrinsic of SystemC-AMS. The experimental results will show that the choice of SystemC-AMS is a winning strategy for building a simultaneous simulation of multiple functional and extra-functional properties of a system. The generated code exposes an accuracy comparable to that of the reference thermal simulator HotSpot. Additionally, the initial overhead due to the general purpose nature of SystemC-AMS is compensated by the surprisingly high performance of transient simulation, with speedups as high as two orders of magnitude. Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2019 | A Cross-level Verification Methodology for Digital IPs Augmented with Embedded Timing MonitorsabstractSmart systems are characterized by the integration in a single device of multi-domain subsystems of different technological domains, namely, analog, digital, discrete and power devices, MEMS, and power sources. Such challenges, emerging from the heterogeneous nature of the whole system, combined with the traditional challenges of digital design, directly impact on performance and on propagation delay of digital components. This article proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems with the automatic insertion of delay sensors, which can detect and correct timing failures. The article then proposes a methodology to verify such added features at system level. The augmented model is abstracted to SystemC TLM, which is automatically injected with mutants (i.e., code mutations) to emulate delays and timing failures. The resulting TLM model is finally simulated to identify timing failures and to verify the correctness of the inserted delay monitors. Experimental results demonstrate the applicability of the proposed design and verification methodology, thanks to an efficient sensor-aware abstraction methodology, by applying the flow to three complex case studies. Sara Vinco, Nicola Bombieri, Daniele Jahier Pagliari, Franco Fummi, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2018 | All-digital embedded meters for on-line power estimationabstractModern low power designs use multiple knobs for concurrent dynamic and leakage power optimization; supply voltage and threshold voltage are the most adopted. An efficient control of these knobs needs management policies aware of the power breakdown. This implies the availability of smart on-chip strategies for dynamic and leakage power estimation at runtime. In this paper, we address this issue proposing the implementation of embedded dynamic/static power meters that use an optimized regression model fed with data collected from in-situ activity monitors. The number of sensors, their bitwidth and optimal placement are obtained through an automated design flow. The methodology works for general logic and applies not just to processor cores, but also to application-specific designs. We apply our solution to a representative class of benchmarks, showing that it can achieve an average estimation error smaller than 3%, with limited area and power overheads. Daniele Jahier Pagliari, Valentino Peluso, Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 6 |
| 2018 | GIS-based optimal photovoltaic panel floorplanning for residential installationsabstractShading is a crucial issue for the placement of PV installations, as it heavily impacts power production and the corresponding return of investment. Nonetheless, residential rooftop installations still rely on rule-of-thumb criteria and on gross estimates of the shading patterns, while more optimized approaches focus solely on the identification of suitable surfaces (e.g., roofs) in a larger geographic area (e.g., city or district). This work addresses the challenge of identifying an optimal (with respect to the overall energy production) placement of PV panels on a roof. The novel aspect of the proposed solution lies in the possibility of having a sparse, irregular placement of individual modules so as to better exploit the variance of solar data. The latter are represented in terms of the distribution of irradiance and temperature values over the roof, as elaborated from historical traces and Geographical Information System (GIS) data. Experimental results will prove the effectiveness of the algorithm through three real world case studies, and that the generated optimal solutions allow to increase power production by up to 28% with respect to rule-of-thumb solutions. Sara Vinco, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Enrico Macii, Massimo Poncino |
DATE | 6 |
| 2018 | Battery-aware Design Exploration of Scheduling Policies for Multi-sensor DevicesabstractLifetime maximization is a key challenge in battery-powered multi-sensor devices. Battery-aware power management strategies combine task scheduling with dynamic voltage scaling (DVS), accounting for the fact that the power drawn by the device is different from that provided by the battery due to its many non-idealities. However, state-of-the-art techniques in this field do not take into account several important aspects, such as the impact of sensing tasks on the overall power demand, the (operating point dependent) losses due to multiple DC-DC conversions, and the dynamic modifications in battery efficiency caused by different distributions of the currents in the temporal and in the frequency domains. In this work, we propose a novel approach to identify optimal power management solutions, that addresses all these limitations. Specifically, using advanced battery and DC-DC converter models, we propose methods to explore the scheduling space both statically (at design time) and dynamically (at runtime), accounting not only for computation tasks, but also for communication and sensing. With this method, we show that the battery lifetime can be increased by as much as 23.36% if an optimal power management strategy is adopted. Yukai Chen, Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Optimal Topology-Aware PV Panel Floorplanning with Hybrid OrientationabstractDespite of being one of the most widespread green energy sources, the efficiency of PV rooftop installations is still repressed by shading and by the absence of a rigorous irradiance-aware placement approach. The goal of this work is to reach optimal energy production via an irregular placement of PV modules, by considering two degrees of freedom: orientation of each PV module and topology. Experimental results will prove the effectiveness of the proposed solution onto two real world case studies, with an increase of power production of up to 40%. Sara Vinco, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | Fundamental Feature Extraction of the Battery Charge Phase from Product DataabstractThe modeling of electrical energy storage systems has received a lot of attention during the last two decades, as a consequence of the remarkable increase of battery-powered devices. However, automated extraction of a cell/pack characteristics from available product data, has been proposed only more recently. Although the automated modeling of a battery discharge performance is already well reported in the literature, extraction of the basic features of the charging phase is less common since most commercial chargers work at the standard constant current-constant voltage (CC-CV) protocol at fixed working conditions. Nevertheless, many rechargeable batteries allow different charging current rates. Therefore, for a full modeling and analysis of the battery behavior during the charging phase, the basic features, like the internal resistance and the open circuit voltage (Rch and VOCch), should be modeled. Unfortunately, only a few data regarding the charging phase are available from time-based plots in datasheets. This work presents a method for modeling the fundamental characteristics of the charging phase of a battery, starting from typical multi-plot time-based charts. Alberto Bocca, Yukai Chen, Alberto Macii, Massimo Poncino |
ISCAS | 4 |
| 2018 | Application-Driven Synthesis of Energy-Efficient Reconfigurable-Precision OperatorsabstractThe increasing performance demands in emerging Internet of Things applications clash with the low energy budgets of end-nodes. Therefore, hardware operators able to reconfigure their computational precision at runtime are increasingly employed in these devices, to obtain good-enough results at minimal energy costs. Among the many methods proposed to implement such operators, Dynamic Voltage and Accuracy Scaling (DVAS) is particularly promising, due to its broad applicability and low overheads. However, a straight-forward application of DVAS conflicts with the optimizations performed by classic EDA algorithms, and does not yield the expected results. In this paper, we propose a novel synthesis algorithm for reconfigurable-precision circuits, that allows to integrate DVAS in a standard implementation flow. Moreover, we show how this algorithm can exploit information about the application, namely on the frequency of usage of each precision, to further reduce the total energy consumption. Applying our method to the popular LeNet neural network for digit recognition, we are able to reduce the energy due to Multiply-And-Accumulate (MAC) operations by 25%, compared to a straight-forward application of DVAS. Daniele Jahier Pagliari, Massimo Poncino |
ISCAS | 2 |
| 2018 | A Compact PV Panel Model for Cyber-Physical Systems in Smart CitiesabstractOne of the ambitious goals of the "Smart city" paradigm is to design zero-energy buildings. Buildings can be considered as connected cyber-physical systems that require the construction of sound methodologies inherited from the Electronic Design Automation (EDA) research. In particular, aiming at autonomous buildings, the effective design of renewable energy sources is a key aspect for which such methodologies have to be developed. In this work, we propose a modeling strategy for the early estimation of the performance of photovoltaic (PV) arrays. Although a plethora of PV panel models there exists, most of these models suffer from accuracy/complexity tradeoffs. On one hand, building fast models forces to ignore either the correlation between temperature and irradiance, or the topology of panels, thus yielding inaccurate estimations. On the other, more accurate models are time consuming and require costly measurements or circuit analysis, that cannot be extracted from the sole datasheet. This paper proposes a compact semi-empirical model, suitable for real time simulation and built solely from information derived from the PV panel datasheet. The model is built by empirically fitting an expression of the panel operating point as a function of both irradiance and temperature, and of the adopted PV system topology. The accuracy and effectiveness of the proposed model have been validated w.r.t. the production traces of the PV systems of a real world industrial building. Sara Vinco, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Massimo Poncino |
ISCAS | 5 |
| 2018 | Battery-Aware Energy Model of Drone Delivery TasksabstractDrones are becoming increasingly popular in the commercial market for various package delivery services. In this scenario, the mostly adopted drones are quad-rotors (i.e., quadcopters). The energy consumed by a drone may become an issue, since it may affect (i) the delivery deadline (quality of service), (ii) the number of packages that can be delivered (throughput) and (iii) the battery lifetime (number of recharging cycles). It is thus fundamental try to find the proper compromise between the energy used to complete the delivery and the speed at which the quadcopter flies to reach the destination. In order to achieve this, we have to consider that the energy required by the drone for completing a given delivery task does not exactly correspond to the energy requested to the battery, since the latter is a non-ideal power supply that is able to deliver power with different efficiencies depending on its state of charge. In this paper, we demonstrate that the proposed battery-aware delivery scheduling algorithm carries more packages than the traditional delivery model with the same battery capacity. Moreover, the battery-aware delivery model is 17% more accurate than the traditional delivery model for the same delivery scheme, which prevents the unexpected drone landing. Donkyu Baek, Yukai Chen, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 6 |
| 2018 | Dynamic Bit-width Reconfiguration for Energy-Efficient Deep Learning HardwareabstractDeep learning models have reached state of the art performance in many machine learning tasks. Benefits in terms of energy, bandwidth, latency, etc., can be obtained by evaluating these models directly within Internet of Things end nodes, rather than in the cloud. This calls for implementations of deep learning tasks that can run in resource limited environments with low energy footprints. Research and industry have recently investigated these aspects, coming up with specialized hardware accelerators for low power deep learning. One effective technique adopted in these devices consists in reducing the bit-width of calculations, exploiting the error resilience of deep learning. However, bit-widths are tipically set statically for a given model, regardless of input data. Unless models are retrained, this solution invariably sacrifices accuracy for energy efficiency. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2018 | LAPSE: Low-Overhead Adaptive Power Saving and Contrast Enhancement for OLEDsabstractOrganic Light Emitting Diode (OLED) display panels are becoming increasingly popular especially in mobile devices; one of the key characteristics of these panels is that their power consumption strongly depends on the displayed image. In this paper we propose LAPSE, a new methodology to concurrently reduce the energy consumed by an OLED display and enhance the contrast of the displayed image, that relies on image-specific pixel-by-pixel transformations. Unlike previous approaches, LAPSE focuses specifically on reducing the overheads required to implement the transformation at runtime. To this end, we propose a transformation that can be executed in real time, either in software, with low time overhead, or in a hardware accelerator with a small area and low energy budget. Despite the significant reduction in complexity, we obtain comparable results to those achieved with more complex approaches in terms of power saving and image quality. Moreover, our method allows to easily explore the full quality-versus-power tradeoff by acting on a few basic parameters; thus, it enables the runtime selection among multiple display quality settings, according to the status of the system. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
IEEE Trans. Image Process. | 3 |
| 2018 | Thermal Management of Batteries Using Supercapacitor Hybrid Architecture With Idle Period Insertion Strategy
Donghwa Shin, Massimo Poncino, Enrico Macii |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | A circuit-equivalent battery model accounting for the dependency on load frequencyabstractCircuit-equivalent battery models are considered defacto standard for modeling and simulation of digital systems due to many practical advantages. In spite of the many variants of models proposed in the literature, none of them accounts for one important feature of the battery dynamics, namely, the dependency on the frequency of current load profile. For a given average current value, current loads with different spectral distributions may have quite different impacts on the battery discharge. This is a very well-know issue in the design of hybrid energy storage systems, where different types of storages devices are used, each with different storage efficiency for different load frequency ranges. We propose a basic modification to a state-of-the-art model that incorporates this load frequency dependency, as well as a methodology to identify the frequency-sensitive parameters of the model from publicly available data (e.g., datasheets). The results show that frequency-agnostic models can significantly overestimate the battery state-of-charge, and that this effect is far from being negligible. Yukai Chen, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2017 | A methodology for the design of dynamic accuracy operators by runtime back biasabstractMobile and IoT applications must balance increasing processing demands with limited power and cost budgets. Approximate computing achieves this goal leveraging the error tolerance features common in many emerging applications to reduce power consumption. In particular, adequate (i.e., energy/quality-configurable) hardware operators are key components in an error tolerant system. Existing implementations of these operators require significant architectural modifications, hence they are often design-specific and tend to have large overheads compared to accurate units. In this paper, we propose a methodology to design adequate data-path operators in an automatic way, which uses threshold voltage scaling as a knob to dynamically control the power/accuracy tradeoff. The method overcomes the limitations of previous solutions based on supply voltage scaling, in that it introduces lower overheads and it allows fine-grain regulation of this tradeoff. We demonstrate our approach on a state-of-the-art 28nm FDSOI technology, exploiting the strong effect of back biasing on threshold voltage. Results show a power consumption reduction of as much as 39% compared to solutions based only on supply voltage scaling, at iso-accuracy. Daniele Jahier Pagliari, Yves Durand, David Coriat, Anca Mariana Molnos, Edith Beigné, Enrico Macii, Massimo Poncino |
DATE | 7 |
| 2017 | Workload-driven frequency-aware battery sizingabstractDespite the wide body of literature on the sizing of energy storage devices available in the domain of electrical energy systems, the problem has not drawn much attention in the area of battery-powered electronic systems. It is well-known that the straightforward method of sizing battery as the product of an expected duration and the average load current always underestimates the actual capacity that the battery can supply. The variability of the workload and of its spectral distribution will in fact affect the effective capacity of battery that cannot be ignored. This paper proposed a methodology to compute the required capacity of a battery based on the properties of the workload; in particular it accounts for both the impact of the distribution of the current load and of its frequencies, and determines corrective factors for both effects to be used for the calculation of the actual capacity. We used a frequency-sensitive circuit-equivalent battery model to validates our method on three synthetic and two real workloads. Simulation results show that even for workload with same average current, the required capacity can be as much as 70% larger than the capacity estimated using a traditional method. Yukai Chen, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2017 | A Layered Methodology for the Simulation of Extra-Functional Properties in Smart SystemsabstractSmart systems represent a broad class of intelligent, miniaturized devices incorporating functionality like sensing, actuation, and control. In order to support these functions, they must include sophisticated and heterogeneous components, such as sensors and actuators, multiple power sources and storage devices, digital signal processing, and wireless connectivity. The high degree of heterogeneity typical of smart systems has a heavy impact on their design: the challenges are not in fact restricted to their functionality, but are also related to a number of extra-functional properties, including power consumption, temperature, and aging. Current simulation- or model-based design approaches do not target a smart system as a whole, but rather single domains (digital, analog, power devices, etc.) or properties. This paper tries to overcome this limitation by proposing a framework for the concurrent simulation of both functionality and such extra-functional properties. The latter are modeled as different information flows, managed by dedicated “virtual buses” and formalized through the adoption of IP-XACT. SystemC, through the support of physical and continuous time modeling provided by its analog and mixed signal extension, is used to implement both functional and extra-functional models. Experimental results show the efficiency, accuracy and modularity of the proposed approach on an example case study, in which substantial speedups with respect to standard model-based design tools go along with a very high degree of accuracy (-5%). Furthermore, the case study highlights that the proposed framework allows to easily capture at run time the mutual impact of properties, e.g., in case of power and temperature. Sara Vinco, Yukai Chen, Franco Fummi, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Approximate Energy-Efficient Encoding for Serial InterfacesabstractSerial buses are ubiquitous interconnections in embedded computing systems that are used to interface processing elements with peripherals, such as sensors, actuators, and I/O controllers. Despite their limited wiring, as off-chip connections they can account for a significant amount of the total power consumption of a system-on-chip device. Encoding the information sent on these buses is the most intuitive and affordable way to reduce their power contribution; moreover, the encoding can be made even more effective by exploiting the fact that many embedded applications can tolerate intermediate approximations without a significant impact on the final quality of results, thus trading off accuracy for power consumption. We propose a simple yet very effective approximate encoding for reducing dynamic energy in serial buses. Our approach uses differential encoding as a baseline scheme and extends it with bounded approximations to overcome the intrinsic limitations of differential encoding for data with low temporal correlation. We show that the proposed scheme, in addition to yielding extremely compact codecs, is superior to all state-of-the-art approximate serial encodings over a wide set of traces representing data received or sent from/to sensor or actuators. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Serial T0: approximate bus encoding for energy-efficient transmission of sensor signalsabstractOff-chip serial buses are common in embedded systems, and due to the long physical lines, can contribute significantly to their energy consumption. However, these buses are often connected to analog sensors, whose data is inherently affected by noise and A/D errors. Thus, communication can tolerate small approximations, without a significant impact on the system outputs quality. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
DAC | 3 |
| 2016 | Low-overhead adaptive constrast enhancement and power reduction for OLEDs
Daniele Jahier Pagliari, Massimo Poncino, Enrico Macii |
DATE | 2 |
| 2016 | CONTREX: Design of Embedded Mixed-Criticality CONTRol Systems under Consideration of EXtra-Functional PropertiesabstractThe increasing processing power of today's HW/SW platforms leads to the integration of more and more functions in a single device. Additional design challenges arise when these functions share computing resources and belong to different criticality levels. The paper presents the CONTREX European project and its preliminary results. CONTREX complements current activities in the area of predictable computing platforms and segregation mechanisms with techniques to consider the extra-functional properties, i.e., timing constraints, power, and temperature. CONTREX enables energy efficient and cost aware design through analysis and optimization of these properties with regard to application demands at different criticality levels. Ralph Görgen, Kim Grüttner, Fernando Herrera, Pablo Peñil, Julio L. Medina, Eugenio Villar, Gianluca Palermo, William Fornaciari, Carlo Brandolese, Davide Gadioli, Sara Bocchio, Luca Ceva, Paolo Azzoni, Massimo Poncino, Sara Vinco, Enrico Macii, Salvatore Cusenza, John M. Favaro, Raúl Valencia, Ingo Sander, Kathrin Rosvall, Davide Quaglia |
DSD | 14 |
| 2016 | IP-XACT for smart systems design: extensions for the integration of functional and extra-functional modelsabstractSmart systems are miniaturized devices integrating computation, communication, sensing and actuation. As such, their design can not focus solely on functional behavior, but it must rather take into account different extra-functional concerns, such as power consumption or reliability. Any smart system can thus be modeled through a number of views, each focusing on a specific concern. Such views may exchange information, and they must thus be simulated simultaneously to reproduce mutual influence of the corresponding concerns. This paper shows how the IP-XACT standard, with some necessary extensions, can effectively support this simultaneous simulation. The extended IP-XACT descriptions allow to model extra-functional properties with a homogeneous format, defined by analysing requirements and characteristic of three main concerns, i.e., power, temperature and reliability. The IP-XACT descriptions are then used to automatically generate a skeleton of the simulation infrastructure in SystemC. The skeleton can be easily populated with models available in the literature, thus reaching simultaneous simulation of multiple concerns. Sara Vinco, Michele Lora, Enrico Macii, Massimo Poncino |
FDL | 4 |
| 2016 | Fast Thermal Simulation using SystemC-AMSabstractOut of the many options available for thermal simulation of digital electronic systems, those based on solving an RC equivalent circuit of the thermal network are the most popular choice in the EDA community, as they provide a reasonable tradeoff between accuracy and complexity. HotSpot, in particular, has become the de-facto standard in these communities, although other simulators are also popular. These tools have many benefits, but they are relatively inefficient when performing thermal analysis for long simulation times, due to the occurrence of a large number of redundant computations intrinsic in the underlying models. Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | Approximate Differential Encoding for Energy-Efficient Serial CommunicationabstractEmbedded computing systems include several off-chip serial links, that are typically used to interface processing elements with peripherals, such as sensors, actuators and I/O controllers. Because of the long physical lines of these connections, they can contribute significantly to the total energy consumption. On the other hand, many embedded applications are error resilient, i.e. they can tolerate intermediate approximations without a significant impact on the final quality of results. This feature can be exploited in serial buses to explore the trade-off between data approximations and energy consumption. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | Graphene-PLA (GPLA): a Compact and Ultra-Low Power Logic Array ArchitectureabstractThe key characteristics of the next generation of ICs for wearable applications include high integration density, small area, low power consumption, high energy-efficiency, reliability and enhanced mechanical properties like stretchability and transparency. The proper mix of new materials and novel integration strategies is the enabling factor to achieve those design specifications. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | A Unified Model of Power Sources for the Simulation of Electrical Energy SystemsabstractModels of power sources are essential elements in the simulation of systems that generate, store and manage energy. In spite of the huge difference in power scale, they perform a common function: converting a primary environmental quantity into power. This paper proposes a unified model of a power source that is applicable to any power scale, and that can be derived solely from data contained in the specification or the datasheet of a device. The key feature of our model is the normalization of the energy generation characteristic of the power source by means of a reduction to a function expressing extracted power vs. the "scavenged" quantity. The proposed model proved to apply to two kinds of power sources, i.e., a wind turbine and a photovoltaic panel, and to provide a good level of accuracy and simulation performance w.r.t. widely adopted models. Sara Vinco, Yukai Chen, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | Enabling quasi-adiabatic logic arrays for silicon and beyond-silicon technologiesabstractAdiabatic logic aims at mimicking an adiabatic (i.e., without energy exchange) charging process in digital circuits. Although regarded as a mostly theoretical computation style, research on the topic has been constantly active over the years, providing several demonstrations of working implementations [1]. The interest in adiabatic circuits recently increased with the introduction of emerging devices, e.g., Nanoelectromechanicals switches (NEMs) [2] and graphene p-n junctions [3], which have been proven to be good technological vehicles for adiabatic computing. Despite their energy efficiency, adiabatic logic faced severe limitations in reaching large scale integration due to the difficulty in logic pipelining and the lack of CAD tools able to cope with today's design complexity. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
ISCAS | 4 |
| 2016 | A Li-Ion Battery Charge Protocol with Optimal Aging-Quality of Service Trade-offabstractThe reduction of usable capacity of rechargeable batteries can be mitigated during the charge process by acting on some stress factors, namely, the average state-of-charge (SOC) and the charge current. Larger values of these quantities cause an increased degradation of battery capacity, so it would be desirable to keep both as low as possible, which is obviously in contrast with the objective of a fast charge. However, by exploiting the fact that in most battery-powered systems the time during which it is plugged for charging largely exceeds the time required to charge, it is possible to devise appropriate charge protocols that achieve a good balance between fast charge and aging. Yukai Chen, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 5 |
| 2016 | Frequency domain characterization of batteries for the design of energy storage subsystemsabstractThe Ragone chart is a pictorial representation to express the well-known the trade-off between available energy vs. power of different classes of energy storage devices (ESDs) like batteries or supercapacitors. Ragone charts, however, do not normally provide information about individual devices, which is an essential requirement for the actual design of the energy storage sub-system. Yukai Chen, Enrico Macii, Massimo Poncino |
VLSI-SoC | 3 |
| 2016 | Multi-function logic synthesis of silicon and beyond-silicon ultra-low power pass-gates circuitsabstractPass-gates logic is known to be intrinsically more energy efficient than static CMOS. This feature attracted the research interest over the years and many working implementations have been demonstrated. Recent works, in particular, have shown that pass-gates logic is well suited for ultra-low power adiabatic circuits mapped on emerging technologies. Despite the progress made, several design issues still prevent pass-gates logic circuits reaching large scale integration. In this work we deal with the lack of synthesis tools and methodologies. We propose a multi-function decomposition engine that yields (i) an efficient abstract circuit modeling through a more compact data-structure, the Multi-Function Pass Diagram (MFPD) and (ii) an effective multi-gate area/delay-driven low-power synthesis&optimization flow. Simulation results conducted on different technologies, i.e., silicon and graphene, demonstrate that logic circuits synthesized with the proposed tool are smaller in size and depth, hence less power consuming and faster than circuits obtained through conventional synthesis flows based on Binary Decision Diagrams. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
VLSI-SoC | 4 |
| 2015 | One-pass logic synthesis for graphene-based Pass-XNOR logic circuitsabstractElectrostatically controlled graphene P-N junctions are devices built on a single layer graphene sheet that can be turned-ON/OFF via external potential difference. Their electrical behavior resembles a CMOS transmission gate with an embedded XNOR Boolean functionality. Recent works presented an efficient design style, the Pass-XNOR logic (PXL), which allows the implementation of adiabatic logic circuits with ultra low-power features. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
DAC | 4 |
| 2015 | Characterizing the Activity Factor in NBTI Aging Models for Embedded CoresabstractIn deeply scaled CMOS technologies, device aging causes cores performance parameters to degrade over time. While accurate models to efficiently assess these degradation exist for devices and circuits, no reliable model for processor cores has gained strong acceptance in the literature. In this work, we propose a methodology for deriving an NBTI aging model for embedded cores. Based on an accurate characterization on the netlist of the core, we were able to (1) prove the independence of the aging on the workload (i.e., executed instructions), and (2) calculate an equivalent average constant aging factor that justifies the use of the baseline model template. We derived and assessed the proposed model by using a RISC-like processor core implemented in a 45nm process technology as a reference architecture, achieving a maximum error of 2.2% against simulated data on the core netlist. Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | Exploiting the Expressive Power of Graphene Reconfigurable Gates via Post-Synthesis OptimizationabstractAs an answer to the new electronics market demands, semiconductor industry is looking for different materials, new process technologies and alternative design solutions that can support Silicon replacement in the VLSI domain. The recent introduction of graphene, together with the option of electrostatically controlling its doping profile, has shown a possible way to implement fast and power efficient Reconfigurable Gates (RGs). Also, and this is the most important feature considered in this work, those graphene RGs show higher expressive power, i.e., they implement more complex functions, like Majority, MUX, XOR, with less area w.r.t. CMOS counterparts. Unfortunately, state-of-the-art synthesis tools, which have been customized for standard NAND/NOR CMOS gates, do not exploit the aforementioned feature of graphene RGs. Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino, Luca G. Amarù, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
ACM Great Lakes Symposium on VLSI | 5 |
| 2015 | Design and Characterization of Analog-to-Digital Converters using Graphene P-N JunctionsabstractElectrostatically controlled graphene p-n junctions are devices built on single-layer graphene sheets whose in-to-out resistance can be dynamically tuned through external voltage potentials. Roberto Giorgio Rizzo, Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 5 |
| 2015 | An aging-aware battery charge scheme for mobile devices exploiting plug-in time patternsabstractThe aging of a rechargeable battery is mainly due to stress during charge-discharge cycles. Although the discharge phase is difficult to control, the charging phase can be performed in a specific way in order to mitigate the aging of the battery during its usage. It therefore becomes important to select the correct charging algorithm. In the case of mobile systems, equipped mainly with lithiumion batteries, the standard widely adopted for charging a battery is the typical constant current/constant voltage (CC-CV) protocol usually based on a linearly regular charge process. In this work, we propose a charging protocol based on the standard CC-CV method in which the charge start time and the value of the charging current can be programmed in such a way that the aging of the battery is mitigated. To validate this charging scheme we use an aging model that includes the charge/discharge current among the major parameters, and an analytical macro-model for the CC-CV charge time analysis. Alberto Bocca, Alessandro Sassone, Alberto Macii, Enrico Macii, Massimo Poncino |
ICCD | 5 |
| 2015 | An automated design flow for approximate circuits based on reduced precision redundancyabstractReduced Precision Redundancy (RPR) is a popular Approximate Computing technique, in which a circuit operated in Voltage Over-Scaling (VOS) is paired to a reduced-bitwidth and faster replica so that VOS-induced timing errors are partially recovered by the replica, and their impact is mitigated. Previous works have provided various examples of effective implementations of RPR, which however suffer from three limitations: first, these circuits are designed using ad-hoc procedures, and no generalization is provided; second, error impact analysis is carried out statistically, thus neglecting issues like non-elementary data distribution and temporal correlation. Last, only dynamic power was considered in the optimization. In this work we propose a new generalized approach to RPR that allows to overcome all these limitations, leveraging the capabilities of state-of-the-art synthesis and simulation tools. By sacrificing theoretical provability in favor of an empirical input-based analysis, we build a design tool able to automatically add RPR to a preexisting gate-level netlist. Thanks to this method, we are able to confute some of the conclusions drawn in previous works, in particular those related to statistical assumptions on inputs; we show that a given inputs distribution may yield extremely different results depending on their temporal behavior. Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino |
ICCD | 4 |
| 2015 | An equation-based battery cycle life model for various battery chemistriesabstractThe evaluation of the cycle life of batteries is an essential task in the assessment of the reliability and cost of battery-operated devices. Several compact cycle life models have been proposed in the literature, that exhibit a general trade-off between generality and accuracy. Some models are based on a compact equation derived from experimental data and try to extract a general relationship between cycle life and the relevant parameters (mostly the depth of discharge), but suffer from poor accuracy. At the other extreme, more accurate models, based on incorporating the aging effect into an equivalent circuit, tend to be focused on a specific device and are seldom applicable to another battery. In this work we propose an equation-based model that tries to overcome the accuracy limits of previous similar models. The model parameters are obtained by fitting the curve based on information reported in datasheets, and can be adapted (with different accuracy levels) to the amount of available information. We applied the model to various commercial batteries for which full information on their cycle life is available. Results show an average estimation error, in terms of the number of cycles, generally smaller than 10%, which is consistent with the typical tolerance provided in the datasheets, and much lower than previous equation-based models. Alberto Bocca, Alessandro Sassone, Donghwa Shin, Alberto Macii, Enrico Macii, Massimo Poncino |
VLSI-SoC | 6 |
| 2015 | A Statistical Model-Based Cell-to-Cell Variability Management of Li-ion Battery PackabstractThe cell-to-cell variability of batteries is a well-known problem particularly when it comes to the assembly of large battery packs. Different battery cells exhibit substantial variability due to manufacturing tolerances, which should be assessed and managed carefully. Such variability has been approached mostly from the point of view of the chemical and physical phenomena, but these solutions are normally too complicated for the system-level design of electric applications. This paper proposes a combined cell-to-cell variability model of the capacity and internal resistance of a Li-ion battery that accounts for the variability effects in the cell manufacturing process. The proposed model allows to verify some known properties, such as the correlation between the capacity and internal resistance, to be verified qualitatively and the amount of variability and its impact on the design of battery packs to be assessed quantitatively. Using this model, the issue of how to consider the variability when constructing battery packs was also addressed. Modern battery packs normally incorporate some cell balancing circuitry, which is meant to balance cell voltages during charging at the expense of a bypassed (unstored) charge. For discharge, the cell-to-cell variability hides a part of the usable capacity of the battery pack. This paper proposes the use of variability information to assemble battery packs with minimal intracolumn variance of capacity. A weight-based variance minimization method, based on the correlation between cell capacity and weight is proposed to avoid resorting to direct battery capacity measurements, which is time-consuming and requires costly measurement equipment. The simulation result shows that the proposed weight-based approach allows an acceptable management of the cell-to-cell variability without the discharging experiment. Donghwa Shin, Massimo Poncino, Enrico Macii, Naehyuck Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Statistical Battery Models and Variation-Aware Battery ManagementabstractCell-to-cell variability of batteries is a well-known problem especially when it comes to assembling large battery packs. Different battery cells exhibit substantial variability among them due to manufacturing tolerances, which should be carefully assessed and managed. Although battery packs usually incorporate some cell balancing circuitry, it is supposed to balance cell voltages dynamically at the expense of bypassed (not stored) charge. Donghwa Shin, Enrico Macii, Massimo Poncino |
DAC | 3 |
| 2014 | A cross-level verification methodology for digital IPs augmented with embedded timing monitorsabstractSmart systems implement the leading technology advances in the context of embedded devices. Current design methodologies are not suitable to deal with tightly interacting subsystems of different technological domains, namely analog, digital, discrete and power devices, MEMS and power sources. The effects of interaction between components and with the environment must be modeled and simulated at system level to achieve high performance. Focusing on the digital domain, additional design constraints have to be considered as a result of the integration of multi-domain subsystems in a single device. The main digital design challenges, combined with those emerging from the heterogeneous nature of the whole system, directly impact on performance and on propagation delay of the digital component. This paper proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems, and a methodology to verify the added features at system-level. The design approach consists of augmenting the RTL model through the automatic insertion of delay sensors, which can detect and correct timing failures. The augmented model is abstracted to SystemC TLM and, then, mutants (i.e., code mutations for emulating timing failures) are automatically injected into the model. Experimental results demonstrate the applicability of the proposed design and verification methodology and the effectiveness of the simulation performance. Valerio Guarnieri, Massimo Petricca, Alessandro Sassone, Sara Vinco, Nicola Bombieri, Franco Fummi, Enrico Macii, Massimo Poncino |
DATE | 8 |
| 2014 | Cache aging reduction with improved performance using dynamically re-sizable cacheabstractAging of transistors is a limiting factor for long term reliability of devices in sub-100nm technologies. It's a worst-case metric where the lifetime of a device is determined by the earliest failing component. Impact is more serious on memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. Previous works have shown that partitioning based strategies based on power management techniques can effectively control aging effects and can extend lifetime of the cache significantly. However, such a benefit comes as a tradeoff with performance which reduces proportionally as the time elapses. To address this problem and provide a single solution to concurrently improve aging, energy and performance of the cache, we propose an architectural solution based on the dynamically re-sizable cache and cache partitioning approaches. By this strategy, cache is dynamically re-sized and reconfigured whenever a cache block becomes unreliable. Coupling such aging mitigation technique along with dynamically re-sizable cache approach provides on average 30% lifetime improvement with less than 0.4x degradation in performance whereas, in previous solutions, performance degradation sometimes goes upto 10x. Haroon Mahmood, Massimo Poncino, Enrico Macii |
DATE | 2 |
| 2014 | Thermal management of batteries using a hybrid supercapacitor architectureabstractThermal analysis and management of batteries have been an important research issue for battery-operated systems such as electric vehicles and mobile devices. Nowadays, battery packs are designed considering heat dissipation, and external cooling devices such as a cooling fan are also widely used to enforce the reliability and extend the lifetime of a battery. This type of approaches that target the enhancement of the cooling efficiency via the reduction of the thermal resistance cannot achieve an immediate temperature drop to avoid a thermal emergency situation. Approaches based on removing the heat from the heat sources via idle period insertion (similar to what is done for silicon devices) would allow faster thermal response; however it is not obvious how to implement these schemes in the context of batteries. In this paper, we propose the use of a simple parallel battery-supercapacitor hybrid architecture with a dual-mode discharging strategy that can provide immediate temperature management, in which the supercapacitor is used as an energy buffer during the idle periods of the battery. Simulation results shows that the proposed method can keep the battery temperature within the safe range without external cooling devices while exploiting the advantage of the battery-supercapacitor parallel connection. Donghwa Shin, Massimo Poncino, Enrico Macii |
DATE | 2 |
| 2014 | Pass-XNOR logic: A new logic style for P-N junction based graphene circuitsabstractIn this work we introduce a new logic style for p-n junctions based digital graphene circuits: the pass-XNOR logic style. The latter enables the realization of compact, energy efficient circuits that better exploit the characteristics of graphene. We first show how a single p-n junction can be conceived as a pass-XNOR gate, i.e., a transmission gate with embedded logic functionality, the XNOR Boolean operator. Secondly, we propose a smart integration strategy in which series/parallel connections of pass-XNOR gates allow to implement AND/OR logical conjunctions, and, therefore, all possible truth tables. Experimental results conducted on a set of representative logic functions show the superior of pass-XNOR logic circuits w.r.t. standard CMOS circuits and graphene circuits that use p-n junctions in a complementary-like structure. Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2014 | Ultra Low-Power Computation via Graphene-Based Adiabatic Logic GatesabstractAs an answer to the difficulties in improving the figures of merit of deeply-scaled CMOS devices, researchers have looked at alternative materials and technologies for implementing new devices that can overcome the limitations of CMOS. Graphene has emerged as one of the most promising candidates among these new materials, recent works have demonstrated the implementation of electrostatically-controlled p-n junctions that can serve as the basic primitive for a new class of compact, fast and energy-efficient graphene-based logic gates. In this work we revisit those gates from a different perspective, namely, as devices that can operate adiabatically, that is, that are able to reuse the dissipated energy. We show how to build the basic logic gates by appropriately interconnecting graphene based p-n junctions and characterize those adiabatic gates for power and performance. The comparison between these adiabatic gates and both adiabatic CMOS and their non-adiabatic graphene-based counterpart shows that the former can operate with 2X to 3X less average power, and about 4X better power-delay product. Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
DSD | 4 |
| 2014 | Modeling of the charging behavior of li-ion batteries based on manufacturer's dataabstractThe market of portable devices, wireless sensors, electric vehicles and storage systems has grown enormously in recent years. As a consequence, batteries and related technologies have become one of the major topics for researchers. Due to the large variety of applications in which batteries are involved, battery modeling is becoming an extremely important research topic. This relevance is witnessed by the number of papers addressing battery modeling. Alessandro Sassone, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 6 |
| 2014 | Automated generation of battery aging models from datasheetsabstractThe de-facto standard approach in battery modeling consists of the definition of a generic model template in terms of an equivalent electric circuit, which is then populated either using data obtained from direct measurements on actual devices or by some extrapolation of battery characteristics available from datasheets. These models typically describe only intra-cycle effects, that is, those manifesting within a single charge/discharge cycle of a battery. However, basic battery dynamics, during a single discharge, cannot provide a true estimate of the actual lifetime of the battery, e.g., how its usability decreases due to long-term and irreversible effects, such as the fading of capacity due to aging or to repeated cycling. While some solutions in the literature provide answers to this problem by proposing suitable models for these effects, they do not provide solutions for how to incorporate them into a generic model template. In this work we propose a method to include inter-cycle battery effects into a reference model template in an automated way, and using solely data reported by battery manufacturers. Flexibility and accuracy of the proposed strategy are demonstrated by modeling a commercial lithium iron phosphate battery, whose datasheet provides long-term capacity fading information. Massimo Petricca, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ICCD | 6 |
| 2014 | A compact macromodel for the charge phase of a battery with typical charging protocolabstractAvailability of a simulation model of a battery is one of the most important requisites in the system-level design of battery-powered systems. The vast majority of the models describe the discharge behavior of the battery; so far, the estimation of charging time has been in fact only marginally studied because the charging phase is regarded as a relatively controlled process compared to discharge. In this paper, we present a compact macro-model for the estimation of charging time under the most widely used charge protocol, i.e., Constant Current-Constant Voltage (CC-CV). This model is derived under the consideration of the context of the existing models including the well-known Peukert's law and equivalent electric circuits. The estimation result with the proposed model based on the manufacturer's data of commercial Li-ion batteries shows fair accuracy, especially when compared to estimates on parameters extracted from discharge characteristics. Donghwa Shin, Alessandro Sassone, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 6 |
| 2014 | An open-source framework for formal specification and simulation of electrical energy systemsabstractElectrical energy systems (EESs) are systems which consume, generate, distribute and store energy at various scales. This paper presents a modeling and simulation framework that uses principles borrowed from the system-level simulation of digital systems and extends them to the case of EESs. The framework relies on open-source standards such as SystemC (and its Analog and Mixed-Signal extensions) for simulation, and IP-XACT for interface definition. Sara Vinco, Alessandro Sassone, Franco Fummi, Enrico Macii, Massimo Poncino |
ISLPED | 5 |
| 2014 | Modeling of Physical Defects in PN Junction Based Graphene Devices
Sandeep Miryala, Matheus Oleiro, Letícia Maria Veiras Bolzani, Andrea Calimera, Enrico Macii, Massimo Poncino |
J. Electron. Test. | 6 |
| 2014 | Dynamic Indexing: Leakage-Aging Co-Optimization for CachesabstractTraditional implementations of low-power states based on voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused by negative bias temperature instability (NBTI), which can be explained in terms of the intuitive correlation between the idleness and the reduced workload of a system. Such a joint benefit has been exploited only partially because of the different nature of energy and aging as cost functions: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually result in limited aging reductions. In this paper, we address this problem in the context of power-managed caches, which represent a critical target for NBTI-reduced aging: given their symmetric structure, SRAM structures are, in particular, sensitive to NBTI effects because they cannot take advantage of the value-dependent recovery typical of NBTI. We propose a strategy called dynamic indexing, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the various power managed units (e.g., lines). This distribution allows fully using the leakage optimization potential and extending the lifetime of a cache. We explore various alternatives, in particular different granularities of the power managed units as well as different reindexing functions. Experimental analysis shows that it is possible to simultaneously reduce leakage power and aging in caches, with minimal power consumption overhead. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Energy/Lifetime Cooptimization by Cache Partitioning With Graceful Performance DegradationabstractAging of transistors can adversely impact the long-term reliability of devices in subnanometric technologies. Without any countermeasure, the first component that becomes unreliable will determine the life span of an entire device. The effect is more susceptible in memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. In this paper, we propose a reliability management technique based on the idea of cache partitioning, which deals with cell failures by gracefully degrading its performance. By this partitioning-based strategy, various subblocks will become unreliable at different times, and the cache will keep functioning with reduced efficiency. A coarse-grain implementation of this approach, with the use of a smart aging-driven partitioning algorithm, provides a lifetime extension of more than 2× . On the other hand, a fine-grain strategy with a single cache line as a unit of power management, stretch the lifetime to its maximum limits with an addition of small hardware overhead. Haroon Mahmood, Mirko Loghi, Massimo Poncino, Enrico Macii |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Energy-optimal SRAM supply voltage scheduling under lifetime and error constraintsabstractThis work addresses the energy efficiency of the memory architecture in safety-critical systems that have to guarantee a given level of service and a minimum lifetime. We specifically target SRAM structures in which decreased reliability manifests itself in terms of the aging induced by NBTI (Negative Bias Temperature Instability), and in which the level of service is represented by the bit-error rate (BER). Andrea Calimera, Enrico Macii, Massimo Poncino |
DAC | 3 |
| 2013 | A verilog-a model for reconfigurable logic gates based on graphene pn-junctionsabstractSingle layer sheets of graphene show special electrical properties that can enable the next generation of smart ICs. Recent works have proven the availability of an electrostatically controlled pn-junction upon which it is possible to design multi-function reconfigurable logic devices that naturally behave as multiplexers. In this work we introduce a stable large-signal Verilog-A model that mimics the behavior of the aforementioned devices. The proposed model, validated through the SPICE characterization of a MUX-based standard cell library we designed as benchmark, represents a first step towards the integration of Electronic Design Automation tools that can support the design of all-graphene ICs. Sandeep Miryala, Mehrdad Montazeri, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 5 |
| 2013 | SMAC: Smart Systems Co-designabstractIn this paper we present the concepts and the organization of the FP7 Project SMAC (Smart systems Co-design), an Integrated Project (IP) of the 7th ICT Call under the Objective 3.2 "Smart components and Smart Systems integration". We describe in particular the project objectives and its organization, and how it addresses the challenges of the integration of heterogeneous and conflicting domains that emerge in the design of smart systems. The main outcome of the SMAC project is the development of flexible software platform (the SMAC platform) for smart subsystems/components design include methodologies and EDA tools enabling multi-disciplinary and multi-scale modeling and design, simulation of multi-domain systems, subsystems and components at all levels of abstraction, system integration and exploration for optimization of functional and non-functional metrics. Nicola Bombieri, Giuliana Drogoudis, Giuliana Gangemi, Renaud Gillon, Enrico Macii, Massimo Poncino, Salvatore Rinaudo, Francesco Stefanni, Dimitrios Trachanis, Mark van Helvoort |
DSD | 6 |
| 2013 | Delay model for reconfigurable logic gates based on graphene PN-junctionsabstractIn this paper we address the problem of modeling the timing behavior of a new class of reconfigurable logic gates based on electrostatically controlled graphene pn-junctions. These gates naturally behave as a 2-to-1 multiplexer in which the polarity of the input select line can be dynamically reconfigured. Interconnection of multiple gates and proper assignments of the inputs signals allow to implement all the basic Boolean logic functions, and, at a larger scale, any digital circuit. Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2013 | Computer-aided design of electrical energy systemsabstractElectrical energy systems (EESs) include energy generation, distribution, storage, and consumption, and involve many diverse components and sub-systems to implement these tasks. This paper represents a first step towards the computer-aided design for EESs, encompassing modeling, simulation, design and optimization of these systems. CAD for EESs is a challenging task that mandates a multidisciplinary and heterogeneous approach. We identify similarities and differences between electrical energy systems and electronics systems in order to inherit as much as possible the profound legacy resources of electronic design automation (EDA). We introduce fundamental concepts, from the general problem formulation to the development and deployment of efficient, scalable, and versatile CAD and EDA methods and framework for the optimal or near-optimal EESs. Younghyun Kim 0001, Donghwa Shin, Massimo Petricca, Sangyoung Park, Massimo Poncino, Naehyuck Chang |
ICCAD | 5 |
| 2013 | An automated framework for generating variable-accuracy battery models from datasheet informationabstractModels based on an electrical circuit equivalent have become the most popular choice for modeling the behavior of batteries, thanks to their ease of co-simulation with other parts of a digital system. Such circuit models are actually model templates: the specific values of their electrical elements must be derived by the analysis of the specific battery devices to be modeled. This process requires either to measure the battery characteristics or to derive them from the datasheet. In the latter case, however, very often not all information are available and the model fitting becomes then unfeasible. In this paper we present a methodology for deriving, in a semi-automatic way, circuit equivalent battery models solely from data available in a battery datasheet. In order to account for the different amount of information available, we introduce the concept of “level” of a model, so that models with different accuracy can be derived depending on the available data. The methodology requires only minimal intervention by the designer and it automatically generates MATLAB models once the required data for the corresponding model level are transcribed from the datasheet. Simulation results show that our methodology allows to accurately reconstruct the information reported in the datasheet as well as to derive missing ones. Massimo Petricca, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 6 |
| 2013 | A statistical model of cell-to-cell variation in Li-ion batteries for system-level designabstractDue to manufacturing tolerances, different battery cells exhibit substantial variability among them, which should be carefully assessed and managed, especially when assembling large battery packs. Cell-to-cell variability has been mostly approached from the point of view of the chemical and physical phenomena, but these studies did not provide a practical solution for the system-level design. In this work, we propose a combined cell-to-cell variation model of the capacity and of the internal resistance of a battery cell that accounts for variability effects in the cell manufacturing process. The model is derived from analytical models for a specific type of Li-ion cell provided in the literature, from which we identify what model parameters can be regarded as true random variables. This allows transforming capacity and internal resistance into the functions of random variables, which can be incorporated into an equivalent circuit model that is suitable for the system-level statistical simulations. The proposed model allows us to qualitatively verify some known properties such as the correlation between capacity and internal resistance, and quantitatively assess the amount of variability and its impact on the design of battery packs. Donghwa Shin, Massimo Poncino, Enrico Macii, Naehyuck Chang |
ISLPED | 2 |
| 2013 | Layout-Driven Post-Placement Techniques for Temperature Reduction and Thermal Gradient MinimizationabstractWith the continuing scaling of CMOS technology, on-chip temperature and thermal-induced variations have become a major design concern. To effectively limit the high temperature in a chip equipped with a cost-effective cooling system, thermal specific approaches, besides low power techniques, are necessary at the chip design level. The high temperature in hotspots and large thermal gradients are caused by the high local power density and the nonuniform power dissipation across the chip. With the objective of reducing power density in hotspots, we propose two placement techniques that spread cells in hotspots over a larger area. Increasing the area occupied by the hotspot directly reduces its power density, leading to a reduction in peak temperature and thermal gradient. To minimize the introduced overhead in delay and dynamic power, we maintain the relative positions of the coupling cells in the new layout. We compare the proposed methods in terms of temperature reduction, timing, and area overhead to the baseline method, which enlarges the circuit area uniformly. The experimental results showed that our methods achieve a larger reduction in both peak temperature and thermal gradient than the baseline method. The baseline method, although reducing peak temperature in most cases, has little impact on thermal gradient. Wei Liu 0016, Andrea Calimera, Alberto Macii, Enrico Macii, Alberto Nannarelli, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2012 | Application-specific memory partitioning for joint energy and lifetime optimizationabstractPower management of caches based on turning idle cache lines into a low-energy state is also beneficial for the aging effects caused by Negative Bias Temperature Instability (NBTI), provided that idleness is correctly exploited; unlike energy, aging, being a measure of delay, is in fact a worst-case metric. Haroon Mahmood, Massimo Poncino, Mirko Loghi, Enrico Macii |
DATE | 2 |
| 2012 | IR-drop analysis of graphene-based power distribution networksabstractElectromigration (EM) has been indicated as the killer effect for copper interconnects. ITRS projections show that for future technologies (22nm and beyond) the on-chip current demand will exceed the physical limit copper metal wires can tolerate. This represents a serious limitation for the design of power distribution networks of next generation ICs. New carbon nanomaterials, governed by ballistic transport, have shown higher immunity to EM, thereby representing potential candidate to replace copper. In this paper we make use of compact conductance models to benchmark Graphene Nanoribbons (GNRs) against copper. The two materials have been used to route a state-of-the-art multi-level power-grid architecture obtained through an industrial 45nm physical design flow. Although the adopted design style is optimized for metal grids, results obtained using our simulation framework show that GNRs, if properly sized, can outperform copper, thus allowing the design of reliable circuits with reduced IR-drop penalties. Sandeep Miryala, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2012 | Investigating the effects of Inverted Temperature Dependence (ITD) on clock distribution networksabstractThe aggressive scaling of CMOS technology toward nanometer lengths contributed to the surfacing of many effects that were not appreciable at the micrometer regime. Among them, Inverted Temperature Dependence (ITD) is certainly the most unusual. It manifests itself as a speed up of CMOS gates when the temperature increases, resulting in a reversal of the worst-case condition, i.e., CMOS gates show the largest delay at low temperatures. On the other hand, for metal interconnects an high temperature still holds as worst case condition. The two contrasting behaviors may invalidate the results obtained through standard design flow which do not consider temperature as an explicit variable in their optimizations. In this paper we focus on the impact of ITD on clock distribution networks (CDN), whose function is vital to guarantee the synchronization among physically spaced sequential components of digital circuits. Using our simulation framework, we characterized the thermal behavior of a clock tree mapped onto an industrial 65nm CMOS technology and obtained using a standard synthesis tool. Results demonstrate the presence of ITD at low operating voltages and open new potential research scenarios into the EDA field. Alessandro Sassone, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino, Richard Goldman, Vazgen Melikyan, Eduard Babayan, Salvatore Rinaudo |
DATE | 5 |
| 2012 | Multiple-source and multiple-destination charge migration in hybrid electrical energy storage systemsabstractHybrid electrical energy storage (HEES) systems consist of multiple banks of heterogeneous electrical energy storage (EES) elements that are connected to each other through the Charge Transfer Interconnect. A HEES system is capable of providing an electrical energy storage means with very high performance by taking advantage of the strengths (while hiding the weaknesses) of individual EES elements used in the system. Charge migration is an operation by which electrical energy is transferred from a group of source EES elements to a group of destination EES elements. It is a necessary process to improve the HEES system's storage efficiency and its responsiveness to load demand changes. This paper is the first to formally describe a more general charge migration problem, involving multiple sources and multiple destinations. The multiple-source, multiple-destination charge migration optimization problem is formulated as a nonlinear programming (NLP) problem where the goal is to deliver a fixed amount of energy to the destination banks while maximizing the overall charge migration efficiency and not depleting the available energy resource of the source banks by more than a given percentage. The constraints for the optimization problem are the energy conservation relation and charging current constraints to ensure that charge migration will meet a given deadline. The formulation correctly accounts for the efficiency of chargers, the rate capacity effect of batteries, self-discharge currents and internal resistances of EES elements, as well as the terminal voltage variation of EES elements as a function of their state of charges (SoC's). An efficient algorithm to find a near-optimal migration control policy by effectively solving the above NLP optimization problem as a series of quasi-convex programming problems is presented. Experimental results show significant gain in migration efficiency up to 35%. Yanzhi Wang 0001, Qing Xie 0001, Massoud Pedram, Younghyun Kim 0001, Naehyuck Chang, Massimo Poncino |
DATE | 6 |
| 2012 | NBTI effects on tree-like clock distribution networksabstractNegative Bias Temperature Instability (NBTI) is considered one of the most critical device reliability concerns in nanometer CMOS technologies, because it causes devices to exhibit a temporal drift of performance over time. Wei Liu 0016, Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 6 |
| 2012 | Energy-optimal caches with guaranteed lifetimeabstractThis work addresses the aging of the memory sub-system due to NBTI (Negative Bias Temperature Instability) in systems that have to provide a guaranteed level of service, and specifically, a guaranteed lifetime. Mirko Loghi, Haroon Mahmood, Andrea Calimera, Massimo Poncino, Enrico Macii |
ISLPED | 4 |
| 2012 | Aging-aware caches with graceful degradation of performanceabstractAging of transistors can substantially shorten the lifetime of devices in sub-nanometric technologies. Without any countermeasure, the first component which becomes unreliable will determine the life span of an entire device. This problem is even more relevant for memory arrays, where failure of a single SRAM cell would cause the failure of the whole system. Traditional implementation of power management by turning idle cache lines into a low-energy state can also mitigate the aging effects caused by Negative Bias Temperature Instability (NBTI) provided that idleness is correctly exploited. In this work, we propose a cache structure which deals with cell failures by gracefully degrading its performance. By this partitioning-based strategy, various sub-blocks will become unreliable at different times, and the cache will keep functioning with reduced efficiency. Coupling such aging mitigation with the resulting energy reduction techniques we can obtain up to 2.5x lifetime extension and 40% energy savings with respect to a power managed cache. Haroon Mahmood, Massimo Poncino, Mirko Loghi, Enrico Macii |
VLSI-SoC | 2 |
| 2011 | System level techniques to improve reliability in high power microcontrollers for automotive applicationsabstractIn high power microcontrollers, a decrease in circuit lifetime is often observed in safetycritical applications where circuitry is subjected to the most severe stresses and reliability has become a major concern. Thus, ad-hoc design solutions become necessary to mitigate the impact of ageing. In this paper we discuss hardware-software approaches that exploit distributed on-chip monitoring of wear-out parameters to perform ageing-aware allocation of computation and recovery periods on the various computational units. Andrea Acquaviva, Massimo Poncino, Marco Otella, Michele Sciolla |
DATE | 2 |
| 2011 | Partitioned cache architectures for reduced NBTI-induced agingabstractConventional power management knobs such as voltage scaling or power gating have been shown to have a beneficial effect on the aging phenomena caused Negative Bias Temperature Instability (NBTI). Such a benefit can be especially exploited in SRAM memories, which are particularly sensitive to NBTI effects: given their symmetric structure, they cannot in fact take advantage of value-dependent recovery. We propose an architectural solutions that is based on the idea of partitioning a memory into multiple banks of identical size. While this organization has been widely used for reducing both dynamic and static power, its exploitation for aging benefits requires proper management of the existing idleness of the various banks. This can be achieved by means of a sort of time-varying addressing scheme in which addresses are mapped to different banks over time in such a way that the idleness is uniformly distributed over all the banks. Experimental analysis shows that it is possible to simultaneously reducing leakage power and aging in caches, with minimal overhead and without modifying the internal structure of the SRAM arrays. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2011 | Moving to Green ICT: From stand-alone power-aware IC design to an integrated approach to energy efficient design for heterogeneous electronic systemsabstractEnergy efficiency is one of the most critical aspects of todays information society. The most obvious benefits of being Green are reduced environmental impact and cost savings. Reducing energy consumption of electronic devices, circuits and heterogeneous systems, however, is not trivial. This requires the development of innovative energy-aware vertical design solutions and EDA technologies for next generations' nanoelectronics circuits and systems, and the related energy generation, conversion and management systems. Salvatore Rinaudo, Giuliana Gangemi, Andrea Calimera, Alberto Macii, Massimo Poncino |
DATE | 5 |
| 2011 | Buffering of frequent accesses for reduced cache agingabstractPrevious works have shown that typical power management knobs such as voltage scaling or power gating can also be exploited to reduce aging phenomena caused by Negative Bias Temperature Instability (NBTI). We propose a scheme for power-managed caches that allows to significantly improving the aging of the cache thanks to the use of a small buffer that stores a copy of the lines that are most critical for aging, that is, the ones with the least opportunity of being power-managed; by using the buffer instead of the cache when accessing these critical lines, the original cache is preserved and its lifetime is significantly prolonged. As a side effect, this scheme improves total power since the less energy-hungry buffer is accessed most of the time. Experimental analysis shows this scheme allows to achieve significant (>3x on average) lifetime extensions for the cache, with a concurrent energy saving between 18 and 24%, depending on cache size. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2011 | Balanced reconfiguration of storage banks in a hybrid electrical energy storage systemabstractCompared with the conventional homogeneous electrical energy storage (EES) systems, hybrid electrical energy storage (HEES) systems provide high output power and energy density as well as high power conversion efficiency and low self-discharge at a low capital cost. Cycle efficiency of a HEES system (which is defined as the ratio of energy which is delivered by the HEES system to the load device to energy which is supplied by the power source to the HEES system) is one of the most important factors in determining the overall operational cost of the system. Therefore, EES banks within the HEES system should be prudently designed in order to maximize the overall cycle efficiency. However, the cycle efficiency is not only dependent on the EES element type, but also the dynamic conditions such as charge and discharge rates and energy efficiency of peripheral power circuitries. Also, due to the practical limitations of the power conversion circuitry, the specified capacity of the EES bank cannot be fully utilized, which in turn results in over-provisioning and thus additional capital expenditure for a HEES system with a specified level of service. This is the first paper that presents an EES bank reconfiguration architecture aiming at cycle efficiency and capacity utilization enhancement. We first provide a formal definition of balanced configurations and provide a general reconfigurable architecture for a HEES system, analyze key properties of the balanced reconfiguration, and propose a dynamic reconfiguration algorithm for optimal, online adaptation of the HEES system configuration to the characteristics of the power sources and the load devices as well as internal states of the EES banks. Experimental results demonstrate an overall cycle efficiency improvement of by up to 108% for a DC power demand profile, and pulse duty cycle improvement of by up to 127% for high-current pulsed power profile. We also present analysis results for capacity utilization improvement for a reconfigurable EES bank. Younghyun Kim 0001, Sangyoung Park, Yanzhi Wang 0001, Qing Xie 0001, Naehyuck Chang, Massimo Poncino, Massoud Pedram |
ICCAD | 6 |
| 2011 | Fast Computation of Discharge Current Upper Bounds for Clustered Power GatingabstractThe capability of accurately estimating an upper bound of the maximum current drawn by a digital macroblock from the ground or power supply line constitutes a major asset of automatic power-gating flows. In fact, the maximum current information is essential to properly size the sleep transistor in such a way that speed degradation and signal integrity violations are avoided. Loose upper bounds can be determined with a reasonable computational cost, but they lead to oversized sleep transistors. On the other hand, exact computation of the maximum drawn current is an NP-hard problem, even when conservative simplifying assumptions are made on gate-level current profiles. In this paper, we present a scalable algorithm for tightening upper bound computation, with a controlled and tunable computational cost. The algorithm exploits state-of-the-art commercial timing analysis engines, and it is tightly integrated into an industrial power-gating flow for leakage power reduction. The results we have obtained on large circuits demonstrate the scalability and effectiveness of our estimation approach. Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2011 | Row-Based Power-Gating: A Novel Sleep Transistor Insertion Methodology for Leakage Power Optimization in Nanometer CMOS CircuitsabstractLeakage power has become a serious concern in nanometer CMOS technologies, and power-gating has shown to offer a viable solution to the problem with a small penalty in performance. This paper focuses on leakage power reduction through automatic insertion of sleep transistors for power-gating. In particular, we propose a novel, layout-aware methodology that facilitates sleep transistor insertion and virtual-ground routing on row-based layouts. We also introduce a clustering algorithm that is able to handle simultaneously timing and area constraints, and we extend it to the case of multi-Vtsleep transistors to increase leakage savings. The results we have obtained on a set of benchmark circuits show that the leakage savings we can achieve are, by far, superior to those obtained using existing power-gating solutions and with much tighter timing and area constraints. Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | Post-placement temperature reduction techniquesabstractWith technology scaled to deep submicron era, temperature and temperature gradient have emerged as important design criteria. We propose two post-placement techniques to reduce peak temperature by intelligently allocating whitespace in the hotspots. Both methods are fully compliant with commercial technologies, and can be easily integrated with state-of-the-art thermal-aware design flow. Experiments in a set of tests on circuits implemented in STM 65nm technologies show that our methods achieve better peak temperature reduction than directly increasing circuit's area. Wei Liu 0016, Alberto Nannarelli, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 5 |
| 2010 | An integrated thermal estimation framework for industrial embedded platformsabstractNext generation industrial embedded platforms require the development of complex power and thermal management solutions. Indeed, an increasingly fine and intrusive thermal control is required because of temperature impact on leakage and reliability. To be effective, the implementation of these policies involves decisions that must be taken during various phases along the design process, to enable the development of architectural level countermeasures and the required hardware knobs, such as power modes, power supply regulation granularity and the number of on-chip temperature sensors. As a consequence, a framework allowing thermal estimation exploiting design-time information is desirable. Andrea Acquaviva, Andrea Calimera, Alberto Macii, Massimo Poncino, Enrico Macii, Matteo Giaconia, Claudio Parrella |
ACM Great Lakes Symposium on VLSI | 4 |
| 2010 | Aging effects of leakage optimizations for cachesabstractBesides static power consumption, sub-90nm devices have to account for NBTI effects, which are one of the major concerns about system reliability. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2010 | Thermal-aware floorplanning exploration for 3D multi-core architecturesabstractThermal effects are becoming increasingly important in today's sub-micron technologies. Thermal issues affect the performance, the reliability and the cooling costs of integrated systems. High peak temperatures are of major concern in modern 3D designs, where the stacking of multiple layers leads to higher power densities. Therefore, the integration of the thermal-aware design during the initial phases of the design can reduce the cost and the time-to-market of the resulting product. An efficient floorplanning in terms of thermal effects will reduce the appearance of critical hotspots and will spread heat across the chip area. David Cuesta, José Luis Ayala, J. Ignacio Hidalgo, Massimo Poncino, Andrea Acquaviva, Enrico Macii |
ACM Great Lakes Symposium on VLSI | 4 |
| 2010 | Analysis of NBTI-induced SNM degradation in power-gated SRAM cellsabstractTemporal reliability degradation mechanism, and NBTI in particular, are especially critical for SRAM cells. In fact, unlike logic gates, which under some conditions can be forced into an NBTI-immune state, SRAM cells are always subject to aging, whatever value they are storing. In this work, we first quantify the aging, in terms of degradation of the signal-to-noise margin (SNM), of an SRAM cell as a function of the value stored in the cell, on a 45nm industrial technology. Then, we show how it is possible, by applying power gating to the memory cell, to further reduce the SNM degradation. Finally, we study the joint effect of power gating and bit control techniques. Andrea Calimera, Enrico Macii, Massimo Poncino |
ISCAS | 3 |
| 2010 | Dynamic indexing: concurrent leakage and aging optimization for cachesabstractPrevious works have shown that the traditional implementations of power management (i.e., using power gating or voltage scaling) can also mitigate the aging effect induced by Negative Bias Temperature Instability (NBTI), due to the partial recovery that occurs during the idle intervals used by power management. However, such a potential has been exploited only partially because of the different nature of energy and aging: as a performance figure, aging is affected by the worst idleness pattern. Therefore, large potential energy savings usually turn into limited aging reductions. We address this problem in the context of caches, for which idleness is related to their access pattern. We propose a dynamic indexing scheme, in which the cache indexing function is changed over time in order to uniformly distribute the idleness over all the cache lines. In this way it is possible to fully use the leakage optimization potential and to extend the lifetime of a cache. Experimental analysis shows that it is possible to obtain caches that are effectively aging-free, without any penalty in leakage energy reduction. Andrea Calimera, Mirko Loghi, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2010 | Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven SubbankingabstractPartitioning a memory into multiple blocks that can be independently accessed is a widely used technique to reduce its dynamic power. For embedded systems, its benefits can be even pushed further by properly matching the partition to the memory access patterns. When leakage energy comes into play, however, idle memory blocks must be put into a proper low-leakage sleep state to actually save energy when not accessed. In this case, the matching becomes an instance of the power management problem, because moving to and from this sleep state requires additional energy. In this work, we propose an effective solution to the problem of the leakage-aware partitioning of a memory into disjoint subblocks; in particular, we target scratchpad memories, which are commonly used in some embedded systems as a replacement for caches. We show that, although the solution space is extremely large (for a N--block partition, all the combinations of N-1 address boundaries) and nonconvex, it is possible to prove a nontrivial property that considerably reduces the number of partition boundaries to be enumerated, therefore, making exhaustive exploration feasible. We are thus able to provide an optimal solution to the leakage-aware partitioning problem. Experiments on a different sets of embedded applications have shown that total energy savings larger than 60 percent on average can be obtained, with a marginal overhead in execution time, thanks to an effective implementation of the low-leakage sleep state. Mirko Loghi, Olga Golubeva, Enrico Macii, Massimo Poncino |
IEEE Trans. Computers | 4 |
| 2010 | NBTI-Aware Clustered Power GatingabstractThe emergence of Negative Bias Temperature Instability (NBTI) as the most relevant source of reliability in sub-90nm technologies has led to a new facet of the traditional trade-off between power and reliability. NBTI effects in fact manifest themselves as an increase of the propagation delay of the devices over time, which adds up to the delay penalty incurred by most low-power design solutions. This implies that, given a desired lifetime of a circuit (i.e., a given performance target at some point in time), a power-managed component will fail earlier than a nonpower-managed one. In this work, we show how it is possible to partially overcome this conflict, by leveraging the benefits in terms of aging provided by power-gating (i.e., by using switches that disconnect a logic block from the ground). Thanks to some electrical properties, it is possible to nullify aging effects during standby periods. Based on this important property, we propose a methodology for a NBTI-aware power gating that allows synthesizing low-leakage circuits with maximum lifetime. Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2010 | Temperature-Insensitive Dual- Vth Synthesis for Nanometer CMOS Technologies Under Inverse Temperature DependenceabstractWith the scaling of CMOS technologies, the gap between nominal supply voltage and threshold voltage has decreased significantly. This trend is further amplified in low-power nanometer libraries, which feature cells with identical size and functionality, but different threshold voltages. As a consequence, different cells may have different delay behaviors as the temperature varies within a circuit. For instance, cells with low-threshold devices may experience an increase in delay when temperature increases, whereas cells using high-threshold devices may experience the opposite behavior. The latter effect, also known as inverse temperature dependence (ITD), poses new challenges to circuit designers. Besides making timing analysis more difficult, ITD has important and unforeseeable consequences for power-aware logic synthesis. This paper describes the impact that ITD may have on the design of nanometer circuits. We also provide a threshold voltage assignment algorithm for dual threshold voltage synthesis, which guarantees temperature-insensitive operation of the circuits, together with a significant reduction of both leakage and total power consumption. Experiments performed on a set of standard benchmarks show timing compliance at any operating temperature, and an average leakage reduction around 28% compared to circuits synthesized with a standard synthesis flow that does not take ITD into account. We also apply our proposed synthesis algorithm to a realistic case study consisting of a 32-bit, IEEE-754 floating point unit. Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | Enabling concurrent clock and power gating in an industrial design flowabstractClock-gating and power-gating have proven to be very effective solutions for reducing dynamic and static power, respectively. The two techniques may be coupled in such a way that the clock-gating information can be used to drive the control signal of the power-gating circuitry, thus providing additional leakage minimization conditions w.r.t. those manually inserted by the designer. This conceptual integration, however, poses several challenges when moved to industrial design flows. Although both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. This paper presents a layout-oriented synthesis flow which integrates the two techniques and that relies on leading-edge, commercial EDA tools. Starting from a gated-clock netlist, we partition the circuit in a number of clusters that are implicitly determined by the groups of cells that are clock-gated by the same register. Using a row-based granularity, we achieve runtime leakage reduction by inserting dedicated sleep transistors for each cluster. The entire flow has been benchmarked on a industrial design mapped onto a commercial, 65 nm CMOS technology library. Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 5 |
| 2009 | NBTI-aware sleep transistor design for reliable power-gatingabstractNegative Bias Temperature Instability (NBTI) has been regarded as most important source of reliability of CMOS devices, and specifically pMOS transistors. Andrea Calimera, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Using soft-edge flip-flops to compensate NBTI-induced delay degradationabstractWe present a low-overhead solution to tackle the delay increase caused by Negative Bias Temperature Instability (NBTI), which has emerged as the most critical reliability issue in sub-90nm technology nodes. Karthik Duraisami, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Energy-optimal synchronization primitives for single-chip multi-processorsabstractSynchronization among tasks accounts for a sizable fraction of the energy consumption and execution time of applications running on Multi-Processor Systems-on-Chips platforms. In order to achieve fast and energy-efficient operations, it is therefore essential to implement efficient and power-frugal synchronization primitives. The design of such primitives is complicated by several software and hardware issues, such as: processors running at different speeds, different implementations of the waiting phase upon entering the critical section, and the ratio between static and dynamic power. In this work, we compare a set of classical implementations (i.e., based on busy waiting, or on sleep states) of mutex semaphores, and propose a hybrid (wait/sleep) semaphore in which the sleep state is entered only after a number of busywait cycles. The proposed scheme provides the best overall energy-delay product with respect to previously proposed schemes. Furthermore, we identify an optimal length of the busy-wait cycles, which is empirically shown to depend on the time required to switch from the sleep to the active state. Cesare Ferri, R. Iris Bahar, Mirko Loghi, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2009 | Placement-aware Clustering for Integrated Clock and Power GatingabstractClock-gating and power-gating are the most widely used solutions for reducing dynamic and static power. They can be potentially integrated so that clock-gating conditions can be used to control the power-gating circuitry thus also reducing static power. This integration becomes however difficult when applied in an industrial design flow. Even if both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 5 |
| 2009 | NBTI-aware power gating for concurrent leakage and aging optimizationabstractPower and reliability are known to be intrinsically conflicting metrics: traditional solutions to improve reliability such as redundancy, increase of voltage levels, and up-sizing of critical devices do contrast with traditional low-power solutions, which rely on small devices and scaled supply voltages. The emergence of Negative Bias Temperature Instability (NBTI) as the most relevant source of unreliability in sub-90nm technologies has even exacerbated this incompatibility of the two metrics: NBTI manifests itself as an increase of the propagation delay over time, which adds up to the delay penalty introduced by most low-power design solutions. In this work, we show how the most widely adopted leakage reduction solution, that is, power-gating, can overcome this conflict, and how it can be used to naturally reduce the effects of NBTI on delay. Based on this important property, we present a methodology for NBTI-aware power gating that allows synthesizing low-leakage circuits with maximum lifetime. Andrea Calimera, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2009 | A cosimulation methodology for HW/SW validation and performance estimationabstractCosimulation strategies allow us to simulate and verify HW/SW embedded systems before the real platform is available. In this field, there is a large variety of approaches that rely on different communication mechanisms to implement an efficient interface between the SW and the HW simulators. However, the literature lacks a comprehensive methodology which addresses the need for integrating and synchronizing heterogeneous simulators, like, for example, the SystemC simulation kernel for HW modules and an instruction set simulator for SW applications, without being intrusive for the HW and SW descriptions involved in the simulation. In this context, this article presents, compares, and integrates in a system-level framework two different co-simulation strategies for modeling, analyzing, and validating the performance of a HW/SW embedded system. Moreover, for both of them, a mechanism is proposed to provide an accurate time synchronization of the HW/SW communication. The first strategy is intended to provide an early cosimulation environment where HW/SW interaction can be validated without involving the operating system. The communication is implemented between a single SW task and a SystemC description of an HW module by exploiting the features of the remote debugging interface of a debugger (the GNU GDB), and by modifying the SystemC simulation kernel. On the other hand, the second strategy is intended to be used in further development steps, when the operating system is introduced to validate the cosimulation between HW modules and multitasking SW applications. In this approach, the communication is implemented via interrupts by using the features offered by the operating system. Experimental results are reported on two different case studies to analyze and compare the effectiveness of both the approaches. Franco Fummi, Mirko Loghi, Massimo Poncino, Graziano Pravadelli |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2009 | Tag Overflow Buffering: Reducing Total Memory Energy by Reduced-Tag MatchingabstractWe propose a novel energy-efficient cache architecture based on a matching mechanism that uses a reduced number of tag bits. The idea behind the proposed architecture is based on moving a large subset of the tag bits from the cache into an external register (called theTagOverflowBuffer) that serves as an identifier of the current locality of the memory references. Dynamic energy efficiency is achieved by accessing, for most of the memory references, a reduced-tag cache; furthermore, because of the reduced number of tag bits, leakage energy is also reduced as a by-product. We achieve average energy savings ranging from 16% to 40% (depending on different cache structural parameters) on total (i.e., static and dynamic) cache energy, and measured on a standard suite of embedded applications. Mirko Loghi, Paolo Azzoni, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | A Scalable Algorithmic Framework for Row-Based Power-GatingabstractLeakage power is a serious concern in nanometer CMOS technologies. In this paper we focus on leakage reduction through automatic insertion of sleep transistors for power gating in standard cell based designs. In particular, we propose clustering algorithms for row- based power-gating methodology which is based on using rows of the layout as the granularity for clustering. Our clustering methodology does timing and area constraint driven power-gating in contrast to only timing driven power-gating as proposed in the previous works. We present two distinct clustering algorithms with different accuracy-efficiency trade-off. An optimal one, which exploits a 0-1 or binary integer programming approach, and a heuristic one, which resorts to an implicit enumeration of the layout rows. Results show that, for all the benchmarks, the leakage power savings, as compared to previous techniques, are more than 75% when we have the same timing constraints but half sleep transistor area and at least 60% when area constraint is set at one fourth. We also show that we can perform clustering with no speed degradation and achieve maximum leakage power savings up-to 83%. Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 6 |
| 2008 | Integrating Clock Gating and Power Gating for Combined Dynamic and Leakage Power Optimization in Digital CMOS CircuitsabstractClock Gating and Power Gating are two of the most effective techniques that are applied today for reducing dynamic and leakage power, respectively, in digital CMOS circuits. The combined use of the two solutions, however, poses some challenges in terms of practical integration of the required control logic and the power/timing overhead associated to it. This paper presents an analysis methodology and a prototype CAD tool that support the designer in understanding when the joint application of Clock Gating and Power Gating may result in significant power savings. Enrico Macii, Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Massimo Poncino |
DSD | 5 |
| 2008 | Temperature-insensitive synthesis using multi-vt librariesabstractTemperature fluctuations can alter the delay in MOS circuits. However, increases in temperature do not always lead to a corresponding increase in circuit delay, specifically when operating at low supply voltages. Instead a temperature inversion effect can be observed on the delay of MOS devices under certain conditions, where the delay actually decreases as temperature increases. Given these non-monotonic effects, guaranteeing timing correctness can no longer be achieved simply by characterizing the design under worst case (i.e., high temperature) conditions. In this paper, we present a synthesis methodology in which multi-Vth design is used to generate temperature-insensitive circuits, while minimizing leakage power dissipation as a side-effect. Our experiments with ISCAS benchmark circuits demonstrate the promise of this approach and show that significant reduction in static power is also possible. Andrea Calimera, Enrico Macii, Massimo Poncino, R. Iris Bahar |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Energy efficiency bounds of pulse-encoded busesabstractPulse-encoded buses, (i.e., in which a transition is encoded as a pulse) have recently emerged as an effective solution to solve crosstalk issues in global interconnects, since they suppress transitions in opposite directions by construction. As a side effect, this also reduces energy, since coupling capacitances in deep-submicron technologies are larger than ground capacitances. Furthermore, a single pulse consumes less energy than a conventional transition, because the limited length does not fully implies extra transitions since all input transitions are encoded with a pulse. Karthik Duraisami, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Optimal sleep transistor synthesis under timing and area constraintsabstractLeakage power reduction in nano-CMOS designs has gained tremendous interest both in academia and industry. Many techniques have been proposed in the literature for leakage power reduction and one of the prominent techniques for leakage power reduction is the use of sleep transistors as power-gating elements to cut-off sub-threshold leakage current in circuits when they are in stand-by mode. Although sleep transistor insertion is very effective in cutting-off leakage, it also incurs timing, area and routing overhead. Since most of the sleep transistor insertion methodologies do post layout insertion, care should be taken such that there is minimal perturbation of the original layout. Over design of sleep transistors cells and sub-optimal sleep transistor placement must be avoided to achieve final design closure. Since the sleep transistor area plays an important and prominent role in this aspect, it necessitates for optimal sleep transistor sizing and synthesis technique under area constraints. In this paper, we first provide a methodology for optimal sleep transistor synthesis under given area constraints. We then apply our technique to the general timing and area constraint driven row-based power-gating methodology proposed in [13] and show how optimal low leakage designs with constraints on timing and area can be designed Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 6 |
| 2008 | On quantifying the figures of merit of power-gating for leakage power minimization in nanometer CMOS circuitsabstractPower-gating has proved to be one of the most effective solutions for reducing stand-by leakage power in nanometer-scale CMOS circuits, and different strategies and algorithms for its application have been proposed recently. Unfortunately, power- gating comes with its own set of costs: Performance degradation, area increase, dynamic power increase and routing congestion. When a decision to power-gate a design has to be taken, pros and cons of power-gating have to be properly weighted to achieve optimal results. In this paper, we define "Figures of Merit" (FoMs) for power-gating, which can be used by designers to better understand the benefits and costs of power-gating, thereby allowing them to achieve optimal results. We then quantify the FoMs by applying a state-of-the-art, industry-strength power- gating flow on a set of designs implemented onto an industrial 65 nm CMOS process, and provide insightful discussion on how optimum power-gating can be achieved. Ashoka Visweswara Sathanur, Andrea Calimera, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 7 |
| 2008 | Reducing leakage power by accounting for temperature inversion dependence in dual-Vt synthesized circuitsabstractThe effects of temperature on delay depend on several parameters, such as cell size, load, supply voltage, and threshold voltage. In particular, variations in Vth can yield a temperature inversion effect causing a decreases of cell delay as temperature increases. This phenomenon, besides affecting timing analysis of a design, has important and unforeseeable consequences on power optimization techniques. In this paper, we focus on the impact of such effects on multi-Vt design; in particular, we show how traditional dual-Vt optimization may yield timing errors in circuits by ignoring temperature effects. Moreover, we present a temperature-aware dual-Vt optimization technique that reduces leakage power and can guarantee that the circuit is timing feasible at the boundary temperatures provided by the technology library. Our experiments show an average 27% leakage reduction with respect to a non temperature-aware design flow. Andrea Calimera, R. Iris Bahar, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2008 | Multiple power-gating domain (multi-VGND) architecture for improved leakage power reductionabstractRow-based power-gating has recently emerged as a meet-in-the-middle sleep transistor insertion paradigm between cell-level and block-level granularity, in which each layout row defines the unit of gating, and different rows can be clustered and share the same sleep transistor. Previous works, however, assume the availability of a single virtual ground voltage, thus making the decision of whether to gate or not a given cluster a binary choice: a cluster is either gated or not. Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 5 |
| 2008 | Implementation of a thermal management unit for canceling temperature-dependent clock skew variations
Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Alberto Macii, Enrico Macii, Massimo Poncino |
Integr. | 7 |
| 2008 | Dynamic Thermal Clock Skew Compensation Using Tunable Delay BuffersabstractThe thermal gradients existing in high-performance circuits may significantly affect their timing behavior, in particular, by increasing the skew of the clock net and/or altering hold/setup constraints, possibly causing the circuit to operate incorrectly. The knowledge of the spatial distribution of temperature can be used to properly design a clock network that is able to compensate such thermal non-uniformities. However, redesign of the clock network is effective only if temperature distribution is stationary, i.e., does not change over time. In this paper, we specifically address the problem of dynamically modifying the clock tree in such a way that it can compensate for temporal variations of temperature. This is achieved by exploiting the buffers that are inserted during the clock network generation, by transforming them into tunable delay elements. Temperature-induced delay variations are then compensated by applying the proper tuning to the tunable buffers, which is computed offline and stored in a tuning table inserted in the design. We propose an algorithm to minimize the number of inserted tunable buffers, as well as their tunable range (which directly relates to complexity). Results show that clock skew is kept within original bounds with worst-case power and area penalty of 3.5% and 5.5% respectively. Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2007 | Architectural leakage-aware management of partitioned scratchpad memories
Olga Golubeva, Mirko Loghi, Massimo Poncino, Enrico Macii |
DATE | 3 |
| 2007 | Interactive presentation: Efficient computation of discharge current upper bounds for clustered sleep transistor sizing
Ashoka Visweswara Sathanur, Andrea Calimera, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 6 |
| 2007 | Design of a family of sleep transistor cells for a clustered power-gating flow in 65nm technologyabstractClustered sleep transistor insertion is an effective leakage power reduction technique that is well-suited for integration in an automated design flow and offers a flexible tradeoff between area, delay overhead and turn-on transition time. In this work, we focus on the design of a family of sleep transistor cells, fully compatible with the physical design rules of a commercial 65nm CMOS library. We describe circuit-level and layout optimizations, as well as the cell characterization procedure required to support automated sleep transistor cell selection and instantiation in a clustered power-gating insertion flow. Andrea Calimera, Antonio Pullini, Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 7 |
| 2007 | On the energy efficiency of synchronization primitives for shared-memory single-chip multiprocessorsabstractApplications running on Multiprocessor Systems-on-Chips (MP-SoCs) exhibit complex interaction patterns, resulting in significant amounts of time spent while synchronizing for mutually exclusive access to shared resources. Such an overhead is expected to increase with the degree of parallelism and with the mutual correlation of concurrent tasks, thus becoming in a severe obstacle to the full exploitation of a system potential. Although the topic has been extensively studied in the literature, in MPSoC architectures, which exhibit different tradeoffs with respect to traditional multi-processors, the available results may not be valid or hold only partially. Furthermore, the strict energy budget of MPSoCs requires also the evaluation of the energy efficiency of such synchronization primitives. In this work we survey various state-of-the-art implementations of synchronization primitives, in order to assess their impact on performance and on energy consumption. The results of our analysis show that some commonly accepted intuitions in the multiprocessor domain do not hold in the context of MPSoCs. Olga Golubeva, Mirko Loghi, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Design Exploration of a Thermal Management Unit for Dynamic Control of Temperature-Induced Clock SkewabstractPower densities and temperatures in today's high performance circuits have reached alarmingly high levels due to increased scaling in feature sizes. Subsequently, the various techniques used to keep them under control have also created "zones" of varying temperatures, thus contributing to temperature gradients inside the chip. These gradients have detrimental effects on the delay of wires, as resistance in metals increases with temperature. Clock nets are extremely susceptible to this effect, since they run through the entire chip. Different techniques have been proposed to counter the impact of temperature on clock speed; they range from re-designing the clock network assuming a stationary profile to more adaptive solutions that allow to dynamically compensate the clock skew through replacement of the original buffers with a specially designed counterpart, called tunable delay buffers (TDBs). Dynamic skew management based on TDBs calls for the presence on the chip of a thermal management unit (TMU), whose purpose is that of periodically choosing the actual delay that each TDB must provide in order to achieve skew optimization. Preliminary implementations of such a unit for basic assumptions on the distribution of sensors and their accuracy have indicated negligible impact on the original design. This work aims at exploring in detail several issues related to TMU design, pivoting on the fact that sensor distribution and its accuracy could in fact impact the design in a significant way depending on the design. We provide the results of a careful exploration we have performed on a meaningful case study, quantifying values for area and power consumption. Karthik Duraisami, Prassanna Sithambaram, Ashoka Visweswara Sathanur, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 6 |
| 2007 | Locality-driven architectural cache sub-banking for leakage energy reductionabstractIn most processors, caches account for the largest fraction of onchip transistors, thus being a primary candidate for tackling the leakage problem. Existing architectural solutions usually rely on customized cache structures, which are needed to implement some kind of power management policy. Memory arrays, however, are carefully developed and finely tuned by foundries, and their internal structure is typically non accessible to system designers. Olga Golubeva, Mirko Loghi, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2007 | Timing-driven row-based power gatingabstractIn this paper we focus on leakage reduction through automatic insertion of sleep transistors using a row-based granularity. In particular, we tackle here the two main issues involved in this methodology: (i) Clustering and (ii) the interfacing of power-gated and non power-gated regions within the same block. The clustering algorithm automatically selects an optimal subset of rows that can be power-gated with a tightly controlled delay overhead. We then address the issue of interfacing different gated regions and propose a novel technique to address this issue with minimal area and power penalty. Our approach is compatible with state-of-the art logic and physical synthesis flows and it does not significantly impact design closure. We achieve leakage power reductions as high as 89% for a set of standard benchmarks, with minimum timing and area overhead. Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 6 |
| 2007 | Energy-Efficient Multiprocessor Systems-on-Chip for Embedded Computing: Exploring Programming Models and Their Architectural SupportabstractIn today's multiprocessor SoCs (MPSoCs), parallel programming models are needed to fully exploit hardware capabilities and to achieve the 100 Gops/W energy efficiency target required for ambient intelligence applications. However, mapping abstract programming models onto tightly power-constrained hardware architectures imposes overheads which might seriously compromise performance and energy efficiency. The objective of this work is to perform a comparative analysis of message passing versus shared memory as programming models for single-chip multiprocessor platforms. Our analysis is carried out from a hardware-software viewpoint: we carefully tune hardware architectures and software libraries for each programming model. We analyze representative application kernels from the multimedia domain, and identify application-level parameters that heavily influence performance and energy efficiency. Then, we formulate guidelines for the selection of the most appropriate programming model and its architectural support Francesco Poletti, Antonio Poggiali, Davide Bertozzi, Luca Benini, Paul Marchal, Mirko Loghi, Massimo Poncino |
IEEE Trans. Computers | 7 |
| 2007 | Power macromodeling of MPSoC message passing primitivesabstractEstimating the energy consumption of software in multiprocessor systems-on-chip (MPSoCs) is crucial for enabling quick evaluations of both software and hardware optimizations. However, high-level estimations should be applicable at software level, possibly constructing effective power models depending on parameters that can be extracted directly from the application characteristics. We propose a methodology for accurate analysis of power consumption of message-passing primitives in a MPSoC, and, in particular, an energy model which, in spite of its simplicity, allows to model the traffic-dependent nature of energy consumption through the use of a single, abstract parameter, namely, the size of the message exchanged. Mirko Loghi, Luca Benini, Massimo Poncino |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2006 | Thermal resilient bounded-skew clock tree optimization methodologyabstractThe existence of non-uniform thermal gradients on the substrate in high performance IC's can significantly impact the performance of global on-chip interconnects. This issue is further exacerbated by the aggressive scaling and other factors such as dynamic power management schemes and non-uniform gate level switching activity. In high-performance systems, one of the most important problems is clock skew minimization since it has a direct impact on the maximum operating frequency of the system. Since clocks are routed across the entire chip, the presence of thermal gradients can significantly alter their characteristics because wire resistance increases linearly as the temperature increases. This often results in failure to meet original timing constraints thereby rendering the original topology unusable. Therefore it is necessary to perform a temperature aware re-embedding of the original topology to meet timing under these temperature effects. This work primarily explores these issues by proposing two algorithms that re-structure an existing clock tree topology to compensate for such temperature effects and as a result also meet timing constraints Ashutosh Chakraborty, Prassanna Sithambaram, Karthik Duraisami, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 6 |
| 2006 | ISS-centric modular HW/SW co-simulationabstractModular design is an important requirement in modern embedded system design flows because of the widespread acceptance of new paradigms such as IP core reuse and platform-based design. Co-simulation frameworks must thus support modular design, since programmable devices, ad-hoc HW components, and the interconnect infrastructure must be easily interchangeable in order to allow design exploration while keeping the SW portion unchanged or only marginally changed. The proposed co-simulation framework implements such a modular approach to co-simulation by means of a novel paradigm in which HW models can be modified on the fly by keeping the SW parts unchanged. This is achieved through an ISS-centric co-simulation strategy in which modularity is provided in terms of (i) the replacement of HW components thanks to the use of a common interface based on the device address space, or (ii) the use of different ISS's, thanks to a re-configurable simulator. We demonstrate our approach onto an industrial-strength embedded application, showing that the proposed co-simulation strategy provides both high speed and accuracy. Franco Fummi, Giovanni Perbellini, Mirko Loghi, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2006 | STV-Cache: a leakage energy-efficient architecture for data cachesabstractWe propose a low-leakage cache architecture based on the observation of the spatio-temporal properties of data caches. In particular, we exploit the fact that during the program lifetime a few data values tend to exhibit both spatial and temporal locality in cache, i.e., values that are simultaneously stored by several lines at the same time. Leakage energy can be reduced by turning off those lines and storing these values in a smaller, separate memory. In this work we introduce an architecture that implements such a scheme, as well as an algorithm to detect these special values. We show that by using as few as four values we can achieve 18.45% leakage energy savings, with an additional 13.85% reduction of dynamic energy as a consequence of a reduced average cache access cost. Kimish Patel, Luca Benini, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2006 | Implications of ultra low-voltage devices on design techniques for controlling leakage in NanoCMOS circuitsabstractEnabled by technology scaling, ultra low-voltage devices have now found wide application in modern VLSI circuits. While low-voltage implies reduced dynamic power, it also signifies increased leakage power, as lower supply voltages are usually paired with lower threshold voltages in order to preserve circuit speed. This originates an increase in sub-threshold leakage currents that constitute, today, one of the most serious bottlenecks to further technology and supply voltage scaling. The need of controlling leakage power in nanometric devices is imposing a significant shift in the way integrated circuits are designed and manufactured. The behavior of devices with nanometric feature sizes is much more sensitive to parameters such as the operating temperature of the circuit, which in the past were neglected. In this paper we quantitatively analyze the leakage control capabilities of some well-established circuit-level design techniques, and assess how the effectiveness of such techniques scales with respect to decreased supply voltages (as induced by technology scaling) and temperature variations, thus providing an interesting insight on how leakage control solutions that are in use today is applicable in future designs Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 7 |
| 2006 | Low-energy pixel approximation for DVI-based LCD interfacesabstractSeveral options are available for the approximating of color images in hardware, both in terms of the type of transformation (e.g., quantization, dithering) as well as in terms of where the approximation takes place (e.g., graphics controller, frame buffer, LCD controller). In this work, we propose a color approximation approach, orthogonal to other color simplification schemes, which is done during the digital transmission of the color data to the LCD. The proposed technique targets a specific digital standard (namely, DVI) and its serial protocols TMDS, to provide a 45-66% savings in the number of bit transitions on the LCD bus (which translate to a corresponding saving of energy), depending on the allowed degradation of image quality A. Nurrachmat, Enrico Macii, Massimo Poncino |
ISCAS | 3 |
| 2006 | Dynamic thermal clock skew compensation using tunable delay buffersabstractThe thermal gradients existing in high-performance circuits may significantly affect their timing behavior, in particular by increasing the skew of the clock net and/or altering hold/setup constraints, possibly causing the circuit to operate incorrectly. The knowledge of the spatial distribution of temperature can be used to properly design a clock network that is able to compensate such thermal non-uniformities. However, re-design of the clock network is effective only if temperature distribution is stationary, i.e., does not change over time. In this work, we specifically address the problem of dynamically modifying the clock tree in such a way that it can compensate for temporal variations of temperature. This is achieved by exploiting the buffers that are inserted during the clock network generation, by transforming them into tunable delay elements. Temperature-induced delay variations are then compensated by applying the proper tuning to the tunable buffers, which is computed off-line and stored in a tuning table inserted in the design. We propose an algorithm to minimize the number of inserted tunable buffers, as well as their tunable range (which directly relates to complexity). Results show that clock skew is kept within original bounds with minimum area and power penalty. The maximum increase in power is 23.2% with most benchmarks exhibiting less than 5% increase in power. Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 8 |
| 2006 | Synchronization-driven dynamic speed scaling for MPSoCsabstractEqualizing the ratios between workloads and speeds of processing elements provides the optimal speed allocation. Based on that principle, this work describes a dynamic speed setting policy for multiprocessor systems-on-chip (MPSoCs) that relies on the estimation of processor idle times specifically due to the synchronization work. The policy provides two advantages: first, it does not rely on any assumption about the communication pattern of the application executed by the system. Second, it is purely architectural; it automatically detects changes in the system workload and sets processors speeds accordingly by means of a custom hardware block.Results on a parallel MPEG video decoding application show an EDP saving above 55%, averaged over several datasets, corresponding to an energy saving above 50%, and a corresponding penalty in performance below 8%. Mirko Loghi, Massimo Poncino, Luca Benini |
ISLPED | 2 |
| 2006 | Reducing Conflict Misses by Application-Specific Reconfigurable IndexingabstractThe predictability of memory access patterns in embedded systems can be successfully exploited to devise effective application-specific cache optimizations. In this paper, an improved indexing scheme for direct-mapped caches, which drastically reduces the number of conflict misses by using application-specific information, is proposed. The indexing scheme is based on the selection of a subset of the address bits. With respect to similar approaches, the solution has two main strengths. First, owing to an analytical model for the conflict-miss conditions of a given trace, it provides a symbolic algorithm to compute the optimum solution (i.e., the subset of address bits to be used as cache index that minimize the number of conflict misses). Second, owing to a reconfigurable bit selector that can be programmed at run time, it allows the optimal cache indexing to fit to a given application. Results show an average reduction of conflict misses of 24%, measured over a set of standard benchmarks, and for different cache configurations Kimish Patel, Luca Benini, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | Cache coherence tradeoffs in shared-memory MPSoCsabstractShared memory is a common interprocessor communication paradigm for single-chip multiprocessor platforms. Snoop-based cache coherence is a very successful technique that provides a clean shared-memory programming abstraction in general-purpose chip multiprocessors, but there is no consensus on its usage in resource-constrained multiprocessor systems on chips (MPSoCs) for embedded applications. This work aims at providing a comparative energy and performance analysis of cache-coherence support schemes in MPSoCs. Thanks to the use of a complete multiprocessor simulation platform, which relies on accurate technology-homogeneous power models, we were able to explore different cache-coherent shared-memory communication schemes for a number of cache configurations and workloads. Mirko Loghi, Massimo Poncino, Luca Benini |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2005 | Virtual Hardware Prototyping through Timed Hardware-Software Co-SimulationabstractDesigners of factory automation applications increasingly demand tools for rapid prototyping of hardware extensions to existing systems and verification of resulting behaviors through hardware and software co-simulation. The paper presents a framework for the timing-accurate co-simulation of HDL models and their verification against hardware and software running on an actual embedded device of which only a minimal knowledge of the current design is required. Experiments on real-life applications show that early architectural and design decisions can be taken by measuring the expected performance on the models realized using the proposed framework. Franco Fummi, Mirko Loghi, Stefano Martini, Marco Monguzzi, Giovanni Perbellini, Massimo Poncino |
DATE | 6 |
| 2005 | Tag Overflow Buffering: An Energy-Efficient Cache ArchitectureabstractWe propose a novel energy-efficient memory architecture which relies on the use of a cache with a reduced number of tag bits. The idea behind the proposed architecture is based on moving a large number of the tag bits from the cache into an external register (tag overflow buffer) that identifies the current locality of the memory references; additional hardware allows us to dynamically update the value of the reference locality contained in the buffer. Energy efficiency is achieved by using, for most of the memory accesses, a reduced-tag cache. This architecture is minimally intrusive for existing designs, since it assumes the use of a regular cache, and does not require any special circuitry internal to the cache such as row or column activation mechanisms. Average energy savings are 51% on tag energy, corresponding to about 20% saving on total cache energy, measured on a set of typical embedded applications. Mirko Loghi, Paolo Azzoni, Massimo Poncino |
DATE | 3 |
| 2005 | Exploring Energy/Performance Tradeoffs in Shared Memory MPSoCs: Snoop-Based Cache Coherence vs. Software SolutionsabstractShared memory is a common interprocessor communication paradigm for single-chip multi-processor platforms. Snoop-based cache coherence is a very successful technique that provides a clean shared-memory programming abstraction in general-purpose chip multiprocessors, but there is no consensus on its usage in resource-constrained multiprocessor systems on chips (MPSoC) for embedded applications. This work aims at providing a comparative energy and performance analysis of cache coherence support schemes in MPSoC. Thanks to the use of a complete multiprocessor simulation platform, which relies on accurate technology-homogeneous power models, we were able to explore different cache-coherent shared-memory communication schemes for a number of cache configurations and workloads. Mirko Loghi, Massimo Poncino |
DATE | 2 |
| 2005 | Exploring the energy efficiency of cache coherence protocols in single-chip multi-processorsabstractThe performance of the various cache coherence protocols proposed in the literature have been extensively analyzed in the context of high-performance multi-processor systems.A similar analysis for Multi-Processor Systems-on-Chips (MP-SoCs), where energy is at least as important as performace, and for which strict constraints on hardware and software resources do exist, has not been done yet.This work provides an effort in that sense, showing energy/performance tradeoffs for different snoop-based protocols on a realistic MPSoC architecture. The analysis leverage a multi-processor simulation platform, augmented with accurate power models, that allows cycle-accurate simulations.Our analysis show that (i) cache write policy is actually more important than the actual cache coherence protocol, and (ii) matching the programming model and style to the architecture may have dramatic effects on the energy and performance of the system. Mirko Loghi, Martin Letis, Luca Benini, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2005 | Zero clustering: an approach to extend zero compression to instruction cachesabstractWe propose an energy-efficient architecture for instruction caches that relies on dynamic zero compression (DZC), that is, the possibility of reading and writing a single bit for every zero-valued byte [5]. We enhance the basic DZC by using a simple bit permutation to increase the number of zero-valued bytes, so that the corresponding overhead is negligible. The derivation of an effective permutation relies on a heuristic zero clustering algorithm that is based on the knowledge of the memory reference access trace, thus making this solution suitable for application-specific embedded systems. The architecture proposed in this work makes possible the application of zero compression to instruction caches; experiments showed an increase of zero clusters of more than 70% on average, which translates into a 10% improvement in dynamic energy savings with respect to DZC. Kimish Patel, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Energy-Efficient Color Approximation for Digital LCD InterfacesabstractThe limited resolution capabilities of color displays, coupled with the limited perceptual resolution of the human eye have been exploited to reduce the actual number of colors that are simultaneously displayed. In this work, we propose a color approximation approach, orthogonal to traditional color simplification schemes done either in the frame buffer or in the LCD controller; that targets the reduction of the energy required for the transmission of color data on digital LCD interfaces. Our approach leverages the serial nature of the transmission so as to approximate RGB pixel values with suitable codes that minimize the number of transitions on the LCD bus, for a given tolerated image quality level. Application of this scheme on a set of images shows energy savings from 60% to 75%, depending on the image quality. Andi Nourrachmat, Sabino Salerno, Enrico Macii, Massimo Poncino |
ICCD | 4 |
| 2005 | Frame Buffer Energy Optimization by Pixel PredictionabstractWe propose a technique to reduce the energy consumption of the frame buffer memory, based on the spatial locality of images and display frames. Our scheme reduces energy by selectively avoiding reads from the frame buffer when identical adjacent pixels are detected. This is made possible by using an auxiliary memory that stores the locality information. The proposed architecture allows to dynamically update the locality information, and, unlike previous approaches, it works virtually independent of the size and position of the updates of the display frames. Experimental results evaluated on a set of typical graphical applications show a reduction of about 40% of frame buffer reads. Kimish Patel, Enrico Macii, Massimo Poncino |
ICCD | 3 |
| 2004 | Modeling and Analysis of Heterogeneous Industrial Networks ArchitecturesabstractThe paper deals with modelling and analysis of heterogeneous industrial networks architectures. At first homogeneous modelling strategy using the network simulator is presented and then the role of systemC and the use of the instruction set simulator is provided. Experimental results shows that the network cannot support the workload generated by the PLCs. Franco Fummi, Stefano Martini, Marco Monguzzi, Giovanni Perbellini, Massimo Poncino |
DATE | 5 |
| 2004 | Native ISS-SystemC Integration for the Co-Simulation of Multi-Processor SoCabstractIn a system-level design flow, the transition from a high-level description entry implies the refinement from an untimed, unpartitioned description to a real architecture where applications are executed on a programmable device and interact with ad-hoc hardware components. Simulation of such architectures requires the capability of efficient co-simulation of a model of hardware with a model of the processor. This paper presents two co-simulation methodologies, based on SystemC as hardware modeling language and on an instruction set simulator (ISS) as a model of the processor. The first one works at the SystemC kernel level and exploits potentialities of the GNU suite, whereas the second uses features offered by the operating system running on the ISS. The two methodologies improve co-simulation performance with respect to state-of the art methods, and provide different trade-offs between the simplicity of the programming model, the modeling power, and co-simulation performance. Franco Fummi, Stefano Martini, Giovanni Perbellini, Massimo Poncino |
DATE | 4 |
| 2004 | Heterogeneous Co-Simulation of Networked Embedded SystemsabstractNetworked embedded systems pose several challenges in the modeling, simulation, and design domains. The presence of the network, in particular, makes an already critical task such as HW/SW co-simulation even more complex, since a three-way (HW/SW/network) co-simulation and co-design capability is required. Modeling of networks and their interaction with hardware and software is thus key for an effective design methodology at early stages of the design flow. In this work, we present a HW/SW/network co-simulation and co-design methodology, based on the integration of heterogeneous simulation environments such as systemC and NS (network simulator). This methodology has been successfully applied to the design of a system-on-chip performing the fast path of IPv4 routing, allowing to explore different HW/SW allocation for different network configurations. Franco Fummi, Stefano Martini, Giovanni Perbellini, Massimo Poncino, Fabio Ricciato, Maura Turolla |
DATE | 4 |
| 2004 | Synthesis of Partitioned Shared Memory Architectures for Energy-Efficient Multi-Processor SoCabstractAccesses to the shared memory in multi-processor systems-on-chip represent a significant performance bottleneck. Multi-port memories are a common solution to this problem, because they allow parallel accesses. However, they are not an energy-efficient solution. We propose an energy-efficient shared-memory architecture that can be used as a substitute for multi-port memories, which is based on an application-driven partitioning of the shared address space into a multi-bank architecture. Experiments on a set of parallel benchmarks show energy savings of about 56% with respect to a dual-port memory architecture, at a very limited performance penalty. Kimish Patel, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2004 | Energy-efficient bus encoding for LCD displaysabstractThis paper presents a low-power bus encoding technique suitable for the digital interface to a Liquid Crystal Display (LCD). In particular, we focus on interfaces that are compliant to the Digital Visual Interface (DVI) standard, in which the three color channels are serially transmitted to achieve high bandwidth.The proposed technique exploits the well-know inter-pixel correlation that exists in typical images by serially transmitting an encoded representation of the difference between adjacent pixels. The encoding is based on the principle of clustering the 1's in the code towards either ends of the pixel data, in such a way that serial transmission of a code yields at most 1 transition per pixel.The application of the encoding to a series of standard images resulted in energy savings of around 60% on average, with respect to a plain transmission of 8-bit pixel data. Alberto Bocca, Sabino Salerno, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2004 | Cycle-accurate power analysis for multiprocessor systems-on-a-chipabstractDeveloping energy-aware software for multiprocessor systems-on-chip (MPSoCs) is a difficult task, which requires the knowledge of the distribution of the power consumption among several heterogeneous devices (cores, memories, busses, etc.). In this work we analyze the power breakdowns of power consumption for a complete MPSoC platform, under several application workloads and operating conditions. We leverage a complete-system simulation platform with accurate power models for all key hardware modules. Our analysis shows that caches and system interconnect dominate in the power breakdown, pointing out how software locality is meaningful not only for performance but also for energy optimization. Mirko Loghi, Massimo Poncino, Luca Benini |
ACM Great Lakes Symposium on VLSI | 2 |
| 2004 | Reducing cache misses by application-specific re-configurable indexingabstractThe predictability of memory access patterns in embedded systems can be successfully exploited to devise effective application-specific cache optimizations. In this work, we propose an improved indexing scheme for direct-mapped caches, which drastically reduces the number of conflict misses by using application-specific information; the scheme is based on the selection of a subset of the address bits. With respect to similar approaches, our solution has two main strengths. First, it models the misses analytically by building a miss equation, and exploits a symbolic algorithm to compute the exact optimum solution (i.e., the subset of address bits to be used as cache index that minimizes conflict misses). Second, we designed a re-configurable bit selector, which can be programmed at run-time to fit the optimal cache indexing to a given application. Results show an average reduction of conflict misses of 24%, measured over a set of standard benchmarks, and for different cache configurations. Kimish Patel, Enrico Macii, Luca Benini, Massimo Poncino |
ICCAD | 4 |
| 2004 | DynamoSim: a trace-based dynamically compiled instruction set simulatorabstractInstruction set simulators are indispensable tools for the architectural exploration and verification of embedded systems. Different techniques have recently been proposed to speed up the simulation over the classical interpretation-based simulators, while maintaining their flexibility. We introduce a suite of techniques inspired by recent advances in dynamic compilers to construct a hybrid simulation framework. Compared with compiled simulators reported earlier, our framework is more flexible, since any instruction can be interpreted; and faster, since only frequently executed instructions are translated on-the-fly into native code for direct execution, and the scope of our translation is extended from basic blocks to traces, and sophisticated register allocation is performed. Comprehensive results on SPEC2000 benchmarks are reported for the standard SimpleScalar processor to demonstrate the efficiency of proposed techniques. Massimo Poncino, Jianwen Zhu |
ICCAD | 1 |
| 2004 | Software/Network Co-Simulation of Heterogeneous Industrial Networks ArchitecturesabstractThis work presents a modeling and analysis framework for heterogeneous industrial networks architectures, which is based on a tight integration of a network simulator with embedded software, middleware and a real-time operating system. This framework is suitable for modeling and simulating the behavior of typical components involved in factory automation applications (e.g., PLCs, remote controllers, operational screens, etc.) when they are connected through heterogeneous industrial networks. Experiments show that the framework allows to take early architectural decisions by evaluating the expected system performance based on the available models. Franco Fummi, Stefano Martini, Marco Monguzzi, Giovanni Perbellini, Massimo Poncino |
ICCD | 5 |
| 2004 | Analyzing Power Consumption of Message Passing Primitives in a Single-Chip MultiprocessorabstractIn this work, we propose a methodology for the accurate analysis of the power consumption of interprocessor communication in an MPSoC, and the construction of high-level power macromodels. The models leverage a complete MPSoC power estimation environment, that allows to evaluate the power consumption of various software functions, including message passing primitives, which could not be fully characterized in single-processor analysis framework developed in the past. Based on this data we built power macromodels that achieve average estimation errors below 5%. Mirko Loghi, Luca Benini, Massimo Poncino |
ICCD | 3 |
| 2004 | Limited intra-word transition codes: an energy-efficient bus encoding for LCD display interfacesabstractWe propose a class of low-power codes, called Limited Intra-Word Transition (LIWT) codes, suitable for the digital interface to Liquid Crystal Displays (LCD).The proposed technique exploits the existing inter-pixel correlation of typical images, by transmitting an encoded representation of the difference between adjacent pixels. Since all standard LCD transmission protocols are serial, the LIWT specifically targets the minimization of intra-word transitions.The application of the encoding to a series of standard images resulted in transitions savings over 60% on average with respect to two standard TMDS and LVDS protocols. Sabino Salerno, Alberto Bocca, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2003 | Energy-aware design techniques for differential power analysis protectionabstractDifferential power analysis is a very effective cryptanalysis technique that extracts information on secret keys by monitoring instantaneous power consumption of cryptoprocessors. To protect against differential power analysis, power supply noise is added in cryptographic computations, at the price of an increase in power consumption. We present a novel technique, based on well-known power-reducing transformations coupled with randomized clock gating, that introduces a significant amount of scrambling in the power profile without increasing (and, in some cases, by even reducing) circuit power consumption. Luca Benini, Alberto Macii, Enrico Macii, Elvira Omerbegovic, Fabrizio Pro, Massimo Poncino |
DAC | 6 |
| 2003 | A timing-accurate modeling and simulation environment for networked embedded systemsabstractThe design of state-of-the-art, complex embedded systems requires the capability of modeling and simulating the complex networked environment in which such systems operate. This implies the availability of both a networking modeling environment and traditional system-level modeling capabilities. In this paper we present a modeling and simulation methodology based on a timing accurate integration of a system-level modeling language (SystemC) and a network simulation environment (NS-2). The efficiency of the proposed design environment has been demonstrated on a description of a 802.11 MAC layer. Franco Fummi, Giovanni Perbellini, Paolo Gallo, Massimo Poncino, Stefano Martini, Fabio Ricciato |
DAC | 4 |
| 2003 | Estimation of Bus Performance for a Tuplespace in an Embedded ArchitectureabstractThis paper describes a design methodology for the estimation of bus performance of a tuplespace for factory automation. The need of a tuplespace is motivated by the characteristics of typical embedded architectures for factory automation. We describe the features of a bus for embedded applications and the problem of estimating its performance, and present a rapid prototyping design methodology developed for a qualitative and quantitative estimation. The methodology is based on a mix of different modeling languages such as Java, C++, SystemC and Network Simulator2 (NS2). Its application allows the estimation of the expected performance of the bus under design in relation to the developed tuplespace. Nicola Drago, Franco Fummi, Marco Monguzzi, Giovanni Perbellini, Massimo Poncino |
DATE | 5 |
| 2003 | Improving the Efficiency of Memory Partitioning by Address Clustering
Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2003 | A novel architecture for power maskable arithmetic unitsabstractPower maskable units have been proposed as a viable solution for preventing side-channel attacks to cryptoprocessors. This paper presents a novel architecture for the implementation of a class of such kinds of units, namely arithmetic components, which find wide usage in cryptographic applications and which are not suitable to traditional masking techniques. Results of extensive exploration and architectural trade-off analysis show the viability of the proposed solution. Luca Benini, Alberto Macii, Enrico Macii, Elvira Omerbegovic, Massimo Poncino, Fabrizio Pro |
ACM Great Lakes Symposium on VLSI | 5 |
| 2003 | Combining wire swapping and spacing for low-power deep-submicron busesabstractWe propose an approach for reducing the energy consumption of address buses that targets both the switching and the crosstalk components of power dissipation.The method is based on the combined application of two techniques. First, selective wire swapping is applied in such a way that bus wires with high coupling activity are kept far away from each other. Then, the slack available in the floorplanning for the routing of the bus wire is exploited to realize a bus with non-uniform inter-wire spacing. Both swapping and placement are driven by the switching data obtained from the analysis of typical address bus traces, and can be successfully applied to any address bus.Results on a set of profiled address streams show the effectiveness of the proposed approach. Enrico Macii, Massimo Poncino, Sabino Salerno |
ACM Great Lakes Symposium on VLSI | 2 |
| 2003 | Energy-efficient data scrambling on memory-processor interfacesabstractCrypto-processors are prone to security attacks based on the observation of their power consumption profile. We propose new techniques for increasing the non-determinism of such profile, which rely on the idea of introducing randomness in the bus data transfers. This is achieved by combining data scrambling with energy-efficient bus encoding, thus providing high information protection at no energy cost.Results on a set of bus traces originated by real-life applications demonstrate the applicability of the proposed solution. Luca Benini, Angelo Galati, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 5 |
| 2003 | Discharge Current Steering for Battery Lifetime OptimizationabstractPortable and wearable computers can be powered by different combinations of two or more battery packs to give the user the possibility of choosing an optimal compromise between lifetime and weight/size. Recent work on battery-driven power management has demonstrated that sequential discharge is suboptimal in multibattery systems and lifetime can be maximized by distributing (steering) the current load on the available batteries, thereby discharging them in a partially concurrent fashion. Based on these observations, we formulate multibattery lifetime maximization as a continuous, constrained optimization problem, which can be efficiently solved by nonlinear optimizers. We show that significant lifetime extensions can be obtained with respect to standard sequential discharge (up to 160 percent), as well to previously proposed battery scheduling algorithms (up to 12 percent). Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Computers | 5 |
| 2003 | Energy-aware design of embedded memories: A survey of technologies, architectures, and optimization techniquesabstractEmbedded systems are often designed under stringent energy consumption budgets, to limit heat generation and battery size. Since memory systems consume a significant amount of energy to store and to forward data, it is then imperative to balance power consumption and performance in memory system design. Contemporary system design focuses on the trade-off between performance and energy consumption in processing and storage units, as well as in their interconnections. Although memory design is as important as processor design in achieving the desired design objectives, the former topic has received less attention than the latter in the literature. This article centers on one of the most outstanding problems in chip design for embedded applications. It guides the reader through different memory technologies and architectures, and it reviews the most successful strategies for optimizing them in the power/performance plane. Luca Benini, Alberto Macii, Massimo Poncino |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2003 | Scheduling battery usage in mobile systemsabstractThe use of multibattery power supplies is becoming common practice in electronic appliances of the latest generations. Economical and manufacturing constraints are at the basis of this choice. Unfortunately, a partitioned battery subsystem is not able to deliver the same amount of charge as a monolithic battery with the same total capacity. In this paper, we define the concept of battery scheduling, we investigate several policies for solving the problem of optimal charge delivery, and we study the relationship of such policies with different configurations of the battery subsystem. Experimental results, obtained for different kinds of current workloads, demonstrate that the choice of the proper scheduling can make system lifetime as close as 1% of the theoretical upper bound, that is, a monolithic power supply of equal capacity. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | Wire Placement for Crosstalk Energy Minimization in Address BusesabstractWe propose a novel approach to bus energy minimization that targets crosstalk effects. Unlike previous approaches, we try to reduce energy through capacitance optimization, by adopting nonuniform spacing between wires. This allows reduction of power and at the same time takes into account signal integrity. Therefore, performance is not degraded. Results show that the method saves up to 30% of total bus energy at no cost in performance or complexity of the design (no encoding-decoding circuitry is needed), and limited cost in area. Luca Macchiarulo, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2002 | Enhanced clustered voltage scaling for low powerabstractThis paper presents a voltage scaling approach that is based on an enhanced variant of clustered voltage scaling originally proposed by Usami and Horowitz ([1]) The results show that subtituting the original depth first strategy with a breadth first one results in improved speed and quality of results. Data are validated through power and timing analysis performed with a commercial tool. Monica Donno, Luca Macchiarulo, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 5 |
| 2002 | Legacy SystemC Co-Simulation of Multi-Processor Systems-on-ChipabstractWe present a co-simulation environment for multiprocessor architectures, that is based on SystemC and allows a transparent integration of instruction set simulators (ISSs) within the SystemC simulation framework. The integration is based on the well-known concept of bus wrapper, that realizes the interface between the ISS and the simulator. The proposed solution uses an ISS-wrapper interface based on the standard gdb remote debugging interface, and implements two alternative schemes that differ in the amount of communication they require. The two approaches provide different degrees of tradeoff between simulation granularity and speed, and show significant speedup with respect to a micro-architectural, full SystemC simulation of the system description. Luca Benini, Davide Bertozzi, Davide Bruni, Nicola Drago, Franco Fummi, Massimo Poncino |
ICCD | 6 |
| 2002 | Discharge current steering for battery lifetime optimizationabstractRecent work on battery-driven power management has demonstrated that sequential discharge is suboptimal in multi-battery systems, and lifetime can be maximized by distributing (steering) the current load on the available batteries, thereby discharging them in a partially concurrent fashion. Based on these observations, we formulate multi-battery lifetime maximization as a continuous, constrained optimization problem, which can be efficiently solved by non-linear optimizers. We show that great lifetime extensions can be obtained with respect to standard sequential discharge, as well to previously proposed battery allocation schemes. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2002 | Layout-driven memory synthesis for embedded systems-on-chipabstractMemory-processor integration offers new opportunities for reducing, the energy of a system. In the case of embedded systems, where memory access patterns can typically be profiled at design time, one solution consists of mapping the most frequently accessed addresses onto the on-chip SRAM to guarantee power and performance efficiency. In this work, we propose an algorithm for the automatic partitioning of on-chip SRAMs into multiple banks. Starting from the dynamic execution profile of an embedded application running on a given processor core, we synthesize a multi-banked SRAM architecture optimally fitted to the execution profile. The algorithm computes an optimal solution to the problem under realistic assumptions on the power cost metrics, and with constraints on the number of memory banks. The partitioning algorithm is integrated with the physical design phase into a complete flow that allows the back annotation of layout information to drive the partitioning process. Results, collected on a set of embedded applications for the ARM processor, have shown average energy savings around 34%. Luca Benini, Luca Macchiarulo, Alberto Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | Minimizing memory access energy in embedded systems by selective instruction compressionabstractWe propose a technique for reducing the energy spent in the memory-processor interface of an embedded system during the execution of firmware code. The method is based on the idea of compressing the most commonly executed instructions so as to reduce the energy dissipated during memory access. Instruction decompression is performed on-the-fly by a hardware block located between processor and memory: No changes to the processor architecture are required. Hence, our technique is well suited for systems employing IP cores whose internal architecture cannot be modified. We describe a number of decompression schemes and architectures that effectively trade off hardware complexity and static code size increase for memory energy and bandwidth reduction, as proved by the experimental data we have collected by executing several test programs on different design templates. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2001 | From Architecture to Layout: Partitioned Memory Synthesis for Embedded Systems-on-ChipabstractWe propose an integrated front-end/back-end flow for the automatic generation of a multi-bank memory architecture for embedded systems. The flow is based on an algorithm for the automatic partitioning of on-chip SRAM. Starting from the dynamic execution profile of an embedded application running on a given processor core, we synthesize a multi-banked SRAM architecture optimally fitted to the execution profile. Luca Benini, Luca Macchiarulo, Alberto Macii, Enrico Macii, Massimo Poncino |
DAC | 5 |
| 2001 | Extending lifetime of portable systems by battery schedulingabstractMulti-battery power supplies are becoming popular in electronic appliances of the latest generations, due to economical and manufacturing constraints. Unfortunately, a partitioned battery subsystem is not able to deliver the same amount of charge as a monolithic battery with the same total capacity. In this paper, we define the concept of battery scheduling, we investigate policies for solving the problem of optimal charge delivery, and we study the relationship of such policies with different configurations of the battery subsystem. Results, obtained for different workloads, demonstrate that the choice of the proper scheduling can make, in the best cease, system lifetime as close as 1% of that guaranteed by a monolithic battery of equal capacity. Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DATE | 5 |
| 2001 | Low-energy for deep-submicron address busesabstractArticle Low-energy for deep-submicron address buses Share on Authors: Luca Macchiarulo Politecnico di Torino, Torino, Italy Politecnico di Torino, Torino, ItalyView Profile , Enrico Macii Politecnico di Torino, Torino, Italy Politecnico di Torino, Torino, ItalyView Profile , Massimo Poncino Politecnico di Torino, Torino, Italy Politecnico di Torino, Torino, ItalyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 176–181https://doi.org/10.1145/383082.383127Online:06 August 2001Publication History 22citation275DownloadsMetricsTotal Citations22Total Downloads275Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Luca Macchiarulo, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2001 | Synthesis of power-managed sequential components based oncomputational kernel extractionabstractThis paper introduces a power optimization paradigm for sequential components based on the concept of computational kernel, a highly simplified logic block whose behavior mimics the steady-state behavior of the original specification. We present a flexible framework that supports a number of algorithmic options for carrying out kernel extraction. We first describe an exact symbolic procedure that is applicable to components for which only a functional specification (i.e., the state transition graph) is available. Due to its computational complexity, this procedure is mainly of theoretical interest and it is not usable for large circuits. We then propose two approximate algorithms that can be adopted in practical situations. The first one is simulation-based and it is suitable to cases where input data streams representing typical operation of the component are available. The second approach performs kernel extraction by iteratively refining a structural representation of the component obtained through synthesis. The impact of the power optimization paradigm based on kernel extraction is demonstrated by the results of extensive experimentation carried out on a number of benchmarks of different characteristics and nature. Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2001 | Discrete-time battery models for system-level low-power designabstractFor portable applications, long battery lifetime is the ultimate design goal. Therefore, the availability of battery and voltage converter models providing accurate estimates of battery lifetime is key for system-level low-power design frameworks. In this paper, we introduce a discrete-time model for the complete power supply subsystem that closely approximates the behavior of its circuit-level continuous-time counterpart. The model is abstract and efficient enough to enable event-driven simulation of digital systems described at a very high level of abstraction and that includes, among their components, also the power supply. The model gives the designer the possibility of estimating battery lifetime during system-level design exploration, as shown by the results we have collected on meaningful case studies. In addition, it is flexible and it can thus be employed for different battery chemistries. Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2001 | Parameterized RTL power models for soft macrosabstractWe propose a new power macromodel for usage in the context of register-transfer level (RTL) power estimation. The model is suitable for reconfigurable, synthesizable, soft macros because it is parameterized with respect to the input data size (i.e., bit width) and can also be automatically scaled with respect to different technology libraries and/or synthesis options. The power model is precharacterized once and for all for each soft macro and then adapted to each specific instance by means of a single additional experiment to be performed by the end user. No intellectual-property disclosure is required for model scaling. The proposed model is derived from empirical analysis of the sensitivity of power consumption on input statistics, input data size, and technology. The experiments prove that with limited approximation, it is possible to decouple the effects on power of these three factors. The proposed solution is innovative since no previous macromodel supports automatic technology scaling and yields average estimation errors around 10%. Alessandro Bogliolo, Roberto Corgnati, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2001 | Stream synthesis for efficient power simulation based on spectral transformsabstractOne way of minimizing the time required to perform simulation-based power estimation is that of reducing the length of the input trace to be fed to the simulator. Obviously, the use of a reduced stream may introduce some errors in the estimation results. The generation (or synthesis) of the short input sequence to be used for power simulation should then be carried out in such a way that the resulting error is minimized. Existing techniques exploit the knowledge of some statistical and correlation characteristics concerning the original input trace to generate a reduced stream that closely matches such characteristics. In this paper, we introduce a new stream synthesis method. Its distinguishing feature is the use of spectral analysis based on the discrete Fourier transform to determine a reduced sequence of vectors that enables us to shorten the overall power simulation time at a very limited penalty in accuracy. The effectiveness and the robustness, in terms of estimation accuracy, of the proposed synthesis procedure are demonstrated by the experimental results we have obtained on standard combinational benchmarks for a variety of input streams with different statistical and correlation properties. Data for sequential circuits are also reported. Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | Synthesis of application-specific memories for power optimization in embedded systemsabstractThis paper presents a novel approach to memory power optimization for embedded systems based on the exploitation of data locality. Locations with highest access frequency are mapped onto a small, low-power application-specific memory which is placed close the processor. Although, in principle, a cache may be used to implement such a memory, more efficient solutions may be adopted. We propose an architecture that outperforms (power-wise) different types of cache memories at no penalty in performance. Power savings (averaged over a number of embedded applications running on ARM processors) range from 12% to 68%. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DAC | 4 |
| 2000 | A Discrete-Time Battery Model for High-Level Power EstimationabstractIn this paper, we introduce a discrete-time model for the complete power supply sub-system that closely approximates the behavior of its circuit-level (i.e., HSpice), continuous-time counterpart. The model is abstract and efficient enough to enable event-driven simulation of digital systems described at a very high level of abstraction and that include, among their components, also the power supply. Therefore, it can be successfully used for the purpose of battery life-time estimation during design optimization, as shown by the results we have collected on a meaningful case study. Experiments prove also that the accuracy of our model is very close to that provided by the corresponding Spice-level model. Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DATE | 5 |
| 2000 | Regression-based RTL power models for controllersabstractPower consumption of the controller of an RTL design, though typically smaller than that of the datapath, cannot be neglected. Controller power modeling is a more challenging task than that of datapath modules, because the controller is usually specified in an abstract fashion that has no structural relationship with its final implementation. This paper presents an extensive experimental study on pre-state assignment power models. Starting from a large set of controllers, we study the relation between power and static parameters, such as number of inputs and outputs, as well as dynamic parameters, such as input and output switching activity. We then quantify the loss of accuracy caused by the high level of abstraction of the models, and we show how previously published results, obtained under favorable experimental conditions, have underestimated model inaccuracy. Finally, we perform extensive exploration on a large number of alternative macro-model structures, and we select an optimal model equation. Luca Benini, Alessandro Bogliolo, Enrico Macii, Massimo Poncino, Mihai Surmei |
ACM Great Lakes Symposium on VLSI | 4 |
| 2000 | Supporting system-level power exploration for DSP applicationsabstractSystem-level power exploration requires tools for estimation of the overall power consumed by a system, as well as a detailed breakdown of the consumption of its main functional blocks. We focus on power estimation for data-dominated systems specified as synchronous data-flows and implemented on a single-processor architecture. Our estimator is integrated within the Ptolemy design environment, and provides information to system designers on the power dissipated by every task in a given specification. Power estimation is based on instruction-level power models. We demonstrate the applicability of our tool on a few design examples and target architectures. Luca Benini, Marco Ferrero, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 5 |
| 2000 | A recursive algorithm for low-power memory partitioningabstractMemory-processor integration offers new opportunities for reducing the energy of a system. In the case of embedded systems, one solution consists of mapping the most frequently accessed addresses onto the on-chip SRAM to guarantee power and performance efficiency. This option is especially effective when memory access patterns can be profiled and studied at design time (as in typical real-time embedded systems). Luca Benini, Alberto Macii, Massimo Poncino |
ISLPED | 3 |
| 2000 | A multilevel engine for fast power simulation of realistic inputstreamsabstractPower estimation for validation and sign-off is a critical step in the design process. In this phase, accuracy is a key requirement, but there are hard constraints on the time that can be dedicated to power estimation. Moreover, it is important to estimate the power dissipated by the system while running typical applications, i.e., extremely long streams of validation patterns provided by the designer. The power dissipated by digital systems under realistic input stimuli is not accurately described by a single average value, but by a waveform that shows how power consumption varies over time as the system responds to the inputs. In this paper, we face the problem of obtaining accurate power waveforms for combinational and sequential circuits under typical usage patterns. We propose a multilevel simulation engine that achieves high accuracy in estimating the time-domain power waveform, as well as the average power with high computational efficiency. Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2000 | Architectures and synthesis algorithms for power-efficient businterfacesabstractIn this paper we present algorithms for the synthesis of encoding and decoding interface logic that minimizes the average number of transitions on heavily-loaded global bus lines at no cost in communication throughput (i.e., one word is transmitted at each cycle). The distinguishing feature of our approach is that it does not rely on designer's intuition, but it automatically constructs low-transition activity codes and hardware implementation of encoders and decoders, given information on word-level statistics. We propose an accurate method that is applicable to low-width buses, as well as approximate methods that scale well with bus width. Furthermore, we introduce an adaptive architecture that automatically adjusts encoding to reduce transition activity on buses whose word-level statistics are not known a priori. Experimental results demonstrate that our approaches out-perform specialized low-power encoding schemes presented in the past. Luca Benini, Alberto Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Symbolic optimization of interacting controllers based onredundancy identification and removalabstractThis paper presents a binary decision diagram (BDD)-based algorithm for the optimization of the driven machine, M/sub 2/, of a finite-state machine (FSM) network with cascade connection, M/sub 1//spl rarr/M/sub 2/. The technique we propose relies on redundant faults identification and removal. A fault, f, located into machine M/sub 2/, is redundant with respect to the overall network if the driving machine M/sub 1/ is not able to generate any test sequence for such a fault. When the state transition graph (STG) specifications of the network components are available, the standard way for checking the redundancy condition for the considered fault requires one to first construct the product machine M/sub 2//spl times/M/sub 2//sup F/, where M/sub 2//sup F/ is the faulty FSM, then to connect it to the driving machine, and finally to perform reachability analysis on the composed machine M/sub 1//spl rarr/M/sub 2//spl times/M/sub 2//sup F/. Clearly, the size of such machine limits the applicability of the approach above to systems whose components have a few tens of states at most, even when symbolic traversal algorithms are used. Since we are interested in dealing with networks of larger FSM's (i.e., machines whose STGs can not be represented explicitly), we propose to use the product automaton P'=A/sub 1//spl times/A/sub f/, where A/sub 1/' is the finite automaton (FA) accepting all the output sequences of M/sub 1/, and A/sub f/ is the FA accepting all the test sequences for fault f, instead of machine M/sub 1//spl rarr/M/sub 2//spl times/M/sub 2//sup F/. This simplifies sensibly the task of the reachability analysis program, since A/sub f/ has considerably less states and less edges than the product machine M/sub 2//spl times/M/sub 2//sup F/ and, thus, the size of the BDD representation of its transition relation is much more easily manageable. In addition, differently from other approaches, automaton A/sub 1/' is not required to be deterministic and state minimal. This allows us to avoid the application of determinization and state minimization procedures whose complexity is exponential. We present experimental results For examples (i.e., network of interacting controllers) on which existing optimization methods are not applicable, due to the size of the component FSM's. We also provide a comparison to the data produced by state-of-the-art FSM network optimizers on small benchmarks in order to show the effectiveness of our approach. Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2000 | Glitch power minimization by selective gate freezingabstractThis paper presents a technique for glitch power minimization in combinational circuits. The total number of glitches is reduced by replacing some existing gates with functionally equivalent ones (called F-Gates) that can be "frozen" by asserting a control signal. A frozen gate cannot propagate glitches to its output. Algorithms for gate selection and clustering that maximize the percentage of filtered glitches and reduce the overhead for generating the control signals are introduced. A power-efficient CMOS implementation of F-Gates is also described. An important feature of the proposed method is that it can be applied in place directly to layout-level descriptions; therefore, it guarantees very predictable results and minimizes the impact of the transformation on circuit size and speed. Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 1999 | Kernel-Based Power Optimization of RTL Components: Exact and Approximate Extraction AlgorithmsabstractArticle Free Access Share on Kernel-based power optimization of RTL components: exact and approximate extraction algorithms Authors: L. Benini Università di Bologna, Bologna, Italy 40136 Università di Bologna, Bologna, Italy 40136View Profile , G. De Micheli Stanford University, Stanford, CA Stanford University, Stanford, CAView Profile , E. Macii Politecnico di Torino, Torino, Italay 10129 Politecnico di Torino, Torino, Italay 10129View Profile , G. Odasso Politecnico di Torino, Torino, Italy 10129 Politecnico di Torino, Torino, Italy 10129View Profile , M. Poncino Politecnico di Torino, Torino, Italy 10129 Politecnico di Torino, Torino, Italy 10129View Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 247–252https://doi.org/10.1145/309847.309922Published:01 June 1999Publication History 0citation244DownloadsMetricsTotal Citations0Total Downloads244Last 12 Months6Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Luca Benini, Giovanni De Micheli, Enrico Macii, Giuseppe Odasso, Massimo Poncino |
DAC | 5 |
| 1999 | Synthesis of Low-Overhead Interfaces for Power-Efficient Communication over Wide BusesabstractIn this paper we present algorithms for the synthesis of encoding and decoding interface l o gic that minimizes the average number of transitions on heavily-loaded global bus lines.The approach automatically constructs low-transition activity codes and hardware implementation of encoders and decoders, given information on word-level statistics.We present an accurate method that is applicable to low-width buses, as well as approximate methods that scale well with bus width.Furthermore, we introduce an adaptive architecture that automatically adjusts encoding to reduce t r ansition activity on buses whose word-level statistics are not known a-priori.Experimental results demonstrate that our approach well outperforms low-power encoding schemes presented in the past. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DAC | 4 |
| 1999 | Glitch Power Minimization by Gate FreezingabstractThis paper presents a technique for glitch power minimization in combinational circuits. The total number of glitches is reduced by replacing some existing gates with functionally equivalent ones (called F-gates) that can be "frozen" by asserting a control signal. A frozen gate cannot propagate glitches to its output. An important feature of the proposed method is that it can be applied in-place directly to layout-level descriptions; therefore, it guarantees very predictable results and minimizes the impact of the transformation on circuit size and speed. Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DATE | 5 |
| 1999 | Clustered Table-Based Macromodels for RTL Power EstimationabstractMacromodeling is considered the most effective approach to RTL power estimation. Among the macromodels presented in the literature, table-based ones have overcome some of the limitations of conventional, equation-based solutions. In this paper we propose some enhancements to the basic implementation of table-based macromodels that improve the estimation accuracy while preserving the intrinsic robustness. Roberto Corgnati, Enrico Macii, Massimo Poncino |
Great Lakes Symposium on VLSI | 3 |
| 1999 | Regression-Based Macromodeling for Delay Estimation of Behavioral ComponentsabstractThis paper presents a methodology for delay estimation of hardware components described at the behavioral-level. The basis of the proposed technique is a well-known theoretical result that relates the entropy of a logic function to the delay of a multi-level implementation of the same function. We propose an improved model for delay estimation, and we prove its validity by means of experiments performed on a set of standard benchmarks. Alberto Macii, Enrico Macii, Giuseppe Odasso, Massimo Poncino, Riccardo Scarsi |
Great Lakes Symposium on VLSI | 4 |
| 1999 | Parameterized RTL power models for combinational soft macrosabstractWe propose a new RTL power macromodel that is suitable for re-configurable, synthesizable soft-macros. The model is parameterized with respect to the input data size (i.e., bit-width), and can be automatically scaled with respect to different technology libraries and/or synthesis options. Scalability is obtained through a single additional characterization run, and does not require the disclosure of any intellectual property. The model is derived from empirical analysis of the sensitivity of power on input statistics, input data size and technology. The experiments prove that, with limited approximation, it is possible to de-couple the effects on power of these three factors. The proposed solution is innovative, since no previous macromodel supports automatic technology scaling, and yields estimation errors within 15%. Alessandro Bogliolo, Roberto Corgnati, Enrico Macii, Massimo Poncino |
ICCAD | 4 |
| 1999 | Selective instruction compression for memory energy reduction in embedded systemsabstractWe propose a technique for reducing the energy required by firmware code to ezecute on embedded systema. The method ia based on the idea of compressing the moat commonly ezecuted instructions 80 a ~ to reduce the energy dissipated in memory acceeees. Instruction decompression is performed on the fly by a hardware module located between processor and memory: No changea to the processor architecture ore required. Hence, our technique is well-suited for systema employing IP cone whose internal architecture cannot be modified. We describe a number of decompnaaion achemea and architec-tura that effectively trade off hardware complezity for memory energy and bandwidth nduction, aa proved by experimental data collected by executing aeveml sample programs. 1 Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 1999 | Automatic Synthesis of Large Telescopic Units Based on Near-Minimum Timed SupersettingabstractIn high-performance systems, variable-latency units are often employed to improve the average throughput when the worst-case delay exceeds the cycle time. Traditionally, units of this type have been hand-designed. In this paper, we propose a technique for the automatic synthesis of variable-latency units that is applicable to large data-path modules. We define and study an optimization problem, timed supersetting, whose solution is at the kernel of the procedure for automatic generation of variable-latency units. We contribute a new algorithm for solving timed supersetting in the most difficult case, that is, when the timing behavior of the circuit is expressed through an accurate delay model. The proposed solution overcomes the computational limitations of previous approaches and its robustness is experimentally demonstrated by obtaining high-throughput, variable-latency implementations for all the largest circuits in the Iscas '85 and Iscas '89 benchmark suites, as well as for some realistic, high-performance arithmetic units. Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino |
IEEE Trans. Computers | 6 |
| 1999 | Symbolic synthesis of clock-gating logic for power optimization of synchronous controllersabstractRecent results have shown that dynamic power management is effective in reducing the total power consumption of sequential circuits. In this paper, we propose a bottom-up approach for the automatic extraction and synthesis of dynamic power management circuitry starting from structural logic-level specifications. Our techniques leverage the compact BDD-based representation of Boolean and pseudo-Boolean functions to detect idle conditions where the clock can be stopped without compromising functional correctness. Moreover, symbolic techniques allow accurate probabilistic computations; in particular, they enable the use of non-equiprobable primary input distributions, a key step in the construction of models that match the behavior of real hardware devices with a high degree of fidelity. The results are encouraging, since power savings of up to 34% have been obtained on standard benchmark circuits. Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 1998 | Computational Kernels and their Application to Sequential Power OptimizationabstractWe introduce a new sequential optimization paradigm based on the extraction of computational kernels, i.e., logic blocks whose behavior mimics the steady-state behavior of the original circuit. We present a procedure for the automatic extraction of such kernels directly from the gate-level description of the design. The advantage of this solution with respect to extraction algorithms based on STG analysis is that it can be applied to large circuits, since it does not require to manipulate the STG specification. Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino |
DAC | 6 |
| 1998 | Power Estimation of Behavioral DescriptionsabstractThis paper presents a methodology for power estimation of designs described at the behavioral-level as the interconnection of functional modules. The input/output behavior of each module is implicitly stored using BDDs, and the power consumed by the network is estimated using a novel and accurate entropy-based approach. As a demonstration example, we have used the proposed power estimation technique to evaluate and compare the effects of some architectural transformations applied to a reference design specification on the power dissipation of the corresponding implementations. Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 1998 | Timed Supersetting and the Synthesis of Telescopic UnitsabstractIn high-performance systems, variable-latency units are often employed to improve the average throughput when the worst-case delay exceeds the cycle time. Although such units have traditionally been hand-designed, recent results have shown that variable-latency units can be automatically generated. Unfortunately, the existing synthesis procedure has limited applicability due to its computational complexity. In this work, we define and study an optimization problem, timed supersetting, whose solution is at the kernel of the procedure for automatic generation of variable-latency units. We contribute a new algorithm for solving timed supersetting in the most difficult case, that is, when the timing behaviour of the circuits is expressed through an accurate delay model. The proposed solution overcomes the complexity limitation of previous approaches, and its robustness is experimentally demonstrated by obtaining high-throughput, variable-latency implementations for all the largest circuits in the Iscas'85 and Iscas'89 benchmark suites. Luca Benini, Giovanni De Micheli, Antonio Lioy, Enrico Macii, Giuseppe Odasso, Massimo Poncino |
Great Lakes Symposium on VLSI | 6 |
| 1998 | Reducing Power Consumption of Dedicated Processors Through Instruction Set EncodingabstractWith the increased clock frequency of modern, high-performance processors (over 500 MHz, in some cases), limiting the power dissipation has become the most stringent design target. It is thus mandatory for processor engineers to resort to a large variety of optimization techniques to reduce the power requirements in the hot zones of the chip. In this paper, we focus on the power dissipated By the instruction fetch and decode logic, a portion of the processor architecture where a lot of capacitance switching normally takes place. We propose a methodology for determining an encoding of the instruction set that guarantees the minimization of the number of bit transitions occurring inside the registers of the pipeline stages involved in instruction fetching and decoding. The assignment of the binary patterns to the op-codes is driven by the statistics concerning instruction adjacency collected through instruction-level simulation of typical software applications; therefore, the technique is best exploited when applied to encode the instruction set of core processors and microcontrollers, since components of these types ore commonly used to execute fixed portions of machine code within embedded systems. We illustrate the effectiveness of the methodology through the experimental data we have obtained on an existing microprocessor. Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino |
Great Lakes Symposium on VLSI | 5 |
| 1998 | Symbolic algorithms for layout-oriented synthesis of pass transistor logic circuitsabstractThis paper presents a nouel methodology for synthesizing PTL circuifg, whose disfincfiue feafures are the use of a symbolic algorifhm for the covem.ng of fhe initial network in ferms of PTL cells, and the eqloifation of layout-level ama and delay models dun.ng fhe selection of fhe be~t couem.ngsolufion.The results produced by the synthesis procedure on the full guite of fhe kcas'85 combinational circuifs are very encouraging. Fabrizio Ferrandi, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi, Fabio Somenzi |
ICCAD | 4 |
| 1998 | Stream synthesis for efficient power simulation based on spectral transformsabstractIn this paper, we present a power estimation technique for control-flow intensive designs that is tailored towards driving iterative high-level synthesis systems, where hundreds of architectural trade-offs are explored and compared. Our method is fast and relatively accurate. The algorithm utilizes the behavioral information to extract branch probabilities, and uses these in conjunction with switching activity and circuit capacitance information, to estimate the power consumption of a given architecture. Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
ISLPED | 3 |
| 1998 | Telescopic units: a new paradigm for performance optimization of VLSI designsabstractThis paper introduces a novel optimization paradigm for increasing the throughput of digital systems. The basic idea consists of transforming fixed-latency units into variable-latency ones that run with a faster clock cycle. The transformation is fully automatic and can be used in conjunction with traditional design techniques to improve the overall performance of speed-critical units. In addition, we introduce procedures for reducing the area overhead of the modified units, and we formulate an algorithm for automatically restructuring the controllers of the data paths in which variable-latency units have been introduced. Results, obtained on a large set of benchmark circuits, show an average throughput improvement exceeding 27%, at the price of a modest area increase (less than 8% on average). Luca Benini, Enrico Macii, Massimo Poncino, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1998 | Power optimization of core-based systems by address bus encodingabstractThis paper presents a solution to the problem of reducing the power dissipated by a digital system containing an intellectual proprietary core processor which repeatedly executes a special-purpose program. The proposed method relies on a novel, application-dependent low-power address bus encoding scheme. The analysis of the execution traces of a given program allows an accurate computation of the correlations that may exist between blocks of bits in consecutive patterns; this information can be successfully exploited to determine an encoding which sensibly reduces the bus transition activity. Experimental results, obtained on a set of special-purpose applications, are very satisfactory; reductions of the bus activity up to 64.8% (41.8% on average) have been achieved over the original address streams. In addition, data concerning the quality and the performance of the automatically synthesized encoding/decoding circuits, as well as the results obtained for a realistic core-based design, indicate the practical usefulness of the proposed power optimization strategy. Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Stefano Quer |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 1997 | Telescopic Units: Increasing the Average Throughput of Pipelined Designs by Adaptive Latency ControlabstractThis paper presents a technique, alternative to performance-drivensynthesis, that allows to drastically increase the averagethroughput of combinational logic blocks by transforming fixed-latencyunits into variable-latency ones that run with a fasterclock cycle.The transformation is fully automatic and can beused in conjunction with traditional design techniques, such aspipelining, to improve the overall performance of speed-criticalsystems.Results, obtained on a large set of benchmark circuits,are very promising. Luca Benini, Enrico Macii, Massimo Poncino |
DAC | 3 |
| 1997 | Accurate Entropy Calculation for Large Logic Circuits Based on Output ClusteringabstractEntropy-based estimation is a promising approach to the problem of predicting the power dissipated by a digital system for which an architectural description is available. For achieving good performance of the power estimation tool, an accurate computation of the input and output entropies of the Boolean functions implemented by the circuit is essential. For small designs, the calculation can be carried out exactly, thanks to the compact representation and ease of manipulation of Boolean and pseudo-Boolean functions provided by BDD-like data structures. For large circuits, on the other hand, resorting to approximate computations is mandatory. Techniques to determine an upper bound on the exact entropy values have been developed in the recent past. Unfortunately, the results provided by such techniques are, in some ceases, not satisfactory; in other words, the assumptions made to simplify the calculation-total absence of correlation among the output signals of a circuit are in many cases too strong to guarantee a reasonable lightness of the approximate entropy values to the exact ones. In this paper, we propose a method to determine the entropy of large logic circuits with a level of accuracy which is far beyond the one provided by existing approaches. We partition the set of output signals according to the information about the functional correlations that may exist among such signals, and we compute the approximate entropy values after performing output clustering. Experimental results, obtained on a large collection of benchmarks, are very promising. Antonio Lioy, Enrico Macii, Massimo Poncino, Massimo Rossello |
Great Lakes Symposium on VLSI | 3 |
| 1997 | Fast power estimation for deterministic input streamsabstractThe power dissipated by digital systems under realistic input stimuli is not accurately described by a single average value, but by a waveform that shows how power consumption varies over time as the system responds to the inputs. We face the problem of obtaining accurate power waveforms for combinational and sequential circuits under typical usage patterns. We propose a multi level simulation engine that achieves high accuracy in estimating the average power as well as the time domain power waveform with high computational efficiency. Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
ICCAD | 4 |
| 1997 | System-level power optimization of special purpose applications: the beach solutionabstractArticle Free Access Share on System-level power optimization of special purpose applications: the beach solution Authors: Luca Benini Stanford University, Computer Systems Laboratory, Stanford, CA Stanford University, Computer Systems Laboratory, Stanford, CAView Profile , Giovanni De Micheli Stanford University, Computer Systems Laboratory, Stanford, CA Stanford University, Computer Systems Laboratory, Stanford, CAView Profile , Enrico Macii Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129 Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129View Profile , Massimo Poncino Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129 Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129View Profile , Stefano Quer Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129 Politecnico di Torino, Dip. di Automatica e Informatica, Torino, Italy 10129View Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 24–29https://doi.org/10.1145/263272.263277Published:01 August 1997Publication History 33citation259DownloadsMetricsTotal Citations33Total Downloads259Last 12 Months17Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Luca Benini, Giovanni De Micheli, Enrico Macii, Massimo Poncino, Stefano Quer |
ISLPED | 4 |
| 1996 | Symbolic Optimization of FSM Networks Based on Sequential ATPG TechniquesabstractThis paper presents a novel optimization algorithm for FSM networks that relies on sequential test generation and redundancy removal. The implementation of the proposed approach, which is based on the exploitation of input don't care sequences through regular language intersection, is fully symbolic. Experimental results, obtained on a large set of standard benchmarks, improve over the ones of state-of-the-art methods. Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto |
DAC | 4 |
| 1996 | Test Generation for Networks of Interacting FSMs Using Symbolic TechniquesabstractThis paper presents a new testing strategy for networks of interacting FSMs. The approach allows us to generate test patterns for faults in the network by separately handling the network's components. The proposed algorithms are fully symbolic; therefore, they allow the manipulation of large designs. Experimental results, though preliminary, are promising. Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto |
Great Lakes Symposium on VLSI | 4 |
| 1996 | Exact Computation of the Entropy of a Logic CircuitabstractComputing the entropy of a digital circuit has proved to be very useful for several applications in the area of VLSI system design. Recently, a method for entropy calculation has been used in the context of power estimation for logic circuits described at the register-transfer level. The technique has shown to be reasonably effective concerning the trade-off between the accuracy of the estimates produced and the execution time. However, the assumptions required to make the computation feasible are such that the obtained results are approximate. In this paper, we propose a symbolic algorithm for the exact calculation of the entropy of a logic circuit which is able to handle reasonably large examples without introducing any approximation. We present experimental data on standard benchmark designs in order to show the effectiveness of the new method; in addition, we compare our results to the ones obtained with the approximate approach. As a result, we observe a marginal penalty in the performance of the symbolic procedure; on the other hand, accuracy in the calculation increases significantly. Enrico Macii, Massimo Poncino |
Great Lakes Symposium on VLSI | 2 |
| 1996 | Enhancing FSM Traversal by Temporary Re-EncodingabstractSynthesis and optimization of large finite-state machines has improved dramatically over the last few years with the introduction and rapid improvement of symbolic-state manipulation techniques. The algorithms efficiently visit each reachable state in the machine while computing and storing information about these states. We propose a new technique for improving the efficacy of traversal algorithms: re-encoding the states of the machine to more efficiently represent state sets or state transitions, or to more efficiently compute the next set of states. Our technique can be embedded in existing traversal algorithms. Experiments reveal that re-encoding can indeed reduce the time and/or space required for traversal. Gianpiero Cabodi, Luciano Lavagno, Enrico Macii, Massimo Poncino, Stefano Quer, Paolo Camurati, Ellen Sentovich |
ICCD | 4 |
| 1996 | Automatic state space decomposition for approximate FSM traversal based on circuit analysisabstractExploiting circuit structure is a key issue in the implementation of algorithms for state space decomposition when the target is approximate FSM traversal. Given the gate-level description of a sequential circuit, the information about its structure can be captured by evaluating the affinity between pairs or groups of latches. Two main factors have to be considered in carrying out the structural analysis of a sequential circuit: latch connectivity and latch correlation. The first one takes into account the mutual dependency of each memory element on the others; the second one tells us how related are the functions realized by the logic feeding each latch. In this paper we estimate the affinity of two latches by combining these two factors, and we use this measure to formulate the state space decomposition problem as a graph partitioning problem. We propose an algorithm to automatically determine "good" partitions of the latch set which induce state space decomposition, and we present approximate FSM traversal and logic optimization results for the largest ISCAS'89 sequential benchmarks. Gary D. Hachtel, Enrico Macii, Massimo Poncino, Fabio Somenzi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 1995 | Computing the Maximum Power Cycles of a Sequential CircuitabstractThis paper studies the problem of estimating worst case power dissipation in a sequential circuit.We approach this problem by nding the maximum average weight cycles in a weighted directed g r aph.In order to handle practical sized examples, we use symbolic methods, based o n A lgebraic Decision Diagrams (ADDs), for computing the maximum average length cycles as well as the number of gate transitions in the circuit, which is necessary to construct the weighted directed g r aph. Srilatha Manne, Abelardo Pardo, R. Iris Bahar, Gary D. Hachtel, Fabio Somenzi, Enrico Macii, Massimo Poncino |
DAC | 7 |
| 1995 | Estimating worst-case power consumption of CMOS circuits modeled as symbolic neural networksabstractIn this paper we propose a new approach to the problem of estimating worst-case power consumption of CMOS combinational circuits based on neural models. Given the gate level description of a circuit, we build the corresponding neural network, we store it, we calculate the energy dissipated by the network and, finally, we derive the power dissipated by the original circuit. All the operations above are executed in the symbolic domain; that is, Algebraic Decision Diagrams are used to represent and manipulate the graph specification of the neural network modeling the circuit. We present preliminary results to show the feasibility of the method. Enrico Macii, Massimo Poncino |
Great Lakes Symposium on VLSI | 2 |
| 1995 | Using symbolic Rademacher-Walsh spectral transforms to evaluate the correlation between Boolean functionsabstractThe use of symbolic techniques to store integer-valued functions has been shown to be extremely effective in handling both transform matrices and spectral representations of large Boolean functions. In this paper we propose a novel application of symbolic Rademacher-Walsh spectral transforms to the evaluation of Boolean function correlation. In particular, we present an ADD-based algorithm to compute the agreement between two Boolean functions starting from their spectral representations. The method, operating in the transform domain, has appeared to be more advantageous than traditional approaches, using operations in the Boolean domain, concerning both memory occupation and execution time on some classes of functions. Enrico Macii, Massimo Poncino |
Great Lakes Symposium on VLSI | 2 |
| 1995 | Using connectivity and spectral methods to characterize the structure of sequential logic circuits
Enrico Macii, Massimo Poncino |
Microprocess. Microprogramming | 2 |
| 1994 | An ADD-based algorithm for shortest path back-tracing of large graphsabstractSymbolic computation techniques play a fundamental role in logic synthesis and formal hardware verification algorithms. Recently, Algebraic Decision Diagrams, i.e., BDDs with a set of constant values different to the set /spl lcub/0,1/spl rcub/, have been used to solve general purpose problems, such as matrix multiplication, shortest path calculation, and solution of linear systems, as well as logic synthesis and formal verification problems, such as timing analysis, probabilistic analysis of finite state machines, and state space decomposition for approximate finite state machine traversal. ADD-based procedures for single-source and all-pairs shortest path weight calculation have appeared to be very effective for the manipulation of large graphs (over 10/sup 27/ vertices and 10/sup 36/ edges). However, for those procedures to be applicable to real problems, for example flow network problems, computing only shortest path weights is not enough; what it is needed is an algorithm that, given the weight of a shortest path between two vertices of a graph, actually determines the sequence of vertices belonging to the shortest path. This paper proposes a symbolic algorithm to execute shortest path back-tracing which exploits the compactness of the ADD data structure to handle large graphs.> R. Iris Bahar, Gary D. Hachtel, Abelardo Pardo, Massimo Poncino, Fabio Somenzi |
Great Lakes Symposium on VLSI | 4 |
| 1994 | Re-encoding sequential circuits to reduce power dissipation
Gary D. Hachtel, Mariano Hermida de la Rica, Abelardo Pardo, Massimo Poncino, Fabio Somenzi |
ICCAD | 4 |
| 1994 | A Structural Approach to State Space Decomposition for Approximate Reachability AnalysisabstractExploiting circuit structure is a key issue in the implementation of algorithms for state space decomposition when the target is approximate FSM traversal. Given the gate-level description of a sequential circuit, the information about its structure can be captured by evaluating the affinity between pairs or groups of latches. Two main factors have to be considered in carrying out the structural analysis of a sequential circuit: latch connectivity and latch correlation. We estimate the affinity of two latches by combining these two factors, and we use this measure to translate the state space decomposition problem into a graph partitioning problem. Traversal results obtained on the largest ISCAS'89 benchmarks show the effectiveness of the method.> Gary D. Hachtel, Enrico Macii, Massimo Poncino, Fabio Somenzi |
ICCD | 4 |
| 1993 | On the Resetability of Synchronous Sequential Circuits
Antonio Lioy, Massimo Poncino |
ISCAS | 2 |
| 1993 | A study of the resetability of synchronous sequential circuits
Antonio Lioy, Massimo Poncino |
Microprocess. Microprogramming | 2 |
| 1991 | A hierarchical multi-level test generation systemabstractThe authors describe a multi-level ATPG system which handles circuits consisting of 'switch' transistors, Boolean gates, and open-output gates (i.e., tristate, open-collector, open-emitter). Both combinational and synchronous sequential circuits are supported, with provision for full-scan, partial-scan, and non-scan design. The most remarkable features of the system are an unified approach to test generation (suitable to compiled-code implementation) and automatic extraction of hierarchy.> Antonio Lioy, Massimo Poncino |
Great Lakes Symposium on VLSI | 2 |