VLDB 2026 Research / reviewers in the wild / expert
Alberto Macii
dblp:96/886
· DBLP profile ↗
77ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-8869-5710ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 16 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural networks for estimating surface solar irradiation from satellite images
Raimondo Gallo, Marco Castangia, Alberto Macii, Edoardo Patti, Alessandro Aliberti |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | An online reinforcement learning approach for HVAC controlabstractHeating, Ventilation and Air Conditioning (HVAC) optimization for energy consumption reduction is becoming ever more a topic of the utmost environmental and energetic concerns. The two most employed methodologies for optimizing HVAC systems are Model Predictive Control (MPC) and Reinforcement Learning (RL). This paper compares three different RL approaches to HVAC optimization: one based on a black-box system identification model trained on historical data, one based on a white-box model of a building and one online method based on an imitation learning pretraining phase on historical data. The three approaches are compared with a literature baseline and an EnergyPlus baseline. Results show that the overall best method in terms of energy consumption reduction (65% decrease) and thermal comfort increase (25% increase) is the approach based on the white-box model. However, the proposed methodology, based on online and imitation learning, demonstrates remarkable efficiency, achieving comparable improvements in energy consumption after just a few months of online training, while maintaining thermal comfort at around the same level as the baseline. These results prove a direct online RL approach, which avoid the use of costly simulations, can provide a reliable and inexpensive solution to the problem of HVAC optimization. Francesco M. Solinas, Alberto Macii, Edoardo Patti, Lorenzo Bottaccioli |
Expert Syst. Appl. | 2 |
| 2024 | A Simulation Framework for Urban Electric Mobility Based on Limited Widespread Data and Spatial InformationabstractElectric Vehicles (EVs) provide an alternative to traditional mobility and a sustainable means of transportation. As a result, electric vehicle sales are increasing across Europe, prompting researchers to wonder about the impact of EVs on smart grids. The proposed framework simulates users’ activities, highly characterising individual behaviour using Time Use Survey (TUS) data to estimate EV usage and consumption. Then, for each trip, the routes between origin and destination are determined, simulating in separate modules i) the driving behaviour, ii) the motion of the EV and its discharge considering spatial data and iii) the charge considering users’ preference. Thanks to the spatial information openly available, it is possible to characterise the simulation and improve EV consumption estimation. Different scenarios are analysed to demonstrate the versatility of the proposed framework by exploiting its modularity. The individuals’ heterogeneity is considered by using an agent-oriented approach. Furthermore, the simulation proceeds on a time-step basis to enable the use of the simulator in a co-simulation environment for future purposes, such as the integration of power networks. The results indicate that achieving a high realism with limited, i.e. containing scarce data for the problem under study, is feasible, enabling researchers to make informed decisions about future mobility. Claudia De Vizia, Daniele Salvatore Schiera, Alberto Macii, Edoardo Patti, Lorenzo Bottaccioli |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Nonlinear Two-Parameter Model for the Spatial Analysis of Solar IrradiationabstractNowadays, energy estimation in various application areas is a major research topic. Additionally, various machine learning techniques, especially regression methods and artificial neural networks, have been developed in recent decades to improve the accuracy of such estimates. This article presents a nonlinear compact regression model for estimating the yearly solar irradiation in Africa and Europe by considering only the latitude and mean temperature of the locations as input parameters. The definition of the values of the coefficients is based on the least-square method constrained by the maximum absolute error. The results of 16 conventional regression models, using the same number of predictors, were compared with the result of the model proposed. Our model minimizes the root mean square error by at least 15%. Alberto Bocca, Alberto Macii, Enrico Macii |
COMPSAC | 2 |
| 2022 | Solar radiation forecasting with deep learning techniques integrating geostationary satellite images
Raimondo Gallo, Marco Castangia, Alberto Macii, Enrico Macii, Edoardo Patti, Alessandro Aliberti |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Forecasting the Grid Power Demand of Charging Stations from EV Drivers' AttitudeabstractIn recent years there has been a significant increase in the production of electric vehicles (EVs), in the global strive to reduce polluting gases produced by conventional fossil-fuel driven vehicles. Therefore, many optimization algorithms have been proposed for EV mobility and the charging of battery packs in the stations connected to power grids. However, there are situations in which experimental results are not sufficient, and simulations are needed. In this work, we address the effects of the charge demands of an EV fleet on the grid by considering the attitude of EV drivers, and especially their range anxiety. This influences their decision of when to recharge the battery pack. To this end, an agent-based model has been developed for the simulation of a power grid considering different scenarios based mainly on the state of charge (SOC) of battery packs at the time of the charging requests of EVs at service stations. The results indicate that in general a high battery SOC at the beginning of charging increases the probability of reaching higher power peaks on the grid. Alberto Bocca, Alberto Macii, Enrico Macii |
COMPSAC | 2 |
| 2021 | Optimizing Quality Inspection and Control in Powder Bed Metal Additive Manufacturing: Challenges and Research DirectionsabstractOne of the key targets of Industry 4.0 and digital production, in general, is the support of faster, cleaner, and increasingly customizable manufacturing processes. Additive manufacturing (AM) is a natural fit in this context, as it offers the possibility to produce complex parts without the design constraints of traditional manufacturing routes, typically reducing both material waste and time to market. Nonetheless, the lack of repeatability of the manufacturing process, which typically translates into a lack of reproducibility and reliability of the quality of the final products compared to traditional subtractive technologies, is currently one of the major barriers to the widespread adoption of AM in mass production. To overcome this limitation, there are growing efforts in recent years toward better integration of advanced information technologies into AM, exploiting the layer-by-layer nature of the build. The consequence of these efforts is twofold: 1) the integration of advanced sensing technologies into the AM systems, making possible the in situ monitoring of huge amounts of data at multiple time scales and resolutions and 2) the ever-increasing role of data-driven approaches [especially machine learning (ML)] in the analysis of such data to provide real-time quality monitoring and process optimization. This article introduces and reviews the key technological developments of this phenomenon, with a special focus on metal powder bed fusion (PBF) technologies that are attracting the highest attention by the industrial AM community. After introducing the main manufacturing quality issues and needs that have to be developed and optimized, we provide a wide overview of the latest progress of in situ monitoring and control in metal PBF, with special regards to sensing technologies and ML approaches. Finally, we identify the open challenges and future research directions in this field. Santa Di Cataldo, Sara Vinco, Gianvito Urgese, Flaviana Calignano, Elisa Ficarra, Alberto Macii, Enrico Macii |
Proc. IEEE | 6 |
| 2018 | Fundamental Feature Extraction of the Battery Charge Phase from Product DataabstractThe modeling of electrical energy storage systems has received a lot of attention during the last two decades, as a consequence of the remarkable increase of battery-powered devices. However, automated extraction of a cell/pack characteristics from available product data, has been proposed only more recently. Although the automated modeling of a battery discharge performance is already well reported in the literature, extraction of the basic features of the charging phase is less common since most commercial chargers work at the standard constant current-constant voltage (CC-CV) protocol at fixed working conditions. Nevertheless, many rechargeable batteries allow different charging current rates. Therefore, for a full modeling and analysis of the battery behavior during the charging phase, the basic features, like the internal resistance and the open circuit voltage (Rch and VOCch), should be modeled. Unfortunately, only a few data regarding the charging phase are available from time-based plots in datasheets. This work presents a method for modeling the fundamental characteristics of the charging phase of a battery, starting from typical multi-plot time-based charts. Alberto Bocca, Yukai Chen, Alberto Macii, Massimo Poncino |
ISCAS | 3 |
| 2018 | Battery-Aware Energy Model of Drone Delivery TasksabstractDrones are becoming increasingly popular in the commercial market for various package delivery services. In this scenario, the mostly adopted drones are quad-rotors (i.e., quadcopters). The energy consumed by a drone may become an issue, since it may affect (i) the delivery deadline (quality of service), (ii) the number of packages that can be delivered (throughput) and (iii) the battery lifetime (number of recharging cycles). It is thus fundamental try to find the proper compromise between the energy used to complete the delivery and the speed at which the quadcopter flies to reach the destination. In order to achieve this, we have to consider that the energy required by the drone for completing a given delivery task does not exactly correspond to the energy requested to the battery, since the latter is a non-ideal power supply that is able to deliver power with different efficiencies depending on its state of charge. In this paper, we demonstrate that the proposed battery-aware delivery scheduling algorithm carries more packages than the traditional delivery model with the same battery capacity. Moreover, the battery-aware delivery model is 17% more accurate than the traditional delivery model for the same delivery scheme, which prevents the unexpected drone landing. Donkyu Baek, Yukai Chen, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2016 | A compact IGBT electro-thermal model in Verilog-A for fast system-level simulationabstractIn this work, the implementation of a compact electro-thermal model of a trench gate field-stop IGBT device is presented. The primary target is to address system-level mixed signal multi-domain simulation where the simulation time is a key factor for the usefulness of results. The methodology adopted provides a model suitable for use on both SPICE-only simulators, where thermal behavior is mapped to electrical quantities, and on a multi-domain environment, where heterogeneous physical domains are managed by a simulation engine. A comparison with real device characteristics is presented to show model accuracy and finally a simulation of an inverter circuit is investigated to highlight achievable performances. Davide Lena, Michelangelo Grosso, Alberto Bocca, Alberto Macii, Salvatore Rinaudo |
IECON | 4 |
| 2016 | A Li-Ion Battery Charge Protocol with Optimal Aging-Quality of Service Trade-offabstractThe reduction of usable capacity of rechargeable batteries can be mitigated during the charge process by acting on some stress factors, namely, the average state-of-charge (SOC) and the charge current. Larger values of these quantities cause an increased degradation of battery capacity, so it would be desirable to keep both as low as possible, which is obviously in contrast with the objective of a fast charge. However, by exploiting the fact that in most battery-powered systems the time during which it is plugged for charging largely exceeds the time required to charge, it is possible to devise appropriate charge protocols that achieve a good balance between fast charge and aging. Yukai Chen, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2015 | An aging-aware battery charge scheme for mobile devices exploiting plug-in time patternsabstractThe aging of a rechargeable battery is mainly due to stress during charge-discharge cycles. Although the discharge phase is difficult to control, the charging phase can be performed in a specific way in order to mitigate the aging of the battery during its usage. It therefore becomes important to select the correct charging algorithm. In the case of mobile systems, equipped mainly with lithiumion batteries, the standard widely adopted for charging a battery is the typical constant current/constant voltage (CC-CV) protocol usually based on a linearly regular charge process. In this work, we propose a charging protocol based on the standard CC-CV method in which the charge start time and the value of the charging current can be programmed in such a way that the aging of the battery is mitigated. To validate this charging scheme we use an aging model that includes the charge/discharge current among the major parameters, and an analytical macro-model for the CC-CV charge time analysis. Alberto Bocca, Alessandro Sassone, Alberto Macii, Enrico Macii, Massimo Poncino |
ICCD | 3 |
| 2015 | An equation-based battery cycle life model for various battery chemistriesabstractThe evaluation of the cycle life of batteries is an essential task in the assessment of the reliability and cost of battery-operated devices. Several compact cycle life models have been proposed in the literature, that exhibit a general trade-off between generality and accuracy. Some models are based on a compact equation derived from experimental data and try to extract a general relationship between cycle life and the relevant parameters (mostly the depth of discharge), but suffer from poor accuracy. At the other extreme, more accurate models, based on incorporating the aging effect into an equivalent circuit, tend to be focused on a specific device and are seldom applicable to another battery. In this work we propose an equation-based model that tries to overcome the accuracy limits of previous similar models. The model parameters are obtained by fitting the curve based on information reported in datasheets, and can be adapted (with different accuracy levels) to the amount of available information. We applied the model to various commercial batteries for which full information on their cycle life is available. Results show an average estimation error, in terms of the number of cycles, generally smaller than 10%, which is consistent with the typical tolerance provided in the datasheets, and much lower than previous equation-based models. Alberto Bocca, Alessandro Sassone, Donghwa Shin, Alberto Macii, Enrico Macii, Massimo Poncino |
VLSI-SoC | 4 |
| 2014 | Modeling of the charging behavior of li-ion batteries based on manufacturer's dataabstractThe market of portable devices, wireless sensors, electric vehicles and storage systems has grown enormously in recent years. As a consequence, batteries and related technologies have become one of the major topics for researchers. Due to the large variety of applications in which batteries are involved, battery modeling is becoming an extremely important research topic. This relevance is witnessed by the number of papers addressing battery modeling. Alessandro Sassone, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2014 | Automated generation of battery aging models from datasheetsabstractThe de-facto standard approach in battery modeling consists of the definition of a generic model template in terms of an equivalent electric circuit, which is then populated either using data obtained from direct measurements on actual devices or by some extrapolation of battery characteristics available from datasheets. These models typically describe only intra-cycle effects, that is, those manifesting within a single charge/discharge cycle of a battery. However, basic battery dynamics, during a single discharge, cannot provide a true estimate of the actual lifetime of the battery, e.g., how its usability decreases due to long-term and irreversible effects, such as the fading of capacity due to aging or to repeated cycling. While some solutions in the literature provide answers to this problem by proposing suitable models for these effects, they do not provide solutions for how to incorporate them into a generic model template. In this work we propose a method to include inter-cycle battery effects into a reference model template in an automated way, and using solely data reported by battery manufacturers. Flexibility and accuracy of the proposed strategy are demonstrated by modeling a commercial lithium iron phosphate battery, whose datasheet provides long-term capacity fading information. Massimo Petricca, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ICCD | 4 |
| 2014 | A compact macromodel for the charge phase of a battery with typical charging protocolabstractAvailability of a simulation model of a battery is one of the most important requisites in the system-level design of battery-powered systems. The vast majority of the models describe the discharge behavior of the battery; so far, the estimation of charging time has been in fact only marginally studied because the charging phase is regarded as a relatively controlled process compared to discharge. In this paper, we present a compact macro-model for the estimation of charging time under the most widely used charge protocol, i.e., Constant Current-Constant Voltage (CC-CV). This model is derived under the consideration of the context of the existing models including the well-known Peukert's law and equivalent electric circuits. The estimation result with the proposed model based on the manufacturer's data of commercial Li-ion batteries shows fair accuracy, especially when compared to estimates on parameters extracted from discharge characteristics. Donghwa Shin, Alessandro Sassone, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2013 | An automated framework for generating variable-accuracy battery models from datasheet informationabstractModels based on an electrical circuit equivalent have become the most popular choice for modeling the behavior of batteries, thanks to their ease of co-simulation with other parts of a digital system. Such circuit models are actually model templates: the specific values of their electrical elements must be derived by the analysis of the specific battery devices to be modeled. This process requires either to measure the battery characteristics or to derive them from the datasheet. In the latter case, however, very often not all information are available and the model fitting becomes then unfeasible. In this paper we present a methodology for deriving, in a semi-automatic way, circuit equivalent battery models solely from data available in a battery datasheet. In order to account for the different amount of information available, we introduce the concept of “level” of a model, so that models with different accuracy can be derived depending on the available data. The methodology requires only minimal intervention by the designer and it automatically generates MATLAB models once the required data for the corresponding model level are transcribed from the datasheet. Simulation results show that our methodology allows to accurately reconstruct the information reported in the datasheet as well as to derive missing ones. Massimo Petricca, Donghwa Shin, Alberto Bocca, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2013 | Layout-Driven Post-Placement Techniques for Temperature Reduction and Thermal Gradient MinimizationabstractWith the continuing scaling of CMOS technology, on-chip temperature and thermal-induced variations have become a major design concern. To effectively limit the high temperature in a chip equipped with a cost-effective cooling system, thermal specific approaches, besides low power techniques, are necessary at the chip design level. The high temperature in hotspots and large thermal gradients are caused by the high local power density and the nonuniform power dissipation across the chip. With the objective of reducing power density in hotspots, we propose two placement techniques that spread cells in hotspots over a larger area. Increasing the area occupied by the hotspot directly reduces its power density, leading to a reduction in peak temperature and thermal gradient. To minimize the introduced overhead in delay and dynamic power, we maintain the relative positions of the coupling cells in the new layout. We compare the proposed methods in terms of temperature reduction, timing, and area overhead to the baseline method, which enlarges the circuit area uniformly. The experimental results showed that our methods achieve a larger reduction in both peak temperature and thermal gradient than the baseline method. The baseline method, although reducing peak temperature in most cases, has little impact on thermal gradient. Wei Liu 0016, Andrea Calimera, Alberto Macii, Enrico Macii, Alberto Nannarelli, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Investigating the effects of Inverted Temperature Dependence (ITD) on clock distribution networksabstractThe aggressive scaling of CMOS technology toward nanometer lengths contributed to the surfacing of many effects that were not appreciable at the micrometer regime. Among them, Inverted Temperature Dependence (ITD) is certainly the most unusual. It manifests itself as a speed up of CMOS gates when the temperature increases, resulting in a reversal of the worst-case condition, i.e., CMOS gates show the largest delay at low temperatures. On the other hand, for metal interconnects an high temperature still holds as worst case condition. The two contrasting behaviors may invalidate the results obtained through standard design flow which do not consider temperature as an explicit variable in their optimizations. In this paper we focus on the impact of ITD on clock distribution networks (CDN), whose function is vital to guarantee the synchronization among physically spaced sequential components of digital circuits. Using our simulation framework, we characterized the thermal behavior of a clock tree mapped onto an industrial 65nm CMOS technology and obtained using a standard synthesis tool. Results demonstrate the presence of ITD at low operating voltages and open new potential research scenarios into the EDA field. Alessandro Sassone, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino, Richard Goldman, Vazgen Melikyan, Eduard Babayan, Salvatore Rinaudo |
DATE | 3 |
| 2011 | Moving to Green ICT: From stand-alone power-aware IC design to an integrated approach to energy efficient design for heterogeneous electronic systemsabstractEnergy efficiency is one of the most critical aspects of todays information society. The most obvious benefits of being Green are reduced environmental impact and cost savings. Reducing energy consumption of electronic devices, circuits and heterogeneous systems, however, is not trivial. This requires the development of innovative energy-aware vertical design solutions and EDA technologies for next generations' nanoelectronics circuits and systems, and the related energy generation, conversion and management systems. Salvatore Rinaudo, Giuliana Gangemi, Andrea Calimera, Alberto Macii, Massimo Poncino |
DATE | 4 |
| 2011 | Fast Computation of Discharge Current Upper Bounds for Clustered Power GatingabstractThe capability of accurately estimating an upper bound of the maximum current drawn by a digital macroblock from the ground or power supply line constitutes a major asset of automatic power-gating flows. In fact, the maximum current information is essential to properly size the sleep transistor in such a way that speed degradation and signal integrity violations are avoided. Loose upper bounds can be determined with a reasonable computational cost, but they lead to oversized sleep transistors. On the other hand, exact computation of the maximum drawn current is an NP-hard problem, even when conservative simplifying assumptions are made on gate-level current profiles. In this paper, we present a scalable algorithm for tightening upper bound computation, with a controlled and tunable computational cost. The algorithm exploits state-of-the-art commercial timing analysis engines, and it is tightly integrated into an industrial power-gating flow for leakage power reduction. The results we have obtained on large circuits demonstrate the scalability and effectiveness of our estimation approach. Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Row-Based Power-Gating: A Novel Sleep Transistor Insertion Methodology for Leakage Power Optimization in Nanometer CMOS CircuitsabstractLeakage power has become a serious concern in nanometer CMOS technologies, and power-gating has shown to offer a viable solution to the problem with a small penalty in performance. This paper focuses on leakage power reduction through automatic insertion of sleep transistors for power-gating. In particular, we propose a novel, layout-aware methodology that facilitates sleep transistor insertion and virtual-ground routing on row-based layouts. We also introduce a clustering algorithm that is able to handle simultaneously timing and area constraints, and we extend it to the case of multi-Vtsleep transistors to increase leakage savings. The results we have obtained on a set of benchmark circuits show that the leakage savings we can achieve are, by far, superior to those obtained using existing power-gating solutions and with much tighter timing and area constraints. Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | An integrated thermal estimation framework for industrial embedded platformsabstractNext generation industrial embedded platforms require the development of complex power and thermal management solutions. Indeed, an increasingly fine and intrusive thermal control is required because of temperature impact on leakage and reliability. To be effective, the implementation of these policies involves decisions that must be taken during various phases along the design process, to enable the development of architectural level countermeasures and the required hardware knobs, such as power modes, power supply regulation granularity and the number of on-chip temperature sensors. As a consequence, a framework allowing thermal estimation exploiting design-time information is desirable. Andrea Acquaviva, Andrea Calimera, Alberto Macii, Massimo Poncino, Enrico Macii, Matteo Giaconia, Claudio Parrella |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | Power-aware partitioning of data convertersabstractSerial data streaming, one of the most important functions in modern communication systems, is becoming more and more power consuming as bit-rate is increasing without standstill. In this work, we propose a novel technique for partitioning conventional N-bit registers in standard data converters, in order to reduce their switching activity, and therefore power consumption. The architecture here presented have a very low area overhead with respect to the standard ones for serializers and, furthermore, it allows different (i.e., custom) configurations for the partitioning. The proposed method even allows to extract idleness conditions of register banks in order to apply the well-known clock-gating technique to the circuit and thus furtherly reducing the total power consumption. This method has been applied to different data converters (i.e., serializers) in a base-band radio within an ultra low-power industrial design and the results highlight the effectiveness of the proposed technique. Alberto Bonanno, Alberto Bocca, Alberto Macii, Enrico Macii |
VLSI-SoC | 3 |
| 2009 | Enabling concurrent clock and power gating in an industrial design flowabstractClock-gating and power-gating have proven to be very effective solutions for reducing dynamic and static power, respectively. The two techniques may be coupled in such a way that the clock-gating information can be used to drive the control signal of the power-gating circuitry, thus providing additional leakage minimization conditions w.r.t. those manually inserted by the designer. This conceptual integration, however, poses several challenges when moved to industrial design flows. Although both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. This paper presents a layout-oriented synthesis flow which integrates the two techniques and that relies on leading-edge, commercial EDA tools. Starting from a gated-clock netlist, we partition the circuit in a number of clusters that are implicitly determined by the groups of cells that are clock-gated by the same register. Using a row-based granularity, we achieve runtime leakage reduction by inserting dedicated sleep transistors for each cluster. The entire flow has been benchmarked on a industrial design mapped onto a commercial, 65 nm CMOS technology library. Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 3 |
| 2009 | Placement-aware Clustering for Integrated Clock and Power GatingabstractClock-gating and power-gating are the most widely used solutions for reducing dynamic and static power. They can be potentially integrated so that clock-gating conditions can be used to control the power-gating circuitry thus also reducing static power. This integration becomes however difficult when applied in an industrial design flow. Even if both clock and power-gating are supported by most commercial synthesis tools, their combined implementation requires some flexibility in the back-end tools that is not currently available. Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 3 |
| 2008 | A Scalable Algorithmic Framework for Row-Based Power-GatingabstractLeakage power is a serious concern in nanometer CMOS technologies. In this paper we focus on leakage reduction through automatic insertion of sleep transistors for power gating in standard cell based designs. In particular, we propose clustering algorithms for row- based power-gating methodology which is based on using rows of the layout as the granularity for clustering. Our clustering methodology does timing and area constraint driven power-gating in contrast to only timing driven power-gating as proposed in the previous works. We present two distinct clustering algorithms with different accuracy-efficiency trade-off. An optimal one, which exploits a 0-1 or binary integer programming approach, and a heuristic one, which resorts to an implicit enumeration of the layout rows. Results show that, for all the benchmarks, the leakage power savings, as compared to previous techniques, are more than 75% when we have the same timing constraints but half sleep transistor area and at least 60% when area constraint is set at one fourth. We also show that we can perform clustering with no speed degradation and achieve maximum leakage power savings up-to 83%. Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2008 | Integrating Clock Gating and Power Gating for Combined Dynamic and Leakage Power Optimization in Digital CMOS CircuitsabstractClock Gating and Power Gating are two of the most effective techniques that are applied today for reducing dynamic and leakage power, respectively, in digital CMOS circuits. The combined use of the two solutions, however, poses some challenges in terms of practical integration of the required control logic and the power/timing overhead associated to it. This paper presents an analysis methodology and a prototype CAD tool that support the designer in understanding when the joint application of Clock Gating and Power Gating may result in significant power savings. Enrico Macii, Letícia Maria Veiras Bolzani, Andrea Calimera, Alberto Macii, Massimo Poncino |
DSD | 4 |
| 2008 | Optimal sleep transistor synthesis under timing and area constraintsabstractLeakage power reduction in nano-CMOS designs has gained tremendous interest both in academia and industry. Many techniques have been proposed in the literature for leakage power reduction and one of the prominent techniques for leakage power reduction is the use of sleep transistors as power-gating elements to cut-off sub-threshold leakage current in circuits when they are in stand-by mode. Although sleep transistor insertion is very effective in cutting-off leakage, it also incurs timing, area and routing overhead. Since most of the sleep transistor insertion methodologies do post layout insertion, care should be taken such that there is minimal perturbation of the original layout. Over design of sleep transistors cells and sub-optimal sleep transistor placement must be avoided to achieve final design closure. Since the sleep transistor area plays an important and prominent role in this aspect, it necessitates for optimal sleep transistor sizing and synthesis technique under area constraints. In this paper, we first provide a methodology for optimal sleep transistor synthesis under given area constraints. We then apply our technique to the general timing and area constraint driven row-based power-gating methodology proposed in [13] and show how optimal low leakage designs with constraints on timing and area can be designed Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 4 |
| 2008 | On quantifying the figures of merit of power-gating for leakage power minimization in nanometer CMOS circuitsabstractPower-gating has proved to be one of the most effective solutions for reducing stand-by leakage power in nanometer-scale CMOS circuits, and different strategies and algorithms for its application have been proposed recently. Unfortunately, power- gating comes with its own set of costs: Performance degradation, area increase, dynamic power increase and routing congestion. When a decision to power-gate a design has to be taken, pros and cons of power-gating have to be properly weighted to achieve optimal results. In this paper, we define "Figures of Merit" (FoMs) for power-gating, which can be used by designers to better understand the benefits and costs of power-gating, thereby allowing them to achieve optimal results. We then quantify the FoMs by applying a state-of-the-art, industry-strength power- gating flow on a set of designs implemented onto an industrial 65 nm CMOS process, and provide insightful discussion on how optimum power-gating can be achieved. Ashoka Visweswara Sathanur, Andrea Calimera, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 5 |
| 2008 | Multiple power-gating domain (multi-VGND) architecture for improved leakage power reductionabstractRow-based power-gating has recently emerged as a meet-in-the-middle sleep transistor insertion paradigm between cell-level and block-level granularity, in which each layout row defines the unit of gating, and different rows can be clustered and share the same sleep transistor. Previous works, however, assume the availability of a single virtual ground voltage, thus making the decision of whether to gate or not a given cluster a binary choice: a cluster is either gated or not. Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2008 | Implementation of a thermal management unit for canceling temperature-dependent clock skew variations
Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Alberto Macii, Enrico Macii, Massimo Poncino |
Integr. | 5 |
| 2008 | Dynamic Thermal Clock Skew Compensation Using Tunable Delay BuffersabstractThe thermal gradients existing in high-performance circuits may significantly affect their timing behavior, in particular, by increasing the skew of the clock net and/or altering hold/setup constraints, possibly causing the circuit to operate incorrectly. The knowledge of the spatial distribution of temperature can be used to properly design a clock network that is able to compensate such thermal non-uniformities. However, redesign of the clock network is effective only if temperature distribution is stationary, i.e., does not change over time. In this paper, we specifically address the problem of dynamically modifying the clock tree in such a way that it can compensate for temporal variations of temperature. This is achieved by exploiting the buffers that are inserted during the clock network generation, by transforming them into tunable delay elements. Temperature-induced delay variations are then compensated by applying the proper tuning to the tunable buffers, which is computed offline and stored in a tuning table inserted in the design. We propose an algorithm to minimize the number of inserted tunable buffers, as well as their tunable range (which directly relates to complexity). Results show that clock skew is kept within original bounds with worst-case power and area penalty of 3.5% and 5.5% respectively. Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2007 | Interactive presentation: Efficient computation of discharge current upper bounds for clustered sleep transistor sizing
Ashoka Visweswara Sathanur, Andrea Calimera, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2007 | Design of a family of sleep transistor cells for a clustered power-gating flow in 65nm technologyabstractClustered sleep transistor insertion is an effective leakage power reduction technique that is well-suited for integration in an automated design flow and offers a flexible tradeoff between area, delay overhead and turn-on transition time. In this work, we focus on the design of a family of sleep transistor cells, fully compatible with the physical design rules of a commercial 65nm CMOS library. We describe circuit-level and layout optimizations, as well as the cell characterization procedure required to support automated sleep transistor cell selection and instantiation in a clustered power-gating insertion flow. Andrea Calimera, Antonio Pullini, Ashoka Visweswara Sathanur, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 5 |
| 2007 | Design Exploration of a Thermal Management Unit for Dynamic Control of Temperature-Induced Clock SkewabstractPower densities and temperatures in today's high performance circuits have reached alarmingly high levels due to increased scaling in feature sizes. Subsequently, the various techniques used to keep them under control have also created "zones" of varying temperatures, thus contributing to temperature gradients inside the chip. These gradients have detrimental effects on the delay of wires, as resistance in metals increases with temperature. Clock nets are extremely susceptible to this effect, since they run through the entire chip. Different techniques have been proposed to counter the impact of temperature on clock speed; they range from re-designing the clock network assuming a stationary profile to more adaptive solutions that allow to dynamically compensate the clock skew through replacement of the original buffers with a specially designed counterpart, called tunable delay buffers (TDBs). Dynamic skew management based on TDBs calls for the presence on the chip of a thermal management unit (TMU), whose purpose is that of periodically choosing the actual delay that each TDB must provide in order to achieve skew optimization. Preliminary implementations of such a unit for basic assumptions on the distribution of sensors and their accuracy have indicated negligible impact on the original design. This work aims at exploring in detail several issues related to TMU design, pivoting on the fact that sensor distribution and its accuracy could in fact impact the design in a significant way depending on the design. We provide the results of a careful exploration we have performed on a meaningful case study, quantifying values for area and power consumption. Karthik Duraisami, Prassanna Sithambaram, Ashoka Visweswara Sathanur, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 4 |
| 2007 | Timing-driven row-based power gatingabstractIn this paper we focus on leakage reduction through automatic insertion of sleep transistors using a row-based granularity. In particular, we tackle here the two main issues involved in this methodology: (i) Clustering and (ii) the interfacing of power-gated and non power-gated regions within the same block. The clustering algorithm automatically selects an optimal subset of rows that can be power-gated with a tightly controlled delay overhead. We then address the issue of interfacing different gated regions and propose a novel technique to address this issue with minimal area and power penalty. Our approach is compatible with state-of-the art logic and physical synthesis flows and it does not significantly impact design closure. We achieve leakage power reductions as high as 89% for a set of standard benchmarks, with minimum timing and area overhead. Ashoka Visweswara Sathanur, Antonio Pullini, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 4 |
| 2006 | Enabling fine-grain leakage management by voltage anchor insertionabstractFunctional unit shutdown based on MTCMOS devices is effective for leakage reduction in aggressively scaled technologies. However, the applicability of MTCMOS-based shutdown in a synthesis-based design flow poses the challenge of interfacing logic blocks in shutdown mode with active units: The outputs of inactive gates can float at intermediate voltages, causing very large short-circuit currents in the active gates they drive. In this paper, we propose two novel low-overhead elementary cells that fully address this issue. These cells can be added to any synthesis library, and they can be inserted into a netlist at the boundary between shutdown and active regions. Our results show that: (i) our cells solve the interfacing problem with minimum overhead; and (ii) a non-intrusive design flow enhancement is sufficient to automatically insert interface cells in post-synthesis netlists Pietro Babighian, Luca Benini, Alberto Macii, Enrico Macii |
DATE | 3 |
| 2006 | Thermal resilient bounded-skew clock tree optimization methodologyabstractThe existence of non-uniform thermal gradients on the substrate in high performance IC's can significantly impact the performance of global on-chip interconnects. This issue is further exacerbated by the aggressive scaling and other factors such as dynamic power management schemes and non-uniform gate level switching activity. In high-performance systems, one of the most important problems is clock skew minimization since it has a direct impact on the maximum operating frequency of the system. Since clocks are routed across the entire chip, the presence of thermal gradients can significantly alter their characteristics because wire resistance increases linearly as the temperature increases. This often results in failure to meet original timing constraints thereby rendering the original topology unusable. Therefore it is necessary to perform a temperature aware re-embedding of the original topology to meet timing under these temperature effects. This work primarily explores these issues by proposing two algorithms that re-structure an existing clock tree topology to compensate for such temperature effects and as a result also meet timing constraints Ashutosh Chakraborty, Prassanna Sithambaram, Karthik Duraisami, Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2006 | Implications of ultra low-voltage devices on design techniques for controlling leakage in NanoCMOS circuitsabstractEnabled by technology scaling, ultra low-voltage devices have now found wide application in modern VLSI circuits. While low-voltage implies reduced dynamic power, it also signifies increased leakage power, as lower supply voltages are usually paired with lower threshold voltages in order to preserve circuit speed. This originates an increase in sub-threshold leakage currents that constitute, today, one of the most serious bottlenecks to further technology and supply voltage scaling. The need of controlling leakage power in nanometric devices is imposing a significant shift in the way integrated circuits are designed and manufactured. The behavior of devices with nanometric feature sizes is much more sensitive to parameters such as the operating temperature of the circuit, which in the past were neglected. In this paper we quantitatively analyze the leakage control capabilities of some well-established circuit-level design techniques, and assess how the effectiveness of such techniques scales with respect to decreased supply voltages (as induced by technology scaling) and temperature variations, thus providing an interesting insight on how leakage control solutions that are in use today is applicable in future designs Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Alberto Macii, Enrico Macii, Massimo Poncino |
ISCAS | 5 |
| 2006 | Dynamic thermal clock skew compensation using tunable delay buffersabstractThe thermal gradients existing in high-performance circuits may significantly affect their timing behavior, in particular by increasing the skew of the clock net and/or altering hold/setup constraints, possibly causing the circuit to operate incorrectly. The knowledge of the spatial distribution of temperature can be used to properly design a clock network that is able to compensate such thermal non-uniformities. However, re-design of the clock network is effective only if temperature distribution is stationary, i.e., does not change over time. In this work, we specifically address the problem of dynamically modifying the clock tree in such a way that it can compensate for temporal variations of temperature. This is achieved by exploiting the buffers that are inserted during the clock network generation, by transforming them into tunable delay elements. Temperature-induced delay variations are then compensated by applying the proper tuning to the tunable buffers, which is computed off-line and stored in a tuning table inserted in the design. We propose an algorithm to minimize the number of inserted tunable buffers, as well as their tunable range (which directly relates to complexity). Results show that clock skew is kept within original bounds with minimum area and power penalty. The maximum increase in power is 23.2% with most benchmarks exhibiting less than 5% increase in power. Ashutosh Chakraborty, Karthik Duraisami, Ashoka Visweswara Sathanur, Prassanna Sithambaram, Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 6 |
| 2005 | Low-overhead state-retaining elements for low-leakage MTCMOS designabstractMulti-threshold CMOS (MTCMOS) has shown to be a very effective technique for reducing sub-threshold leakage currents in DSM CMOS designs. Application of the MTC-MOS paradigm to sequential circuits requires the availability of data-retaining elements for storing circuit state during stand-by mode. In this paper we propose two novel circuit schemes for sequential elements featuring low leakage currents in stand-by mode and high-speed/low-dynamic power in active mode. We present post-layout simulation results obtained after parasitic extraction for delay and power of circuits built in 130nm CMOS technology. Our experiments demonstrate several advantages of the proposed schemes over the best previously published solutions. Pietro Babighian, Luca Benini, Alberto Macii, Enrico Macii |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Exploring the impact of architectural parameters on energy efficiency of application-specific block-enabled SRAMsabstractApplication-Specific Block-Enabled (ASBE) SRAMs represent a viable solution for reducing energy consumption in embedded memories. The basic idea behind ASBE architectures is that of partitioning the memory array into a number of non-uniformly sized blocks, such that memory access cost is reduced. The number and sizes of the partitions yielding a minimum power implementation of the SRAM macro is determined by the partitioning algorithm based on the memory access profile obtained as a result of the application (or application mix) executed by the processor. Given the complexity of the design space we are dealing with, there are several degrees of freedom that the partitioning engine may exploit to come up with the most energy-efficient memory architecture. In this paper, we investigate how the quality of the partitioned memory depends on the architectural parameters that define the memory structure (e.g., min and max number of lines per partition, min and max number of words per line, granularity of the partitions); such parameters, in turn, are constrained by the technology and process of choice. We believe that the results presented in this work will provide very useful guidelines for a succesfull adoption of the ASBE approach in practice, as this design paradigm is gaining a lot of attention for the new generations of embedded systems. Prassanna Sithambaram, Alberto Macii, Enrico Macii |
ACM Great Lakes Symposium on VLSI | 2 |
| 2004 | Block-Enabled Memory Macros: Design Space Exploration and Application-Specific TuningabstractIn this paper, we propose a combined solution that allows us to customize the architecture of internally partitioned SRAM macros according to the given application be executed. Energy savings with respect to monolithic memory configurations are above 40%, without access time violation. Luca Benini, Alessandro Ivaldi, Alberto Macii, Enrico Macii |
DATE | 3 |
| 2004 | Post-layout leakage power minimization based on distributed sleep transistor insertionabstractThis paper introduces a new approach to sub-threshold leakage power reduction in CMOS circuits. Our technique is based on automatic insertion of sleep transistors for cutting sub-threshold current when CMOS gates are in stand-by mode. Area and speed overhead caused by sleep transistor insertion are tightly controlled thanks to: (i) a post-layout incremental modification step that inserts sleep transistors in an existing row-based layout; (ii) an innovative algorithm that selects the subset of cells that can be gated for maximal leakage power reduction, while meeting user-provided constraints on area and delay increase. The presented technique is highly effective and fully compatible with industrial back-end flows, as demonstrated by post-layout analysis on several benchmarks placed and routed with state-of-the art commercial tools for physical design. Pietro Babighian, Luca Benini, Alberto Macii, Enrico Macii |
ISLPED | 3 |
| 2004 | Memory energy minimization by data compression: algorithms, architectures and implementationabstractStoring data in compressed form is becoming common practice in high-performance systems, where memory bandwidth constitutes a serious bottleneck to program execution speed. In this paper, we suggest hardware-assisted data compression as a tool for reducing energy consumption of processor-based systems. We propose a novel and efficient architecture for on-the-fly data compression and decompression whose field of operation is the cache-to-memory path. Uncompressed cache lines are compressed before they are written back to main memory, and decompressed when cache refills take place. We explore two classes of table-based compression schemes. The first, based on offline data profiling, is particularly suitable to embedded systems, where predictability of the data set is usually higher than in general-purpose systems. The second solution we introduce is adaptive, that is, it takes decisions on whether data words should be compressed according to the data statistics of the program being executed. We describe in details the architecture of the compression/decompression unit and we provide an insight about its implementation as a hardware (HW) block. We present experimental results concerning memory traffic and energy consumption in the cache-to-memory path of a core-based system running standard benchmark programs. The obtained energy savings range from 8%-39% when profile-driven compression is adopted, and from 7%-26% when the adaptive scheme is used. Performance improvements are also achieved as a by-product, showing the practical applicability of the proposed approach. Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Energy-aware design techniques for differential power analysis protectionabstractDifferential power analysis is a very effective cryptanalysis technique that extracts information on secret keys by monitoring instantaneous power consumption of cryptoprocessors. To protect against differential power analysis, power supply noise is added in cryptographic computations, at the price of an increase in power consumption. We present a novel technique, based on well-known power-reducing transformations coupled with randomized clock gating, that introduces a significant amount of scrambling in the power profile without increasing (and, in some cases, by even reducing) circuit power consumption. Luca Benini, Alberto Macii, Enrico Macii, Elvira Omerbegovic, Fabrizio Pro, Massimo Poncino |
DAC | 2 |
| 2003 | A New Algorithm for Energy-Driven Data Compression in VLIW Embedded Processors
Alberto Macii, Enrico Macii, Fabrizio Crudo, Roberto Zafalon |
DATE | 1 |
| 2003 | Improving the Efficiency of Memory Partitioning by Address Clustering
Alberto Macii, Enrico Macii, Massimo Poncino |
DATE | 1 |
| 2003 | A novel architecture for power maskable arithmetic unitsabstractPower maskable units have been proposed as a viable solution for preventing side-channel attacks to cryptoprocessors. This paper presents a novel architecture for the implementation of a class of such kinds of units, namely arithmetic components, which find wide usage in cryptographic applications and which are not suitable to traditional masking techniques. Results of extensive exploration and architectural trade-off analysis show the viability of the proposed solution. Luca Benini, Alberto Macii, Enrico Macii, Elvira Omerbegovic, Massimo Poncino, Fabrizio Pro |
ACM Great Lakes Symposium on VLSI | 2 |
| 2003 | Energy-efficient data scrambling on memory-processor interfacesabstractCrypto-processors are prone to security attacks based on the observation of their power consumption profile. We propose new techniques for increasing the non-determinism of such profile, which rely on the idea of introducing randomness in the bus data transfers. This is achieved by combining data scrambling with energy-efficient bus encoding, thus providing high information protection at no energy cost.Results on a set of bus traces originated by real-life applications demonstrate the applicability of the proposed solution. Luca Benini, Angelo Galati, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 3 |
| 2003 | Discharge Current Steering for Battery Lifetime OptimizationabstractPortable and wearable computers can be powered by different combinations of two or more battery packs to give the user the possibility of choosing an optimal compromise between lifetime and weight/size. Recent work on battery-driven power management has demonstrated that sequential discharge is suboptimal in multibattery systems and lifetime can be maximized by distributing (steering) the current load on the available batteries, thereby discharging them in a partially concurrent fashion. Based on these observations, we formulate multibattery lifetime maximization as a continuous, constrained optimization problem, which can be efficiently solved by nonlinear optimizers. We show that significant lifetime extensions can be obtained with respect to standard sequential discharge (up to 160 percent), as well to previously proposed battery scheduling algorithms (up to 12 percent). Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Computers | 3 |
| 2003 | Energy-aware design of embedded memories: A survey of technologies, architectures, and optimization techniquesabstractEmbedded systems are often designed under stringent energy consumption budgets, to limit heat generation and battery size. Since memory systems consume a significant amount of energy to store and to forward data, it is then imperative to balance power consumption and performance in memory system design. Contemporary system design focuses on the trade-off between performance and energy consumption in processing and storage units, as well as in their interconnections. Although memory design is as important as processor design in achieving the desired design objectives, the former topic has received less attention than the latter in the literature. This article centers on one of the most outstanding problems in chip design for embedded applications. It guides the reader through different memory technologies and architectures, and it reviews the most successful strategies for optimizing them in the power/performance plane. Luca Benini, Alberto Macii, Massimo Poncino |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2003 | Scheduling battery usage in mobile systemsabstractThe use of multibattery power supplies is becoming common practice in electronic appliances of the latest generations. Economical and manufacturing constraints are at the basis of this choice. Unfortunately, a partitioned battery subsystem is not able to deliver the same amount of charge as a monolithic battery with the same total capacity. In this paper, we define the concept of battery scheduling, we investigate several policies for solving the problem of optimal charge delivery, and we study the relationship of such policies with different configurations of the battery subsystem. Experimental results, obtained for different kinds of current workloads, demonstrate that the choice of the proper scheduling can make system lifetime as close as 1% of the theoretical upper bound, that is, a monolithic power supply of equal capacity. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Hardware-Assisted Data Compression for Energy Minimization in Systems with Embedded ProcessorsabstractIn this paper, we suggest hardware-assisted data compression as a tool for reducing energy consumption of core-based embedded systems. We propose a novel and efficient architecture for on-the-fly data compression and decompression whose field of operation is the cache-to-memory path. Uncompressed cache lines are compressed before they are written back to main memory, and decompressed when cache refills take place. We explore two classes of compression methods, profile-driven and differential, since they are characterized by compact HW implementations, and we compare their performance to those provided by some state-of-the-art compression methods (e.g., we have considered a few variants of the Lempel-Ziv encoder). We present experimental results about memory traffic and energy consumption in the cache-to-memory path of a core-based system running standard benchmark programs. The achieved average energy savings range from 4.2% to 35.2%, depending on the selected compression algorithm. Luca Benini, Davide Bruni, Alberto Macii, Enrico Macii |
DATE | 3 |
| 2002 | Enhanced clustered voltage scaling for low powerabstractThis paper presents a voltage scaling approach that is based on an enhanced variant of clustered voltage scaling originally proposed by Usami and Horowitz ([1]) The results show that subtituting the original depth first strategy with a breadth first one results in improved speed and quality of results. Data are validated through power and timing analysis performed with a commercial tool. Monica Donno, Luca Macchiarulo, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2002 | Discharge current steering for battery lifetime optimizationabstractRecent work on battery-driven power management has demonstrated that sequential discharge is suboptimal in multi-battery systems, and lifetime can be maximized by distributing (steering) the current load on the available batteries, thereby discharging them in a partially concurrent fashion. Based on these observations, we formulate multi-battery lifetime maximization as a continuous, constrained optimization problem, which can be efficiently solved by non-linear optimizers. We show that great lifetime extensions can be obtained with respect to standard sequential discharge, as well to previously proposed battery allocation schemes. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 2 |
| 2002 | Layout-driven memory synthesis for embedded systems-on-chipabstractMemory-processor integration offers new opportunities for reducing, the energy of a system. In the case of embedded systems, where memory access patterns can typically be profiled at design time, one solution consists of mapping the most frequently accessed addresses onto the on-chip SRAM to guarantee power and performance efficiency. In this work, we propose an algorithm for the automatic partitioning of on-chip SRAMs into multiple banks. Starting from the dynamic execution profile of an embedded application running on a given processor core, we synthesize a multi-banked SRAM architecture optimally fitted to the execution profile. The algorithm computes an optimal solution to the problem under realistic assumptions on the power cost metrics, and with constraints on the number of memory banks. The partitioning algorithm is integrated with the physical design phase into a complete flow that allows the back annotation of layout information to drive the partitioning process. Results, collected on a set of embedded applications for the ARM processor, have shown average energy savings around 34%. Luca Benini, Luca Macchiarulo, Alberto Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2002 | Minimizing memory access energy in embedded systems by selective instruction compressionabstractWe propose a technique for reducing the energy spent in the memory-processor interface of an embedded system during the execution of firmware code. The method is based on the idea of compressing the most commonly executed instructions so as to reduce the energy dissipated during memory access. Instruction decompression is performed on-the-fly by a hardware block located between processor and memory: No changes to the processor architecture are required. Hence, our technique is well suited for systems employing IP cores whose internal architecture cannot be modified. We describe a number of decompression schemes and architectures that effectively trade off hardware complexity and static code size increase for memory energy and bandwidth reduction, as proved by the experimental data we have collected by executing several test programs on different design templates. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | From Architecture to Layout: Partitioned Memory Synthesis for Embedded Systems-on-ChipabstractWe propose an integrated front-end/back-end flow for the automatic generation of a multi-bank memory architecture for embedded systems. The flow is based on an algorithm for the automatic partitioning of on-chip SRAM. Starting from the dynamic execution profile of an embedded application running on a given processor core, we synthesize a multi-banked SRAM architecture optimally fitted to the execution profile. Luca Benini, Luca Macchiarulo, Alberto Macii, Enrico Macii, Massimo Poncino |
DAC | 3 |
| 2001 | Extending lifetime of portable systems by battery schedulingabstractMulti-battery power supplies are becoming popular in electronic appliances of the latest generations, due to economical and manufacturing constraints. Unfortunately, a partitioned battery subsystem is not able to deliver the same amount of charge as a monolithic battery with the same total capacity. In this paper, we define the concept of battery scheduling, we investigate policies for solving the problem of optimal charge delivery, and we study the relationship of such policies with different configurations of the battery subsystem. Results, obtained for different workloads, demonstrate that the choice of the proper scheduling can make, in the best cease, system lifetime as close as 1% of that guaranteed by a monolithic battery of equal capacity. Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DATE | 3 |
| 2001 | Cached-code compression for energy minimization in embedded processorsabstractThis paper contributes a novel approach for reducing static code size and instruction fetch energy for cache-based core processors running embedded applications. Our irnplemen-tation of the decompression unit guarantees fast and low-energy, on-the-fly instruction decompression at each cache lookup. The decompressor is placed outside the core bound-aries; therefore, processor architecture does not need any modification, making the proposed compression approach suitable to IP-based designs. Viability of our solution is assessed through extensive benchmarking performed on a number of typical embedded programs. 1. Luca Benini, Alberto Macii, Alberto Nannarelli |
ISLPED | 2 |
| 2001 | Discrete-time battery models for system-level low-power designabstractFor portable applications, long battery lifetime is the ultimate design goal. Therefore, the availability of battery and voltage converter models providing accurate estimates of battery lifetime is key for system-level low-power design frameworks. In this paper, we introduce a discrete-time model for the complete power supply subsystem that closely approximates the behavior of its circuit-level continuous-time counterpart. The model is abstract and efficient enough to enable event-driven simulation of digital systems described at a very high level of abstraction and that includes, among their components, also the power supply. The model gives the designer the possibility of estimating battery lifetime during system-level design exploration, as shown by the results we have collected on meaningful case studies. In addition, it is flexible and it can thus be employed for different battery chemistries. Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | Stream synthesis for efficient power simulation based on spectral transformsabstractOne way of minimizing the time required to perform simulation-based power estimation is that of reducing the length of the input trace to be fed to the simulator. Obviously, the use of a reduced stream may introduce some errors in the estimation results. The generation (or synthesis) of the short input sequence to be used for power simulation should then be carried out in such a way that the resulting error is minimized. Existing techniques exploit the knowledge of some statistical and correlation characteristics concerning the original input trace to generate a reduced stream that closely matches such characteristics. In this paper, we introduce a new stream synthesis method. Its distinguishing feature is the use of spectral analysis based on the discrete Fourier transform to determine a reduced sequence of vectors that enables us to shorten the overall power simulation time at a very limited penalty in accuracy. The effectiveness and the robustness, in terms of estimation accuracy, of the proposed synthesis procedure are demonstrated by the experimental results we have obtained on standard combinational benchmarks for a variety of input streams with different statistical and correlation properties. Data for sequential circuits are also reported. Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Synthesis of application-specific memories for power optimization in embedded systemsabstractThis paper presents a novel approach to memory power optimization for embedded systems based on the exploitation of data locality. Locations with highest access frequency are mapped onto a small, low-power application-specific memory which is placed close the processor. Although, in principle, a cache may be used to implement such a memory, more efficient solutions may be adopted. We propose an architecture that outperforms (power-wise) different types of cache memories at no penalty in performance. Power savings (averaged over a number of embedded applications running on ARM processors) range from 12% to 68%. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
DAC | 2 |
| 2000 | A Discrete-Time Battery Model for High-Level Power EstimationabstractIn this paper, we introduce a discrete-time model for the complete power supply sub-system that closely approximates the behavior of its circuit-level (i.e., HSpice), continuous-time counterpart. The model is abstract and efficient enough to enable event-driven simulation of digital systems described at a very high level of abstraction and that include, among their components, also the power supply. Therefore, it can be successfully used for the purpose of battery life-time estimation during design optimization, as shown by the results we have collected on a meaningful case study. Experiments prove also that the accuracy of our model is very close to that provided by the corresponding Spice-level model. Luca Benini, Giuliano Castelli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DATE | 3 |
| 2000 | Supporting system-level power exploration for DSP applicationsabstractSystem-level power exploration requires tools for estimation of the overall power consumed by a system, as well as a detailed breakdown of the consumption of its main functional blocks. We focus on power estimation for data-dominated systems specified as synchronous data-flows and implemented on a single-processor architecture. Our estimator is integrated within the Ptolemy design environment, and provides information to system designers on the power dissipated by every task in a given specification. Power estimation is based on instruction-level power models. We demonstrate the applicability of our tool on a few design examples and target architectures. Luca Benini, Marco Ferrero, Alberto Macii, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 3 |
| 2000 | A recursive algorithm for low-power memory partitioningabstractMemory-processor integration offers new opportunities for reducing the energy of a system. In the case of embedded systems, one solution consists of mapping the most frequently accessed addresses onto the on-chip SRAM to guarantee power and performance efficiency. This option is especially effective when memory access patterns can be profiled and studied at design time (as in typical real-time embedded systems). Luca Benini, Alberto Macii, Massimo Poncino |
ISLPED | 2 |
| 2000 | Architectures and synthesis algorithms for power-efficient businterfacesabstractIn this paper we present algorithms for the synthesis of encoding and decoding interface logic that minimizes the average number of transitions on heavily-loaded global bus lines at no cost in communication throughput (i.e., one word is transmitted at each cycle). The distinguishing feature of our approach is that it does not rely on designer's intuition, but it automatically constructs low-transition activity codes and hardware implementation of encoders and decoders, given information on word-level statistics. We propose an accurate method that is applicable to low-width buses, as well as approximate methods that scale well with bus width. Furthermore, we introduce an adaptive architecture that automatically adjusts encoding to reduce transition activity on buses whose word-level statistics are not known a priori. Experimental results demonstrate that our approaches out-perform specialized low-power encoding schemes presented in the past. Luca Benini, Alberto Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Glitch power minimization by selective gate freezingabstractThis paper presents a technique for glitch power minimization in combinational circuits. The total number of glitches is reduced by replacing some existing gates with functionally equivalent ones (called F-Gates) that can be "frozen" by asserting a control signal. A frozen gate cannot propagate glitches to its output. Algorithms for gate selection and clustering that maximize the percentage of filtered glitches and reduce the overhead for generating the control signals are introduced. A power-efficient CMOS implementation of F-Gates is also described. An important feature of the proposed method is that it can be applied in place directly to layout-level descriptions; therefore, it guarantees very predictable results and minimizes the impact of the transformation on circuit size and speed. Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | Synthesis of Low-Overhead Interfaces for Power-Efficient Communication over Wide BusesabstractIn this paper we present algorithms for the synthesis of encoding and decoding interface l o gic that minimizes the average number of transitions on heavily-loaded global bus lines.The approach automatically constructs low-transition activity codes and hardware implementation of encoders and decoders, given information on word-level statistics.We present an accurate method that is applicable to low-width buses, as well as approximate methods that scale well with bus width.Furthermore, we introduce an adaptive architecture that automatically adjusts encoding to reduce t r ansition activity on buses whose word-level statistics are not known a-priori.Experimental results demonstrate that our approach well outperforms low-power encoding schemes presented in the past. Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DAC | 2 |
| 1999 | Glitch Power Minimization by Gate FreezingabstractThis paper presents a technique for glitch power minimization in combinational circuits. The total number of glitches is reduced by replacing some existing gates with functionally equivalent ones (called F-gates) that can be "frozen" by asserting a control signal. A frozen gate cannot propagate glitches to its output. An important feature of the proposed method is that it can be applied in-place directly to layout-level descriptions; therefore, it guarantees very predictable results and minimizes the impact of the transformation on circuit size and speed. Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
DATE | 3 |
| 1999 | Regression-Based Macromodeling for Delay Estimation of Behavioral ComponentsabstractThis paper presents a methodology for delay estimation of hardware components described at the behavioral-level. The basis of the proposed technique is a well-known theoretical result that relates the entropy of a logic function to the delay of a multi-level implementation of the same function. We propose an improved model for delay estimation, and we prove its validity by means of experiments performed on a set of standard benchmarks. Alberto Macii, Enrico Macii, Giuseppe Odasso, Massimo Poncino, Riccardo Scarsi |
Great Lakes Symposium on VLSI | 1 |
| 1999 | Selective instruction compression for memory energy reduction in embedded systemsabstractWe propose a technique for reducing the energy required by firmware code to ezecute on embedded systema. The method ia based on the idea of compressing the moat commonly ezecuted instructions 80 a ~ to reduce the energy dissipated in memory acceeees. Instruction decompression is performed on the fly by a hardware module located between processor and memory: No changea to the processor architecture ore required. Hence, our technique is well-suited for systema employing IP cone whose internal architecture cannot be modified. We describe a number of decompnaaion achemea and architec-tura that effectively trade off hardware complezity for memory energy and bandwidth nduction, aa proved by experimental data collected by executing aeveml sample programs. 1 Luca Benini, Alberto Macii, Enrico Macii, Massimo Poncino |
ISLPED | 2 |
| 1998 | Reducing Power Consumption of Dedicated Processors Through Instruction Set EncodingabstractWith the increased clock frequency of modern, high-performance processors (over 500 MHz, in some cases), limiting the power dissipation has become the most stringent design target. It is thus mandatory for processor engineers to resort to a large variety of optimization techniques to reduce the power requirements in the hot zones of the chip. In this paper, we focus on the power dissipated By the instruction fetch and decode logic, a portion of the processor architecture where a lot of capacitance switching normally takes place. We propose a methodology for determining an encoding of the instruction set that guarantees the minimization of the number of bit transitions occurring inside the registers of the pipeline stages involved in instruction fetching and decoding. The assignment of the binary patterns to the op-codes is driven by the statistics concerning instruction adjacency collected through instruction-level simulation of typical software applications; therefore, the technique is best exploited when applied to encode the instruction set of core processors and microcontrollers, since components of these types ore commonly used to execute fixed portions of machine code within embedded systems. We illustrate the effectiveness of the methodology through the experimental data we have obtained on an existing microprocessor. Luca Benini, Giovanni De Micheli, Alberto Macii, Enrico Macii, Massimo Poncino |
Great Lakes Symposium on VLSI | 3 |
| 1998 | Symbolic algorithms for layout-oriented synthesis of pass transistor logic circuitsabstractThis paper presents a nouel methodology for synthesizing PTL circuifg, whose disfincfiue feafures are the use of a symbolic algorifhm for the covem.ng of fhe initial network in ferms of PTL cells, and the eqloifation of layout-level ama and delay models dun.ng fhe selection of fhe be~t couem.ngsolufion.The results produced by the synthesis procedure on the full guite of fhe kcas'85 combinational circuifs are very encouraging. Fabrizio Ferrandi, Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi, Fabio Somenzi |
ICCAD | 2 |
| 1998 | Stream synthesis for efficient power simulation based on spectral transformsabstractIn this paper, we present a power estimation technique for control-flow intensive designs that is tailored towards driving iterative high-level synthesis systems, where hundreds of architectural trade-offs are explored and compared. Our method is fast and relatively accurate. The algorithm utilizes the behavioral information to extract branch probabilities, and uses these in conjunction with switching activity and circuit capacitance information, to estimate the power consumption of a given architecture. Alberto Macii, Enrico Macii, Massimo Poncino, Riccardo Scarsi |
ISLPED | 1 |