Fabio Frustaci

dblp:09/1379 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0001-5795-4321ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI Workloads
abstract
Energy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts.
Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci
DATE16
2024 KIT: Kernel Isotropic Transformation of Bilateral Filters for Image Denoising on FPGA
abstract
A Bilateral filter (BF) is commonly adopted as a pre-processing stage in several computer vision tasks because of its ability to denoise images. In contrast to the traditional image convolution that adopts a static kernel, a BF computes adaptive weights on-the-fly by applying exponentiation and division operations to the current pixel window. Prior works dealing with hardware acceleration of the BF rely on straightforward implementations that approximate the exponential function through look-up-tables (LUTs). This paper presents a new approximation technique to efficiently deploy a BF within real-time and low-energy intelligent systems based on FPGAs. The proposed strategy replaces the adaptive filter with its inexact isotropic version. This choice allows dropping a certain number of operations, thus resulting in enhanced speed and energy performances with respect to state-of-the-art hardware accelerators. When implemented on the AMD Xilinx Zynq XC7Z020 FPGA device, the proposed $5 \times 5$ BF design elaborates ∽237 Mega pixels per second and consumes at most 174 mW, with a Peak Signal-to-Noise Ratio (PSNR) degradation of just 0.55% at a noise standard deviation equal to 30.
Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri
FPL3
2024 An explainable embedded neural system for on-board ship detection from optical satellite imagery
abstract
Automatic ship detection from spaceborne systems such as satellites or aircrafts, raises considerable attention in sea surface monitoring because of the several applications in military and civilian field. In this context, processing satellite images on-board would reduce the latency time especially for emergency situations. In this paper, an hardware-oriented (HO) ship detection system based on a customized Convolutional Neural Network (CNN), here referred to as HO-ShipNet, is proposed and tested on a revised version of the “Ships in Satellite Imagery” (SSI) Kaggle dataset, reporting detection accuracy of up to 95%. Furthermore, the explainability of HO-ShipNet is investigated by means of explainable Artificial Intelligence (xAI) techniques (i.e., Local Interpretable Model-Agnostic Explanation (LIME) and Occlusion Sensitivuty Analysis (OSA)), in order to understand the reasoning behind the HO-ShipNet decisions by detecting the most important input features and consequently ensure the trustworthiness of the model itself. Finally, HO-ShipNet is also implemented on the heterogeneous Xilinx xc7z045ffg900-2 SoC Field Programmable Gate Array (FPGA) outperforming state-of-the-art FPGA-based accelerators dealing with high-resolution frames. The promising results encourage the potential deployment of the proposed system for on-board applications.
Cosimo Ieracitano, Nadia Mammone, Fanny Spagnolo, Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Francesco Carlo Morabito
Eng. Appl. Artif. Intell.4
2024 Approximate bilateral filters for real-time and low-energy imaging applications on FPGAs
abstract
Abstract Bilateral filtering is an image processing technique commonly adopted as intermediate step of several computer vision tasks. Opposite to the conventional image filtering, which is based on convolving the input pixels with a static kernel, the bilateral filtering computes its weights on the fly according to the current pixel values and some tuning parameters. Such additional elaborations involve nonlinear weighted averaging operations, which make difficult the deployment of bilateral filtering within existing vision technologies based on real-time and low-energy hardware architectures. This paper presents a new approximation strategy that aims to improve the energy efficiency of circuits implementing the bilateral filtering function, while preserving their real-time performances and elaboration accuracy. In contrast to the state-of-the-art, the proposed technique allows the filtering action to be on the fly adapted to both the current pixel values and to the tuning parameters, thus avoiding any architectural modification or tables update. When hardware implemented within the Xilinx Zynq XC7Z020 FPGA device, a 5 × 5 filter based on the proposed method processes 237.6 Mega pixels per second and consumes just 0.92 nJ per pixel, thus improving the energy efficiency by up to 2.8 times over the competitors. The impact of the proposed approximation on three different imaging applications has been also evaluated. Experiments demonstrate reasonable accuracy penalties over the accurate counterparts.
Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri
J. Supercomput.3
2024 Exploring the Usage of Fast Carry Chains to Implement Multistage Ring Oscillators on FPGAs: Design and Characterization
abstract
Ring oscillators (ROs) serve as basic building blocks in a lot of application scenarios, where they must ensure high reliability, flexibility, and low-area/energy footprint. With the recent advances of the Internet-of-Things (IoT) technology, in particular, the necessity to endow interconnected devices with security facilities has increased as well. In this context, the efficient implementation of ROs on field-programmable gate arrays (FPGAs) is crucial, even though it hides some pitfalls. This article presents a new design strategy for multistage ROs relying on the carry chains (CCs) available into modern FPGA devices. Several configurations of ROs designed as proposed here have been characterized in terms of hardware costs, jitter, and temperature/voltage sensitivity. In all the evaluated cases, the proposed design allows to achieve predictable routing schemes through the automatic place and route (P&R), while reducing slice occupancy and energy consumption by up to 50% and 44%, respectively, in comparison with the traditional lookup table (LUT)-based ROs. When realized on a Artix-7 device, the basic version of the proposed oscillator realized using 33 inverting stages allows obtaining multiphase outputs oscillating at 29.7 MHz with a standard deviation less than 10 kHz. The analysis conducted also demonstrates the high flexibility of the novel circuits, such as the possibility to easily change their behavior depending on the target application requirements. As an example, by exploiting additional pass-through elements, the proposed scheme achieves a sensitivity of 49 kHz/°C that is more than 4 times higher than that shown by the corresponding traditional LUT-based competitor, thus making it more suitable for thermal monitoring applications.
Fanny Spagnolo, Stefania Perri, Massimo Vatalaro, Fabio Frustaci, Felice Crupi, Pasquale Corsonello
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Energy-Quality Scalable Adders Based on Nonzeroing Bit Truncation
abstract
Approximate addition is a technique to trade off energy consumption and output quality in error-tolerant applications. In prior art, bit truncation has been explored as a lever to dynamically trade off energy and quality. In this brief, an innovative bit truncation strategy is proposed to achieve more graceful quality degradation compared to state-of-the-art truncation schemes. This translates into energy reduction at a given quality target. When applied to a ripple-carry adder, the proposed bit truncation approach improves quality by up to 8.5 dB in terms of peak signal-to-noise ratio, compared to traditional bit truncation. As a case study, the proposed approach was applied to a discrete cosine transform engine. In comparison with prior art, the proposed approach reduces energy by 20%, at insignificant delay and silicon area overhead.
Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Massimo Alioto
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Design of Real-Time FPGA-based Embedded System for Stereo Vision
abstract
This paper describes a novel heterogeneous SoC FPGA-based embedded system for stereo vision. Two complete implementations are presented and characterized. In both designs the auxiliary computations, such as the image rectification and the disparity map refinement, are performed by the custom hardware module purpose-designed to compute disparity maps, thus achieving very high speeds. The software routine run by the on-chip general-purpose processor is used to control configuration and communication. Obtained results show that, in comparison with several existing hardware designs, the proposed system reaches higher performances, competitive accuracies, lower complexity and higher flexibility.
Stefania Perri, Fabio Frustaci, Fanny Spagnolo, Pasquale Corsonello
ISCAS2
2016 Approximate SRAMs With Dynamic Energy-Quality Management
abstract
In this paper, approximate SRAMs are explored in the context of error-tolerant applications, in which energy is saved at the cost of the occurrence of read/write errors (i.e., signal quality degradation). This analysis investigates variation-resilient techniques that enable dynamic management of the energy-quality tradeoff down to the bit level. In these techniques, the different impacts of errors on quality at different bit positions are explicitly considered as key enabler of energy savings that are far larger than a simple voltage scaling. The analysis is based on the experimental results in an energy-quality scalable 28-nm SRAM and the extrapolation to a wide range of conditions through the models that combine the individual energy contributions. Results show that the joint adoption of multiple bit-level techniques provides substantially larger energy gains than individual techniques. Compared with the simple voltage scaling at isoquality, the joint adoption of these techniques can provide more than $2\times $ energy reduction at negligible area penalty. Energy savings turn out to be highly sensitive to the choice of joint techniques, thus showing the crucial importance of dynamic energy-quality management in approximate SRAMs.
Fabio Frustaci, David T. Blaauw, Dennis Sylvester, Massimo Alioto
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Exploring well configurations for voltage level converter design in 28 nm UTBB FDSOI technology
abstract
Voltage level converters are critical components in multi supply ultra-low voltage designs, especially when signals need to be converted from the sub-threshold to the above-threshold domain. In these designs, advanced technology processes, such as the Ultra-Thin Body and Buried oxide (UTBB) Fully-Depleted SOI (FDSOI), are greatly desired since they intrinsically allow controlling the Drain Induced Barrier Lowering effect (DIBL) and the Gate Induced Drain Leakage (GIDL), in addition to the reduction of the effects of process variations. Moreover, these technologies provide a group of architectural and device-level techniques for threshold voltage adjustment that can be efficiently adopted to combine high performances and low energy consumption. However, specific design strategies should be applied to efficiently exploit all these potentialities. This paper investigates how the physical design of level converters can benefit from the synergistic adoption of the knobs available in the UTBB FDSOI technology (poly biasing, flip-well, single-well, back biasing). In particular, three mixed single well configurations have been implemented and analyzed. This research work demonstrates that the specific selected approach allows decreasing the energy per cycle consumption, the leakage current and the delay by up to 35.3%, 70.4%, and 6.2%, respectively, with respect to the basic conventional design strategy. Furthermore, statistical analysis confirmed that these advantages are maintained for a wide range of process variations, also improving the functional yield and the minimum input voltage causing the level converter failure.
Pasquale Corsonello, Stefania Perri, Fabio Frustaci
ICCD3
2015 Power supply noise in accurate delay model for the sub-threshold domain
Pasquale Corsonello, Fabio Frustaci, Stefania Perri
Integr.2
2015 Low-Leakage SRAM Wordline Drivers for the 28-nm UTBB FDSOI Technology
abstract
This brief deals with a new design of low-power SRAM wordline decoder in the 28-nm ultrathin body and buried oxide (UTBB) fully depleted silicon-on-insulator (FDSOI) technology. The proposed approach synergistically adopts the poly biasing technique in conjunction with single-well/flip-well configurations and body biasing to opportunely tune the threshold voltage of the devices in the standby and active mode. A tuning methodology is described to optimize the static energy consumption. Post-layout simulations, done at power supply voltages ranging between 1 V and 0.5 V, have shown that, in comparison with the state-of-the-art techniques based on the same UTBB FDSOI technology, the proposed design achieves a maximum leakage up to 85% lower without paying significant delay penalties.
Pasquale Corsonello, Fabio Frustaci, Stefania Perri
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Tapered-VTH CMOS buffer design for improved energy efficiency in deep nanometer technology
abstract
In this paper, the novel "tapered-Vth" approach to design energy-efficient CMOS buffers is introduced. In this approach, the substantial energy consumption due to leakage is reduced by tapering the threshold voltage throughout the buffer stages, other than tapering the transistor size. More specifically, the threshold voltage is progressively reduced when going from the last to the first stage. This enables a considerable leakage reduction in the last stages (which contribute most to the overall leakage) at the price of a higher delay. The resulting delay penalty is then compensated by reducing the transistor threshold voltage in the first stages, with an insignificant leakage increase (they contribute very little to the overall buffer leakage). Simulation results based on a commercial 45-nm 1-V CMOS technology show that the proposed "tapered-VTH" approach can considerably improve the energy efficiency of CMOS buffers over the entire spectrum of possible energy-delay tradeoffs, from high speed to low power.
Fabio Frustaci, Pasquale Corsonello, Massimo Alioto
ISCAS1
2010 A new low-power high-speed single-clock-cycle binary comparator
abstract
This paper presents a new ultra-low power high-speed single-clock-cycle binary comparator. It is based on a novel parallel-prefix algorithm which drastically reduces the switching activity of the internal nodes of the circuit. When implemented by using the ST 90nm-1V technology, the proposed 64-bit comparator exhibits an energy dissipation of only 0.77μW/MHz and a delay of 258ps. With respect to a recently published low-power high-speed parallel-prefix adder, the proposed design shows an energy dissipation reduction of 23% and a speed improvement of 7%.
Fabio Frustaci, Stefania Perri, Marco Lanuzza, Pasquale Corsonello
ISCAS1
2010 Exploiting Self-Reconfiguration Capability to Improve SRAM-based FPGA Robustness in Space and Avionics Applications
abstract
This article presents a novel configuration scrubbing core, used for internal detection and correction of radiation-induced configuration single and multiple bit errors, without requiring external scrubbing. The proposed technique combines the benefits of fast radiation-induced fault detection with fast restoration of the device functionality and small area and power overheads. Experimental results demonstrate that the novel approach significantly improves the availability in hostile radiation environments of FPGA-based designs. When implemented using a Xilinx XC2V1000 Virtex-II device, the presented technique detects and corrects single bit upsets and double, triple and quadruple multi bit upsets, occupying just 1488 slices and dissipating less than 30 mW at a 50MHz running frequency.
Marco Lanuzza, Paolo Zicari, Fabio Frustaci, Stefania Perri, Pasquale Corsonello
ACM Trans. Reconfigurable Technol. Syst.3
2006 Leakage energy reduction techniques in deep submicron cache memories: a comparative study
abstract
Static energy consumption due to subthreshold leakage current is one of the main concern in on-chip level-1 and level-2 cache. In the last few years several techniques have been proposed to limit the subthreshold current in a SRAM cell. Unfortunately, these techniques also increase the dynamic energy during the cell access operation, with respect to the conventional SRAM architecture. In this paper the actual energy saving offered by low leakage approaches is investigated, within the context of a microprocessor memory hierarchy, taking into account their dynamic energy overheads. Simulation based on UMC 0.18mum-1.8V and ST 90nm-1V process models have been performed. Results show that, for both the technologies, the leakage energy saving achieved by the analyzed techniques in the first cache level turns out to be inadequate, owing to the extra dynamic energy dissipation. Only in UL2 they assure a net energy saving due to the smaller number of accesses
Fabio Frustaci, Pasquale Corsonello, Stefania Perri, Giuseppe Cocorullo
ISCAS1
2006 Techniques for Leakage Energy Reduction in Deep Submicrometer Cache Memories
abstract
The techniques known in literature for the design of SRAM structures with low standby leakage typically exploit an additional operation mode, named the sleep mode or the standby mode. In this paper, existing low leakage SRAM structures are analyzed by several SPEC2000 benchmarks. As expected, the examined SRAM architectures have static power consumption lower than the conventional 6-T SRAM cell. However, the additional activities performed to enter and to exit the sleep mode also lead to higher dynamic energy. Our study demonstrates that, due to this, the overall energy consumption achieved by the known low-leakage techniques is greater than the conventional approach. In the second part of this paper, a novel low-leakage SRAM cell is presented. The proposed structure establishes when to enter and to exit the sleep mode, on the basis of the data stored in it, without introducing time and energy penalties with respect to the conventional 6-T cell. The new SRAM structure was realized using the UMC 0.18-mum, 1.8-V, and the ST 90-nm 1-V CMOS technologies. Tests performed with a set of SPEC2000 benchmarks have shown that the proposed approach is actually energy efficient
Fabio Frustaci, Pasquale Corsonello, Stefania Perri, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.1