EDBT 2026 Demo / reviewers in the wild / expert
Fanny Spagnolo
dblp:228/3374
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-2197-4563ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI WorkloadsabstractEnergy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
DATE | 13 |
| 2025 | Multi-Partner Project: Architectures and Design Methodologies to Accelerate AI Workloads. The ICSC Flagship 2 ProjectabstractRecent pre-exascale and exascale supercomputers have driven the development of increasingly sophisticated AI models for diverse applications, including image recognition and classification, natural language processing, and generative AI. These applications require specialized hardware accelerators, to handle the heavy computational demands of AI algorithms in an energy-efficient manner. Today, AI accelerators are deployed across various systems, from low-power edge devices to large-scale servers, high-performance computing (HPC) infrastructures, and data centers. The primary objective of the ICSC Flagship 2 project, discussed in this paper, is to develop heterogeneous hardware platforms optimized to accelerate HPC and big data applications. Specifically, this paper provides an overview of the key challenges addressed and the achievements realized at the current intermediate stage of the ICSC Flagship 2 project focused on architectures, technologies, and design methodologies to design efficient hardware accelerators for AI workloads, such as deep learning (DL) and transformer models. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Stefania Perri, Fanny Spagnolo, Pasquale Corsonello, Sebastiano Fabio Schifano, Cristian Zambelli, Angelo Garofalo, Francesco Conti 0001, Luca Benini |
DATE | 6 |
| 2025 | C4TERO: Configurable Cascaded Carry Chains for High Reliability TERO PUFs on FPGAsabstractIn this paper we present a novel Transient Effect Ring Oscillator Physical Unclonable Function for FPGAs. It exploits in an original way the carry chain resources available in modern devices. The basic cell adopted in the proposed architecture can be runtime configured to implement different oscillation paths. This property enables the possibility to output more than one bit response per cell by choosing among the configurations those that exhibit the highest reliability. Such results are achieved by adopting a specific calibration process able to identify configurations of the cells showing the highest stability and the most uncorrelated responses. When implemented on several Series 7 Xilinx devices, no unstable bits were observed at 1 V and$25~^{\circ }$C. Under voltage variation in the manufacturer recommended ranges, a worst case bit error rate of 0.046% is achieved. The circuit designed as here described consists of 64 cells, produces 128 response bits and consumes just 535 look-up-tables and 256 carry chains. Fanny Spagnolo, Massimo Vatalaro, Stefania Perri, Felice Crupi, Pasquale Corsonello |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | A Novel Compressive Sensing Method for Secure and Energy Efficient ECG Signal Transmission ApplicationsabstractThis paper introduces a novel Compressive Sensing (CS)-based cryptosystem tailored for the secure and efficient transmission of Electrocardiogram (ECG) signals in Internet of Medical Things (IoMT) environments. It leverages the inherent sparsity of ECG signals in the wavelet domain and ensures both data privacy and integrity during transmission through a simple but effective additional encryption stage. We have evaluated the performance of four distinct sensing matrices and three reconstruction algorithms across multiple wavelet families. Through extensive simulations using the MIT-BIH Arrhythmia Database we demonstrate that the Low-Density Parity-Check matrix combined with the L1 optimization algorithm achieves the highest Quality Score, with Compression Ratio up 50%. The proposed approach also shows an excellent ability in preserving important pathological features in presence of abnormal beats. The proposed CS encoder has been hardware implemented on low-resource microcontroller and FPGA devices. When realized on a Xilinx Artix 7 XC7A12 T FPGA, such a prototype allows real-time operations to be sustained running at 1 MHz clock frequency and dissipating only 0.8nJ per sample. Fanny Spagnolo, Bharat Lal, Pasquale Corsonello, Raffaele Gravina |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | KIT: Kernel Isotropic Transformation of Bilateral Filters for Image Denoising on FPGAabstractA Bilateral filter (BF) is commonly adopted as a pre-processing stage in several computer vision tasks because of its ability to denoise images. In contrast to the traditional image convolution that adopts a static kernel, a BF computes adaptive weights on-the-fly by applying exponentiation and division operations to the current pixel window. Prior works dealing with hardware acceleration of the BF rely on straightforward implementations that approximate the exponential function through look-up-tables (LUTs). This paper presents a new approximation technique to efficiently deploy a BF within real-time and low-energy intelligent systems based on FPGAs. The proposed strategy replaces the adaptive filter with its inexact isotropic version. This choice allows dropping a certain number of operations, thus resulting in enhanced speed and energy performances with respect to state-of-the-art hardware accelerators. When implemented on the AMD Xilinx Zynq XC7Z020 FPGA device, the proposed $5 \times 5$ BF design elaborates ∽237 Mega pixels per second and consumes at most 174 mW, with a Peak Signal-to-Noise Ratio (PSNR) degradation of just 0.55% at a noise standard deviation equal to 30. Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri |
FPL | 1 |
| 2024 | An explainable embedded neural system for on-board ship detection from optical satellite imageryabstractAutomatic ship detection from spaceborne systems such as satellites or aircrafts, raises considerable attention in sea surface monitoring because of the several applications in military and civilian field. In this context, processing satellite images on-board would reduce the latency time especially for emergency situations. In this paper, an hardware-oriented (HO) ship detection system based on a customized Convolutional Neural Network (CNN), here referred to as HO-ShipNet, is proposed and tested on a revised version of the “Ships in Satellite Imagery” (SSI) Kaggle dataset, reporting detection accuracy of up to 95%. Furthermore, the explainability of HO-ShipNet is investigated by means of explainable Artificial Intelligence (xAI) techniques (i.e., Local Interpretable Model-Agnostic Explanation (LIME) and Occlusion Sensitivuty Analysis (OSA)), in order to understand the reasoning behind the HO-ShipNet decisions by detecting the most important input features and consequently ensure the trustworthiness of the model itself. Finally, HO-ShipNet is also implemented on the heterogeneous Xilinx xc7z045ffg900-2 SoC Field Programmable Gate Array (FPGA) outperforming state-of-the-art FPGA-based accelerators dealing with high-resolution frames. The promising results encourage the potential deployment of the proposed system for on-board applications. Cosimo Ieracitano, Nadia Mammone, Fanny Spagnolo, Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Francesco Carlo Morabito |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Approximate bilateral filters for real-time and low-energy imaging applications on FPGAsabstractAbstract Bilateral filtering is an image processing technique commonly adopted as intermediate step of several computer vision tasks. Opposite to the conventional image filtering, which is based on convolving the input pixels with a static kernel, the bilateral filtering computes its weights on the fly according to the current pixel values and some tuning parameters. Such additional elaborations involve nonlinear weighted averaging operations, which make difficult the deployment of bilateral filtering within existing vision technologies based on real-time and low-energy hardware architectures. This paper presents a new approximation strategy that aims to improve the energy efficiency of circuits implementing the bilateral filtering function, while preserving their real-time performances and elaboration accuracy. In contrast to the state-of-the-art, the proposed technique allows the filtering action to be on the fly adapted to both the current pixel values and to the tuning parameters, thus avoiding any architectural modification or tables update. When hardware implemented within the Xilinx Zynq XC7Z020 FPGA device, a 5 × 5 filter based on the proposed method processes 237.6 Mega pixels per second and consumes just 0.92 nJ per pixel, thus improving the energy efficiency by up to 2.8 times over the competitors. The impact of the proposed approximation on three different imaging applications has been also evaluated. Experiments demonstrate reasonable accuracy penalties over the accurate counterparts. Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri |
J. Supercomput. | 1 |
| 2024 | Exploring the Usage of Fast Carry Chains to Implement Multistage Ring Oscillators on FPGAs: Design and CharacterizationabstractRing oscillators (ROs) serve as basic building blocks in a lot of application scenarios, where they must ensure high reliability, flexibility, and low-area/energy footprint. With the recent advances of the Internet-of-Things (IoT) technology, in particular, the necessity to endow interconnected devices with security facilities has increased as well. In this context, the efficient implementation of ROs on field-programmable gate arrays (FPGAs) is crucial, even though it hides some pitfalls. This article presents a new design strategy for multistage ROs relying on the carry chains (CCs) available into modern FPGA devices. Several configurations of ROs designed as proposed here have been characterized in terms of hardware costs, jitter, and temperature/voltage sensitivity. In all the evaluated cases, the proposed design allows to achieve predictable routing schemes through the automatic place and route (P&R), while reducing slice occupancy and energy consumption by up to 50% and 44%, respectively, in comparison with the traditional lookup table (LUT)-based ROs. When realized on a Artix-7 device, the basic version of the proposed oscillator realized using 33 inverting stages allows obtaining multiphase outputs oscillating at 29.7 MHz with a standard deviation less than 10 kHz. The analysis conducted also demonstrates the high flexibility of the novel circuits, such as the possibility to easily change their behavior depending on the target application requirements. As an example, by exploiting additional pass-through elements, the proposed scheme achieves a sensitivity of 49 kHz/°C that is more than 4 times higher than that shown by the corresponding traditional LUT-based competitor, thus making it more suitable for thermal monitoring applications. Fanny Spagnolo, Stefania Perri, Massimo Vatalaro, Fabio Frustaci, Felice Crupi, Pasquale Corsonello |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | ERMES: Efficient Racetrack Memory Emulation System based on FPGAabstractWith the scaling of CMOS technology almost over, non-volatile memories based on emerging technologies are gaining considerable popularity. Particularly, spintronic-based Racetrack memories (RTMs) exhibit unprecedented storage capacity, as well as reduced energy per operation and high write endurance, which make them promising candidates to revolutionize the architecture of memory sub-systems. However, since RTM exploits shifting of magnetic domains to align the required data with the access port, its read/write latency is not constant. Due to this behaviour, several performance optimizations related to the target application may be introduced either on memory architecture or data placement or both. To this purpose, specific tools able to emulate the timing characteristics of RTMs are highly desired. Unfortunately, existing software-based simulators show poor flexibility and run-time. To address such limitations, this paper presents a new emulation system for RTMs based on heterogeneous FPGA-CPU Systems-on-Chips (SoCs). Thanks to its high flexibility, the proposed emulator can be easily configured to evaluate different memory architectures. In addition, the CPU can be used to stimulate the RTM architecture under test with appropriate benchmarks, thus providing a fast self-contained evaluation environment. As case study, ERMES has been implemented within the Xilinx Zynq Ultrascale XCUZ9EG SoC to evaluate performances of several memory configurations when running benchmark applications from the MiBench suite, experiencing a speed-up higher than × 146 over software-based simulators. Fanny Spagnolo, Salim Ullah, Pasquale Corsonello, Akash Kumar 0001 |
FPL | 1 |
| 2020 | An Efficient Convolution Engine based on the À-trous Spatial Pyramid PoolingabstractThis paper presents an efficient hardware architecture able to perform 2D dilated convolutions and suitable for the integration within modern heterogeneous embedded systems targeting semantic image segmentation. The proposed design supports multiple dilation rates. Moreover, it uses limited amounts of resources even when large convolution windows are processed. As a case study, the novel circuit has been integrated within a Xilinx Zynq-7000 FPSoC device to accelerate a state-of-the-art CNN model for medical images segmentation. Obtained results demonstrate that higher computational capabilities, reduced resources utilization and lower power consumption are achieved with respect to the competitors existing in literature. Cristian Sestito, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri |
ASAP | 2 |
| 2020 | Design of a real-time face detection architecture for heterogeneous systems-on-chips
Fanny Spagnolo, Stefania Perri, Pasquale Corsonello |
Integr. | 1 |
| 2018 | Design of Real-Time FPGA-based Embedded System for Stereo VisionabstractThis paper describes a novel heterogeneous SoC FPGA-based embedded system for stereo vision. Two complete implementations are presented and characterized. In both designs the auxiliary computations, such as the image rectification and the disparity map refinement, are performed by the custom hardware module purpose-designed to compute disparity maps, thus achieving very high speeds. The software routine run by the on-chip general-purpose processor is used to control configuration and communication. Obtained results show that, in comparison with several existing hardware designs, the proposed system reaches higher performances, competitive accuracies, lower complexity and higher flexibility. Stefania Perri, Fabio Frustaci, Fanny Spagnolo, Pasquale Corsonello |
ISCAS | 3 |