Stefania Perri

dblp:22/6492 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0003-1363-9201ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI Workloads
abstract
Energy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts.
Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci
DATE15
2025 Multi-Partner Project: Architectures and Design Methodologies to Accelerate AI Workloads. The ICSC Flagship 2 Project
abstract
Recent pre-exascale and exascale supercomputers have driven the development of increasingly sophisticated AI models for diverse applications, including image recognition and classification, natural language processing, and generative AI. These applications require specialized hardware accelerators, to handle the heavy computational demands of AI algorithms in an energy-efficient manner. Today, AI accelerators are deployed across various systems, from low-power edge devices to large-scale servers, high-performance computing (HPC) infrastructures, and data centers. The primary objective of the ICSC Flagship 2 project, discussed in this paper, is to develop heterogeneous hardware platforms optimized to accelerate HPC and big data applications. Specifically, this paper provides an overview of the key challenges addressed and the achievements realized at the current intermediate stage of the ICSC Flagship 2 project focused on architectures, technologies, and design methodologies to design efficient hardware accelerators for AI workloads, such as deep learning (DL) and transformer models.
Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Stefania Perri, Fanny Spagnolo, Pasquale Corsonello, Sebastiano Fabio Schifano, Cristian Zambelli, Angelo Garofalo, Francesco Conti 0001, Luca Benini
DATE5
2025 C4TERO: Configurable Cascaded Carry Chains for High Reliability TERO PUFs on FPGAs
abstract
In this paper we present a novel Transient Effect Ring Oscillator Physical Unclonable Function for FPGAs. It exploits in an original way the carry chain resources available in modern devices. The basic cell adopted in the proposed architecture can be runtime configured to implement different oscillation paths. This property enables the possibility to output more than one bit response per cell by choosing among the configurations those that exhibit the highest reliability. Such results are achieved by adopting a specific calibration process able to identify configurations of the cells showing the highest stability and the most uncorrelated responses. When implemented on several Series 7 Xilinx devices, no unstable bits were observed at 1 V and$25~^{\circ }$C. Under voltage variation in the manufacturer recommended ranges, a worst case bit error rate of 0.046% is achieved. The circuit designed as here described consists of 64 cells, produces 128 response bits and consumes just 535 look-up-tables and 256 carry chains.
Fanny Spagnolo, Massimo Vatalaro, Stefania Perri, Felice Crupi, Pasquale Corsonello
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 KIT: Kernel Isotropic Transformation of Bilateral Filters for Image Denoising on FPGA
abstract
A Bilateral filter (BF) is commonly adopted as a pre-processing stage in several computer vision tasks because of its ability to denoise images. In contrast to the traditional image convolution that adopts a static kernel, a BF computes adaptive weights on-the-fly by applying exponentiation and division operations to the current pixel window. Prior works dealing with hardware acceleration of the BF rely on straightforward implementations that approximate the exponential function through look-up-tables (LUTs). This paper presents a new approximation technique to efficiently deploy a BF within real-time and low-energy intelligent systems based on FPGAs. The proposed strategy replaces the adaptive filter with its inexact isotropic version. This choice allows dropping a certain number of operations, thus resulting in enhanced speed and energy performances with respect to state-of-the-art hardware accelerators. When implemented on the AMD Xilinx Zynq XC7Z020 FPGA device, the proposed $5 \times 5$ BF design elaborates ∽237 Mega pixels per second and consumes at most 174 mW, with a Peak Signal-to-Noise Ratio (PSNR) degradation of just 0.55% at a noise standard deviation equal to 30.
Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri
FPL4
2024 An explainable embedded neural system for on-board ship detection from optical satellite imagery
abstract
Automatic ship detection from spaceborne systems such as satellites or aircrafts, raises considerable attention in sea surface monitoring because of the several applications in military and civilian field. In this context, processing satellite images on-board would reduce the latency time especially for emergency situations. In this paper, an hardware-oriented (HO) ship detection system based on a customized Convolutional Neural Network (CNN), here referred to as HO-ShipNet, is proposed and tested on a revised version of the “Ships in Satellite Imagery” (SSI) Kaggle dataset, reporting detection accuracy of up to 95%. Furthermore, the explainability of HO-ShipNet is investigated by means of explainable Artificial Intelligence (xAI) techniques (i.e., Local Interpretable Model-Agnostic Explanation (LIME) and Occlusion Sensitivuty Analysis (OSA)), in order to understand the reasoning behind the HO-ShipNet decisions by detecting the most important input features and consequently ensure the trustworthiness of the model itself. Finally, HO-ShipNet is also implemented on the heterogeneous Xilinx xc7z045ffg900-2 SoC Field Programmable Gate Array (FPGA) outperforming state-of-the-art FPGA-based accelerators dealing with high-resolution frames. The promising results encourage the potential deployment of the proposed system for on-board applications.
Cosimo Ieracitano, Nadia Mammone, Fanny Spagnolo, Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Francesco Carlo Morabito
Eng. Appl. Artif. Intell.5
2024 Approximate bilateral filters for real-time and low-energy imaging applications on FPGAs
abstract
Abstract Bilateral filtering is an image processing technique commonly adopted as intermediate step of several computer vision tasks. Opposite to the conventional image filtering, which is based on convolving the input pixels with a static kernel, the bilateral filtering computes its weights on the fly according to the current pixel values and some tuning parameters. Such additional elaborations involve nonlinear weighted averaging operations, which make difficult the deployment of bilateral filtering within existing vision technologies based on real-time and low-energy hardware architectures. This paper presents a new approximation strategy that aims to improve the energy efficiency of circuits implementing the bilateral filtering function, while preserving their real-time performances and elaboration accuracy. In contrast to the state-of-the-art, the proposed technique allows the filtering action to be on the fly adapted to both the current pixel values and to the tuning parameters, thus avoiding any architectural modification or tables update. When hardware implemented within the Xilinx Zynq XC7Z020 FPGA device, a 5 × 5 filter based on the proposed method processes 237.6 Mega pixels per second and consumes just 0.92 nJ per pixel, thus improving the energy efficiency by up to 2.8 times over the competitors. The impact of the proposed approximation on three different imaging applications has been also evaluated. Experiments demonstrate reasonable accuracy penalties over the accurate counterparts.
Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri
J. Supercomput.4
2024 Exploring the Usage of Fast Carry Chains to Implement Multistage Ring Oscillators on FPGAs: Design and Characterization
abstract
Ring oscillators (ROs) serve as basic building blocks in a lot of application scenarios, where they must ensure high reliability, flexibility, and low-area/energy footprint. With the recent advances of the Internet-of-Things (IoT) technology, in particular, the necessity to endow interconnected devices with security facilities has increased as well. In this context, the efficient implementation of ROs on field-programmable gate arrays (FPGAs) is crucial, even though it hides some pitfalls. This article presents a new design strategy for multistage ROs relying on the carry chains (CCs) available into modern FPGA devices. Several configurations of ROs designed as proposed here have been characterized in terms of hardware costs, jitter, and temperature/voltage sensitivity. In all the evaluated cases, the proposed design allows to achieve predictable routing schemes through the automatic place and route (P&R), while reducing slice occupancy and energy consumption by up to 50% and 44%, respectively, in comparison with the traditional lookup table (LUT)-based ROs. When realized on a Artix-7 device, the basic version of the proposed oscillator realized using 33 inverting stages allows obtaining multiphase outputs oscillating at 29.7 MHz with a standard deviation less than 10 kHz. The analysis conducted also demonstrates the high flexibility of the novel circuits, such as the possibility to easily change their behavior depending on the target application requirements. As an example, by exploiting additional pass-through elements, the proposed scheme achieves a sensitivity of 49 kHz/°C that is more than 4 times higher than that shown by the corresponding traditional LUT-based competitor, thus making it more suitable for thermal monitoring applications.
Fanny Spagnolo, Stefania Perri, Massimo Vatalaro, Fabio Frustaci, Felice Crupi, Pasquale Corsonello
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Accuracy Evaluation of Transposed Convolution-Based Quantized Neural Networks
abstract
Several modern applications in the field of Artificial Intelligence exploit deep learning to make accurate decisions. Recent work on compression techniques allows for deep learning applications, such as computer vision, to run on Edge Computing devices. For instance, quantizing the precision of deep learning architectures allows Edge Computing devices to achieve high throughput at low power. Quantization has been mainly focused on multilayer perceptrons and convolution-based models for classification problems. However, its impact over more complex scenarios, such as image up-sampling, is still underexplored. This paper presents a systematic evaluation of the accuracy achieved by quantized neural networks when performing image up-sampling in three different applications: image compression/decompression, synthetic image generation and semantic segmentation. Taking into account the promising attitude of learnable filters to predict pixels, transposed convolutional layers are used for up-sampling. Experimental results based on analytical metrics show that acceptable accuracies are reached with quantization spanning between 3 and 7 bits. Based on the visual inspection, the range 2–6 bits guarantees appropriate accuracy.
Cristian Sestito, Stefania Perri, Robert J. Stewart 0001
IJCNN2
2020 An Efficient Convolution Engine based on the À-trous Spatial Pyramid Pooling
abstract
This paper presents an efficient hardware architecture able to perform 2D dilated convolutions and suitable for the integration within modern heterogeneous embedded systems targeting semantic image segmentation. The proposed design supports multiple dilation rates. Moreover, it uses limited amounts of resources even when large convolution windows are processed. As a case study, the novel circuit has been integrated within a Xilinx Zynq-7000 FPSoC device to accelerate a state-of-the-art CNN model for medical images segmentation. Obtained results demonstrate that higher computational capabilities, reduced resources utilization and lower power consumption are achieved with respect to the competitors existing in literature.
Cristian Sestito, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri
ASAP4
2020 Design of a real-time face detection architecture for heterogeneous systems-on-chips
Fanny Spagnolo, Stefania Perri, Pasquale Corsonello
Integr.2
2019 Energy-Quality Scalable Adders Based on Nonzeroing Bit Truncation
abstract
Approximate addition is a technique to trade off energy consumption and output quality in error-tolerant applications. In prior art, bit truncation has been explored as a lever to dynamically trade off energy and quality. In this brief, an innovative bit truncation strategy is proposed to achieve more graceful quality degradation compared to state-of-the-art truncation schemes. This translates into energy reduction at a given quality target. When applied to a ripple-carry adder, the proposed bit truncation approach improves quality by up to 8.5 dB in terms of peak signal-to-noise ratio, compared to traditional bit truncation. As a case study, the proposed approach was applied to a discrete cosine transform engine. In comparison with prior art, the proposed approach reduces energy by 20%, at insignificant delay and silicon area overhead.
Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Massimo Alioto
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Design of Real-Time FPGA-based Embedded System for Stereo Vision
abstract
This paper describes a novel heterogeneous SoC FPGA-based embedded system for stereo vision. Two complete implementations are presented and characterized. In both designs the auxiliary computations, such as the image rectification and the disparity map refinement, are performed by the custom hardware module purpose-designed to compute disparity maps, thus achieving very high speeds. The software routine run by the on-chip general-purpose processor is used to control configuration and communication. Obtained results show that, in comparison with several existing hardware designs, the proposed system reaches higher performances, competitive accuracies, lower complexity and higher flexibility.
Stefania Perri, Fabio Frustaci, Fanny Spagnolo, Pasquale Corsonello
ISCAS1
2015 Exploring well configurations for voltage level converter design in 28 nm UTBB FDSOI technology
abstract
Voltage level converters are critical components in multi supply ultra-low voltage designs, especially when signals need to be converted from the sub-threshold to the above-threshold domain. In these designs, advanced technology processes, such as the Ultra-Thin Body and Buried oxide (UTBB) Fully-Depleted SOI (FDSOI), are greatly desired since they intrinsically allow controlling the Drain Induced Barrier Lowering effect (DIBL) and the Gate Induced Drain Leakage (GIDL), in addition to the reduction of the effects of process variations. Moreover, these technologies provide a group of architectural and device-level techniques for threshold voltage adjustment that can be efficiently adopted to combine high performances and low energy consumption. However, specific design strategies should be applied to efficiently exploit all these potentialities. This paper investigates how the physical design of level converters can benefit from the synergistic adoption of the knobs available in the UTBB FDSOI technology (poly biasing, flip-well, single-well, back biasing). In particular, three mixed single well configurations have been implemented and analyzed. This research work demonstrates that the specific selected approach allows decreasing the energy per cycle consumption, the leakage current and the delay by up to 35.3%, 70.4%, and 6.2%, respectively, with respect to the basic conventional design strategy. Furthermore, statistical analysis confirmed that these advantages are maintained for a wide range of process variations, also improving the functional yield and the minimum input voltage causing the level converter failure.
Pasquale Corsonello, Stefania Perri, Fabio Frustaci
ICCD2
2015 Power supply noise in accurate delay model for the sub-threshold domain
Pasquale Corsonello, Fabio Frustaci, Stefania Perri
Integr.3
2015 Low-Leakage SRAM Wordline Drivers for the 28-nm UTBB FDSOI Technology
abstract
This brief deals with a new design of low-power SRAM wordline decoder in the 28-nm ultrathin body and buried oxide (UTBB) fully depleted silicon-on-insulator (FDSOI) technology. The proposed approach synergistically adopts the poly biasing technique in conjunction with single-well/flip-well configurations and body biasing to opportunely tune the threshold voltage of the devices in the standby and active mode. A tuning methodology is described to optimize the static energy consumption. Post-layout simulations, done at power supply voltages ranging between 1 V and 0.5 V, have shown that, in comparison with the state-of-the-art techniques based on the same UTBB FDSOI technology, the proposed design achieves a maximum leakage up to 85% lower without paying significant delay penalties.
Pasquale Corsonello, Fabio Frustaci, Stefania Perri
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Fast and Wide Range Voltage Conversion in Multisupply Voltage Designs
abstract
Multisupply voltage design technique is widely used in modern system-on-chips to tradeoff energy and speed. Level shifters (LSs) allow different voltage domains to be interfaced. In this brief, a new LS is presented for fast and wide range voltage conversion. Because of a novel architecture combined with the use of multithreshold CMOS technique, the proposed circuit guarantees robust voltage shifting from the deep subthreshold to the above-threshold domain while exhibiting fast response and low energy consumption. When implemented in a 90-nm technology node, considering process-voltage-temperature variations, the proposed design reliably converts 100-mV input signals into 1 V output signals. Post-layout simulation results demonstrate that the new LS shows a propagation delay of 16.6 ns, a static power dissipation of 8.7 nW and a total energy per transition of only 77 fJ for a 0.2 V 1-MHz input pulse.
Marco Lanuzza, Pasquale Corsonello, Stefania Perri
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Area-Delay Efficient Binary Adders in QCA
abstract
As transistors decrease in size more and more of them can be accommodated in a single die, thus increasing chip computational capabilities. However, transistors cannot get much smaller than their current size. The quantum-dot cellular automata (QCA) approach represents one of the possible solutions in overcoming this physical limit, even though the design of logic modules in QCA is not always straightforward. In this brief, we propose a new adder that outperforms all state-of-the-art competitors and achieves the best area-delay tradeoff. The above advantages are obtained by using an overall area similar to the cheaper designs known in literature. The 64-bit version of the novel adder spans over 18.72 μ2of active area and shows a delay of only nine clock cycles, that is just 36 clock phases.
Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Adaptive Census Transform: A novel hardware-oriented stereovision algorithm
Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo
Comput. Vis. Image Underst.1
2010 A new low-power high-speed single-clock-cycle binary comparator
abstract
This paper presents a new ultra-low power high-speed single-clock-cycle binary comparator. It is based on a novel parallel-prefix algorithm which drastically reduces the switching activity of the internal nodes of the circuit. When implemented by using the ST 90nm-1V technology, the proposed 64-bit comparator exhibits an energy dissipation of only 0.77μW/MHz and a delay of 258ps. With respect to a recently published low-power high-speed parallel-prefix adder, the proposed design shows an energy dissipation reduction of 23% and a speed improvement of 7%.
Fabio Frustaci, Stefania Perri, Marco Lanuzza, Pasquale Corsonello
ISCAS2
2010 Exploiting Self-Reconfiguration Capability to Improve SRAM-based FPGA Robustness in Space and Avionics Applications
abstract
This article presents a novel configuration scrubbing core, used for internal detection and correction of radiation-induced configuration single and multiple bit errors, without requiring external scrubbing. The proposed technique combines the benefits of fast radiation-induced fault detection with fast restoration of the device functionality and small area and power overheads. Experimental results demonstrate that the novel approach significantly improves the availability in hostile radiation environments of FPGA-based designs. When implemented using a Xilinx XC2V1000 Virtex-II device, the presented technique detects and corrects single bit upsets and double, triple and quadruple multi bit upsets, occupying just 1488 slices and dissipating less than 30 mW at a 50MHz running frequency.
Marco Lanuzza, Paolo Zicari, Fabio Frustaci, Stefania Perri, Pasquale Corsonello
ACM Trans. Reconfigurable Technol. Syst.4
2007 Design and Implementation of a 90nm Low bit-rate Image Compression Core
abstract
This paper presents a low-cost, high throughput discrete wavelet transform-based image compressor. The hardware solution proposed here exploits a modified set partitioning in hierarchical trees (SPIHT) algorithm and ensures that appropriate reconstructed image qualities can be achieved also for compression ratios over 100:1. Obtained results demonstrate that a maximum data rate of about 23 Mpixels/s can be sustained on a 64x64 size tile. In 90 nm technology, the required area is only 1.77 mm2. To obtain higher performance, multiples cores can be used in a parallel implementation.
Pasquale Corsonello, Stefania Perri, Giovanni Staino, Marco Lanuzza, Giuseppe Cocorullo
DSD2
2007 An efficient and optimized FPGA Feedback M-PSK Symbol Timing Recovery Architecture based on the Gardner Timing Error Detector
abstract
This paper presents an efficient and optimized FPGA implementation of a complete digital Symbol Timing Recovery (STR) architecture based on a digital PLL loop structure. Matlab modelling and then a complete hardware communication system test, reveal that the implemented STR circuit offers the best performances compared with the other implemented works present in literature. When implemented on a Xilinx Virtex-2P XC2VP7 FF672 FPGA chip the proposed STR circuit occupies just 138 slices, uses 2 embedded multipliers and reaches a clock frequency of 106 MHz; a symbol rate of 10 Msymbol/sec can be reached when 10 samples per symbol are employed. The obtained results are promising for its use in software defined radio system applications.
Emanuele Sciagura, Paolo Zicari, Stefania Perri, Pasquale Corsonello
DSD3
2006 An integrated countermeasure against differential power analysis for secure smart-cards
abstract
This paper presents a new hardware technique for the realization of secure smart-cards. The proposed strategy represents a valid countermeasure against non-invasive attacks, such as power analysis. It is based on a simple sub-circuit (Kocher et al., 1999) that can be easily integrated into the smart-card chip. It has been proven that the new technique decorrelates the power consumed by any digital circuit from the internally elaborated data, thus avoiding extraction of secret information from smart cards during the execution of their internal computations
Pasquale Corsonello, Stefania Perri, Martin Margala
ISCAS2
2006 Leakage energy reduction techniques in deep submicron cache memories: a comparative study
abstract
Static energy consumption due to subthreshold leakage current is one of the main concern in on-chip level-1 and level-2 cache. In the last few years several techniques have been proposed to limit the subthreshold current in a SRAM cell. Unfortunately, these techniques also increase the dynamic energy during the cell access operation, with respect to the conventional SRAM architecture. In this paper the actual energy saving offered by low leakage approaches is investigated, within the context of a microprocessor memory hierarchy, taking into account their dynamic energy overheads. Simulation based on UMC 0.18mum-1.8V and ST 90nm-1V process models have been performed. Results show that, for both the technologies, the leakage energy saving achieved by the analyzed techniques in the first cache level turns out to be inadequate, owing to the extra dynamic energy dissipation. Only in UL2 they assure a net energy saving due to the smaller number of accesses
Fabio Frustaci, Pasquale Corsonello, Stefania Perri, Giuseppe Cocorullo
ISCAS3
2006 Low bit rate image compression core for onboard space applications
abstract
This paper presents low-cost, purpose optimized discrete wavelet transform-based image compressors for future spacecrafts and microsatellites. The hardware solution proposed here exploits a modified set partitioning in hierarchical trees algorithm and ensures that appropriate reconstructed image qualities can be achieved also for compression ratios over 100:1. Several implementations are presented varying the parallelism level and the tile size. Obtained results demonstrate that, using a parallel implementation operating on a 64 /spl times/ 64 size tile, a maximum data rate of about 18 Mpixels/s can be sustained. In this case, only 4500 slices and 24 BlockRAMs of a XILINX Virtex II device are required.
Pasquale Corsonello, Stefania Perri, Giovanni Staino, Marco Lanuzza, Giuseppe Cocorullo
IEEE Trans. Circuits Syst. Video Technol.2
2006 Techniques for Leakage Energy Reduction in Deep Submicrometer Cache Memories
abstract
The techniques known in literature for the design of SRAM structures with low standby leakage typically exploit an additional operation mode, named the sleep mode or the standby mode. In this paper, existing low leakage SRAM structures are analyzed by several SPEC2000 benchmarks. As expected, the examined SRAM architectures have static power consumption lower than the conventional 6-T SRAM cell. However, the additional activities performed to enter and to exit the sleep mode also lead to higher dynamic energy. Our study demonstrates that, due to this, the overall energy consumption achieved by the known low-leakage techniques is greater than the conventional approach. In the second part of this paper, a novel low-leakage SRAM cell is presented. The proposed structure establishes when to enter and to exit the sleep mode, on the basis of the data stored in it, without introducing time and energy penalties with respect to the conventional 6-T cell. The new SRAM structure was realized using the UMC 0.18-mum, 1.8-V, and the ST 90-nm 1-V CMOS technologies. Tests performed with a set of SPEC2000 benchmarks have shown that the proposed approach is actually energy efficient
Fabio Frustaci, Pasquale Corsonello, Stefania Perri, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Low-Cost Fully Reconfigurable Data-Path for FPGA-Based Multimedia Processor
abstract
This paper describes novel data-path architecture for FPGA-based multimedia processors. The proposed circuit can adapt itself at run-time to different operations and data wordlengths avoiding time and power consuming reconfiguration. The new data-path can operate in SIMD fashion and guarantees high parallelism levels when operations on lower precisions are executed. It also supports IEEE-754 compliant single precision floating-point addition and multiplication. The proposed circuit has been characterized using VIRTEXII XILINX devices, but it can be efficiently used also in other FPGA families.
Marco Lanuzza, Stefania Perri, Martin Margala, Pasquale Corsonello
FPL2
2004 Variable precision arithmetic circuits for FPGA-based multimedia processors
abstract
This brief describes new efficient variable precision arithmetic circuits for field programmable gate array (FPGA)-based processors. The proposed circuits can adapt themselves to different data word lengths, avoiding time and power consuming reconfiguration. This is made possible thanks to the introduction of on purpose designed auxiliary logic, which enables the new circuits to operate in single instruction multiple data (SIMD) fashion and allows high parallelism levels to be guaranteed when operations on lower precisions are executed. The new SIMD structures have been designed to optimally exploit the resources of a widely used family of SRAM-based FPGAs, but their architectures can be easily adapted to any either SRAM-based or antifuse-based FPGA chips.
Stefania Perri, Pasquale Corsonello, Maria Antonia Iachino, Marco Lanuzza, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.1
2003 Variable Precision Multipliers for FPGA-Based Reconfigurable Computing Systems
Pasquale Corsonello, Stefania Perri, Maria Antonia Iachino, Giuseppe Cocorullo
FPL2
2003 A high-speed energy-efficient 64-bit reconfigurable binary adder
abstract
Datapaths for media signal processing are typically built using programmable computational elements such as adders and multipliers, which can be run-time reconfigured to operate on simple integers with 8, 16, or 32 bits of precision. In this brief, a new high-speed energy-efficient reconfigurable adder for media signal processing is presented. The proposed circuit is based on carry-propagation schemes and can be partitioned to perform one 64-, two 32-, four 16-, and eight 8-bit additions. When the Austria Mikro System (AMS) 0.35 /spl mu/m 2-poly 3-metal 3.3 V CMOS (CSD) process is used to produce layout, a worst propagation delay of about 4.9 ns and an average energy dissipation of about 181 /spl mu/W/MHz are obtained.
Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.1
2002 VLSI circuits for low-power high-speed asynchronous addition
abstract
This paper presents a new low-power high-speed fully static CMOS variable-time adder. The VLSI implementation proposed here is based on the statistical carry look-ahead addition technique. The new circuit takes advantage of an innovative way of using a composition of propagate signals and of appropriately designed overlapped execution modules to reduce average addition time, layout area, and power dissipation. A 56-bit adder designed as described here and realized using AMS 0.35-/spl mu/m CMOS standard cells at 3.3V supply voltage shows an average addition time of about 4.3 ns and a maximum power dissipation of only 50 mW at 200-MHz repetitive frequency using a silicon area of less than 0.23 mm/sup 2/.
Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.1
2000 Area-time-power tradeoff in cellular arrays VLSI implementations
abstract
Designing pipelined cellular arrays for arithmetical purposes, the choice of circuit design style is crucial. Usually, this choice is made by establishing an optimal area-time-power tradeoff. In order to achieve this result, analysis and simulations of the whole designed array have to be repeatedly performed for several design styles. This paper presents a methodology that allows the same result to be obtained avoiding time-consuming simulations of an entire array. The proposed technique is based on an appropriate partitioning of the arrays into small subcircuits. The features of the latter are analytically recomposed to evaluate performances and costs of an array of any size for various design approaches.
Pasquale Corsonello, Stefania Perri, G. Cororullo
IEEE Trans. Very Large Scale Integr. Syst.2