VLDB 2026 Research / reviewers in the wild / expert
Pasquale Corsonello
dblp:52/4718
· DBLP profile ↗
40ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-9528-1110ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI WorkloadsabstractEnergy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
DATE | 14 |
| 2025 | Multi-Partner Project: Architectures and Design Methodologies to Accelerate AI Workloads. The ICSC Flagship 2 ProjectabstractRecent pre-exascale and exascale supercomputers have driven the development of increasingly sophisticated AI models for diverse applications, including image recognition and classification, natural language processing, and generative AI. These applications require specialized hardware accelerators, to handle the heavy computational demands of AI algorithms in an energy-efficient manner. Today, AI accelerators are deployed across various systems, from low-power edge devices to large-scale servers, high-performance computing (HPC) infrastructures, and data centers. The primary objective of the ICSC Flagship 2 project, discussed in this paper, is to develop heterogeneous hardware platforms optimized to accelerate HPC and big data applications. Specifically, this paper provides an overview of the key challenges addressed and the achievements realized at the current intermediate stage of the ICSC Flagship 2 project focused on architectures, technologies, and design methodologies to design efficient hardware accelerators for AI workloads, such as deep learning (DL) and transformer models. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Stefania Perri, Fanny Spagnolo, Pasquale Corsonello, Sebastiano Fabio Schifano, Cristian Zambelli, Angelo Garofalo, Francesco Conti 0001, Luca Benini |
DATE | 7 |
| 2025 | C4TERO: Configurable Cascaded Carry Chains for High Reliability TERO PUFs on FPGAsabstractIn this paper we present a novel Transient Effect Ring Oscillator Physical Unclonable Function for FPGAs. It exploits in an original way the carry chain resources available in modern devices. The basic cell adopted in the proposed architecture can be runtime configured to implement different oscillation paths. This property enables the possibility to output more than one bit response per cell by choosing among the configurations those that exhibit the highest reliability. Such results are achieved by adopting a specific calibration process able to identify configurations of the cells showing the highest stability and the most uncorrelated responses. When implemented on several Series 7 Xilinx devices, no unstable bits were observed at 1 V and$25~^{\circ }$C. Under voltage variation in the manufacturer recommended ranges, a worst case bit error rate of 0.046% is achieved. The circuit designed as here described consists of 64 cells, produces 128 response bits and consumes just 535 look-up-tables and 256 carry chains. Fanny Spagnolo, Massimo Vatalaro, Stefania Perri, Felice Crupi, Pasquale Corsonello |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A Novel Compressive Sensing Method for Secure and Energy Efficient ECG Signal Transmission ApplicationsabstractThis paper introduces a novel Compressive Sensing (CS)-based cryptosystem tailored for the secure and efficient transmission of Electrocardiogram (ECG) signals in Internet of Medical Things (IoMT) environments. It leverages the inherent sparsity of ECG signals in the wavelet domain and ensures both data privacy and integrity during transmission through a simple but effective additional encryption stage. We have evaluated the performance of four distinct sensing matrices and three reconstruction algorithms across multiple wavelet families. Through extensive simulations using the MIT-BIH Arrhythmia Database we demonstrate that the Low-Density Parity-Check matrix combined with the L1 optimization algorithm achieves the highest Quality Score, with Compression Ratio up 50%. The proposed approach also shows an excellent ability in preserving important pathological features in presence of abnormal beats. The proposed CS encoder has been hardware implemented on low-resource microcontroller and FPGA devices. When realized on a Xilinx Artix 7 XC7A12 T FPGA, such a prototype allows real-time operations to be sustained running at 1 MHz clock frequency and dissipating only 0.8nJ per sample. Fanny Spagnolo, Bharat Lal, Pasquale Corsonello, Raffaele Gravina |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | KIT: Kernel Isotropic Transformation of Bilateral Filters for Image Denoising on FPGAabstractA Bilateral filter (BF) is commonly adopted as a pre-processing stage in several computer vision tasks because of its ability to denoise images. In contrast to the traditional image convolution that adopts a static kernel, a BF computes adaptive weights on-the-fly by applying exponentiation and division operations to the current pixel window. Prior works dealing with hardware acceleration of the BF rely on straightforward implementations that approximate the exponential function through look-up-tables (LUTs). This paper presents a new approximation technique to efficiently deploy a BF within real-time and low-energy intelligent systems based on FPGAs. The proposed strategy replaces the adaptive filter with its inexact isotropic version. This choice allows dropping a certain number of operations, thus resulting in enhanced speed and energy performances with respect to state-of-the-art hardware accelerators. When implemented on the AMD Xilinx Zynq XC7Z020 FPGA device, the proposed $5 \times 5$ BF design elaborates ∽237 Mega pixels per second and consumes at most 174 mW, with a Peak Signal-to-Noise Ratio (PSNR) degradation of just 0.55% at a noise standard deviation equal to 30. Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri |
FPL | 2 |
| 2024 | An explainable embedded neural system for on-board ship detection from optical satellite imageryabstractAutomatic ship detection from spaceborne systems such as satellites or aircrafts, raises considerable attention in sea surface monitoring because of the several applications in military and civilian field. In this context, processing satellite images on-board would reduce the latency time especially for emergency situations. In this paper, an hardware-oriented (HO) ship detection system based on a customized Convolutional Neural Network (CNN), here referred to as HO-ShipNet, is proposed and tested on a revised version of the “Ships in Satellite Imagery” (SSI) Kaggle dataset, reporting detection accuracy of up to 95%. Furthermore, the explainability of HO-ShipNet is investigated by means of explainable Artificial Intelligence (xAI) techniques (i.e., Local Interpretable Model-Agnostic Explanation (LIME) and Occlusion Sensitivuty Analysis (OSA)), in order to understand the reasoning behind the HO-ShipNet decisions by detecting the most important input features and consequently ensure the trustworthiness of the model itself. Finally, HO-ShipNet is also implemented on the heterogeneous Xilinx xc7z045ffg900-2 SoC Field Programmable Gate Array (FPGA) outperforming state-of-the-art FPGA-based accelerators dealing with high-resolution frames. The promising results encourage the potential deployment of the proposed system for on-board applications. Cosimo Ieracitano, Nadia Mammone, Fanny Spagnolo, Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Francesco Carlo Morabito |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Approximate bilateral filters for real-time and low-energy imaging applications on FPGAsabstractAbstract Bilateral filtering is an image processing technique commonly adopted as intermediate step of several computer vision tasks. Opposite to the conventional image filtering, which is based on convolving the input pixels with a static kernel, the bilateral filtering computes its weights on the fly according to the current pixel values and some tuning parameters. Such additional elaborations involve nonlinear weighted averaging operations, which make difficult the deployment of bilateral filtering within existing vision technologies based on real-time and low-energy hardware architectures. This paper presents a new approximation strategy that aims to improve the energy efficiency of circuits implementing the bilateral filtering function, while preserving their real-time performances and elaboration accuracy. In contrast to the state-of-the-art, the proposed technique allows the filtering action to be on the fly adapted to both the current pixel values and to the tuning parameters, thus avoiding any architectural modification or tables update. When hardware implemented within the Xilinx Zynq XC7Z020 FPGA device, a 5 × 5 filter based on the proposed method processes 237.6 Mega pixels per second and consumes just 0.92 nJ per pixel, thus improving the energy efficiency by up to 2.8 times over the competitors. The impact of the proposed approximation on three different imaging applications has been also evaluated. Experiments demonstrate reasonable accuracy penalties over the accurate counterparts. Fanny Spagnolo, Pasquale Corsonello, Fabio Frustaci, Stefania Perri |
J. Supercomput. | 2 |
| 2024 | Exploring the Usage of Fast Carry Chains to Implement Multistage Ring Oscillators on FPGAs: Design and CharacterizationabstractRing oscillators (ROs) serve as basic building blocks in a lot of application scenarios, where they must ensure high reliability, flexibility, and low-area/energy footprint. With the recent advances of the Internet-of-Things (IoT) technology, in particular, the necessity to endow interconnected devices with security facilities has increased as well. In this context, the efficient implementation of ROs on field-programmable gate arrays (FPGAs) is crucial, even though it hides some pitfalls. This article presents a new design strategy for multistage ROs relying on the carry chains (CCs) available into modern FPGA devices. Several configurations of ROs designed as proposed here have been characterized in terms of hardware costs, jitter, and temperature/voltage sensitivity. In all the evaluated cases, the proposed design allows to achieve predictable routing schemes through the automatic place and route (P&R), while reducing slice occupancy and energy consumption by up to 50% and 44%, respectively, in comparison with the traditional lookup table (LUT)-based ROs. When realized on a Artix-7 device, the basic version of the proposed oscillator realized using 33 inverting stages allows obtaining multiphase outputs oscillating at 29.7 MHz with a standard deviation less than 10 kHz. The analysis conducted also demonstrates the high flexibility of the novel circuits, such as the possibility to easily change their behavior depending on the target application requirements. As an example, by exploiting additional pass-through elements, the proposed scheme achieves a sensitivity of 49 kHz/°C that is more than 4 times higher than that shown by the corresponding traditional LUT-based competitor, thus making it more suitable for thermal monitoring applications. Fanny Spagnolo, Stefania Perri, Massimo Vatalaro, Fabio Frustaci, Felice Crupi, Pasquale Corsonello |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | ERMES: Efficient Racetrack Memory Emulation System based on FPGAabstractWith the scaling of CMOS technology almost over, non-volatile memories based on emerging technologies are gaining considerable popularity. Particularly, spintronic-based Racetrack memories (RTMs) exhibit unprecedented storage capacity, as well as reduced energy per operation and high write endurance, which make them promising candidates to revolutionize the architecture of memory sub-systems. However, since RTM exploits shifting of magnetic domains to align the required data with the access port, its read/write latency is not constant. Due to this behaviour, several performance optimizations related to the target application may be introduced either on memory architecture or data placement or both. To this purpose, specific tools able to emulate the timing characteristics of RTMs are highly desired. Unfortunately, existing software-based simulators show poor flexibility and run-time. To address such limitations, this paper presents a new emulation system for RTMs based on heterogeneous FPGA-CPU Systems-on-Chips (SoCs). Thanks to its high flexibility, the proposed emulator can be easily configured to evaluate different memory architectures. In addition, the CPU can be used to stimulate the RTM architecture under test with appropriate benchmarks, thus providing a fast self-contained evaluation environment. As case study, ERMES has been implemented within the Xilinx Zynq Ultrascale XCUZ9EG SoC to evaluate performances of several memory configurations when running benchmark applications from the MiBench suite, experiencing a speed-up higher than × 146 over software-based simulators. Fanny Spagnolo, Salim Ullah, Pasquale Corsonello, Akash Kumar 0001 |
FPL | 3 |
| 2020 | An Efficient Convolution Engine based on the À-trous Spatial Pyramid PoolingabstractThis paper presents an efficient hardware architecture able to perform 2D dilated convolutions and suitable for the integration within modern heterogeneous embedded systems targeting semantic image segmentation. The proposed design supports multiple dilation rates. Moreover, it uses limited amounts of resources even when large convolution windows are processed. As a case study, the novel circuit has been integrated within a Xilinx Zynq-7000 FPSoC device to accelerate a state-of-the-art CNN model for medical images segmentation. Obtained results demonstrate that higher computational capabilities, reduced resources utilization and lower power consumption are achieved with respect to the competitors existing in literature. Cristian Sestito, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri |
ASAP | 3 |
| 2020 | Design of a real-time face detection architecture for heterogeneous systems-on-chips
Fanny Spagnolo, Stefania Perri, Pasquale Corsonello |
Integr. | 3 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2019 | Energy-Quality Scalable Adders Based on Nonzeroing Bit TruncationabstractApproximate addition is a technique to trade off energy consumption and output quality in error-tolerant applications. In prior art, bit truncation has been explored as a lever to dynamically trade off energy and quality. In this brief, an innovative bit truncation strategy is proposed to achieve more graceful quality degradation compared to state-of-the-art truncation schemes. This translates into energy reduction at a given quality target. When applied to a ripple-carry adder, the proposed bit truncation approach improves quality by up to 8.5 dB in terms of peak signal-to-noise ratio, compared to traditional bit truncation. As a case study, the proposed approach was applied to a discrete cosine transform engine. In comparison with prior art, the proposed approach reduces energy by 20%, at insignificant delay and silicon area overhead. Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Design of Real-Time FPGA-based Embedded System for Stereo VisionabstractThis paper describes a novel heterogeneous SoC FPGA-based embedded system for stereo vision. Two complete implementations are presented and characterized. In both designs the auxiliary computations, such as the image rectification and the disparity map refinement, are performed by the custom hardware module purpose-designed to compute disparity maps, thus achieving very high speeds. The software routine run by the on-chip general-purpose processor is used to control configuration and communication. Obtained results show that, in comparison with several existing hardware designs, the proposed system reaches higher performances, competitive accuracies, lower complexity and higher flexibility. Stefania Perri, Fabio Frustaci, Fanny Spagnolo, Pasquale Corsonello |
ISCAS | 4 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2015 | Exploring well configurations for voltage level converter design in 28 nm UTBB FDSOI technologyabstractVoltage level converters are critical components in multi supply ultra-low voltage designs, especially when signals need to be converted from the sub-threshold to the above-threshold domain. In these designs, advanced technology processes, such as the Ultra-Thin Body and Buried oxide (UTBB) Fully-Depleted SOI (FDSOI), are greatly desired since they intrinsically allow controlling the Drain Induced Barrier Lowering effect (DIBL) and the Gate Induced Drain Leakage (GIDL), in addition to the reduction of the effects of process variations. Moreover, these technologies provide a group of architectural and device-level techniques for threshold voltage adjustment that can be efficiently adopted to combine high performances and low energy consumption. However, specific design strategies should be applied to efficiently exploit all these potentialities. This paper investigates how the physical design of level converters can benefit from the synergistic adoption of the knobs available in the UTBB FDSOI technology (poly biasing, flip-well, single-well, back biasing). In particular, three mixed single well configurations have been implemented and analyzed. This research work demonstrates that the specific selected approach allows decreasing the energy per cycle consumption, the leakage current and the delay by up to 35.3%, 70.4%, and 6.2%, respectively, with respect to the basic conventional design strategy. Furthermore, statistical analysis confirmed that these advantages are maintained for a wide range of process variations, also improving the functional yield and the minimum input voltage causing the level converter failure. Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
ICCD | 1 |
| 2015 | Novel Varactor-Loaded Phasing Line for Reflectarray Unit Cell with Large Reconfigurability Frequency Range
Sandra Costanzo, Francesca Venneri, Antonio Raffo, Giuseppe Di Massa, Pasquale Corsonello |
WorldCIST (2) | 5 |
| 2015 | Power supply noise in accurate delay model for the sub-threshold domain
Pasquale Corsonello, Fabio Frustaci, Stefania Perri |
Integr. | 1 |
| 2015 | Low-Leakage SRAM Wordline Drivers for the 28-nm UTBB FDSOI TechnologyabstractThis brief deals with a new design of low-power SRAM wordline decoder in the 28-nm ultrathin body and buried oxide (UTBB) fully depleted silicon-on-insulator (FDSOI) technology. The proposed approach synergistically adopts the poly biasing technique in conjunction with single-well/flip-well configurations and body biasing to opportunely tune the threshold voltage of the devices in the standby and active mode. A tuning methodology is described to optimize the static energy consumption. Post-layout simulations, done at power supply voltages ranging between 1 V and 0.5 V, have shown that, in comparison with the state-of-the-art techniques based on the same UTBB FDSOI technology, the proposed design achieves a maximum leakage up to 85% lower without paying significant delay penalties. Pasquale Corsonello, Fabio Frustaci, Stefania Perri |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Fast and Wide Range Voltage Conversion in Multisupply Voltage DesignsabstractMultisupply voltage design technique is widely used in modern system-on-chips to tradeoff energy and speed. Level shifters (LSs) allow different voltage domains to be interfaced. In this brief, a new LS is presented for fast and wide range voltage conversion. Because of a novel architecture combined with the use of multithreshold CMOS technique, the proposed circuit guarantees robust voltage shifting from the deep subthreshold to the above-threshold domain while exhibiting fast response and low energy consumption. When implemented in a 90-nm technology node, considering process-voltage-temperature variations, the proposed design reliably converts 100-mV input signals into 1 V output signals. Post-layout simulation results demonstrate that the new LS shows a propagation delay of 16.6 ns, a static power dissipation of 8.7 nW and a total energy per transition of only 77 fJ for a 0.2 V 1-MHz input pulse. Marco Lanuzza, Pasquale Corsonello, Stefania Perri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Area-Delay Efficient Binary Adders in QCAabstractAs transistors decrease in size more and more of them can be accommodated in a single die, thus increasing chip computational capabilities. However, transistors cannot get much smaller than their current size. The quantum-dot cellular automata (QCA) approach represents one of the possible solutions in overcoming this physical limit, even though the design of logic modules in QCA is not always straightforward. In this brief, we propose a new adder that outperforms all state-of-the-art competitors and achieves the best area-delay tradeoff. The above advantages are obtained by using an overall area similar to the cheaper designs known in literature. The 64-bit version of the novel adder spans over 18.72 μ2of active area and shows a delay of only nine clock cycles, that is just 36 clock phases. Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Adaptive Census Transform: A novel hardware-oriented stereovision algorithm
Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo |
Comput. Vis. Image Underst. | 2 |
| 2011 | Tapered-VTH CMOS buffer design for improved energy efficiency in deep nanometer technologyabstractIn this paper, the novel "tapered-Vth" approach to design energy-efficient CMOS buffers is introduced. In this approach, the substantial energy consumption due to leakage is reduced by tapering the threshold voltage throughout the buffer stages, other than tapering the transistor size. More specifically, the threshold voltage is progressively reduced when going from the last to the first stage. This enables a considerable leakage reduction in the last stages (which contribute most to the overall leakage) at the price of a higher delay. The resulting delay penalty is then compensated by reducing the transistor threshold voltage in the first stages, with an insignificant leakage increase (they contribute very little to the overall buffer leakage). Simulation results based on a commercial 45-nm 1-V CMOS technology show that the proposed "tapered-VTH" approach can considerably improve the energy efficiency of CMOS buffers over the entire spectrum of possible energy-delay tradeoffs, from high speed to low power. Fabio Frustaci, Pasquale Corsonello, Massimo Alioto |
ISCAS | 2 |
| 2010 | A new low-power high-speed single-clock-cycle binary comparatorabstractThis paper presents a new ultra-low power high-speed single-clock-cycle binary comparator. It is based on a novel parallel-prefix algorithm which drastically reduces the switching activity of the internal nodes of the circuit. When implemented by using the ST 90nm-1V technology, the proposed 64-bit comparator exhibits an energy dissipation of only 0.77μW/MHz and a delay of 258ps. With respect to a recently published low-power high-speed parallel-prefix adder, the proposed design shows an energy dissipation reduction of 23% and a speed improvement of 7%. Fabio Frustaci, Stefania Perri, Marco Lanuzza, Pasquale Corsonello |
ISCAS | 4 |
| 2010 | Exploiting Self-Reconfiguration Capability to Improve SRAM-based FPGA Robustness in Space and Avionics ApplicationsabstractThis article presents a novel configuration scrubbing core, used for internal detection and correction of radiation-induced configuration single and multiple bit errors, without requiring external scrubbing. The proposed technique combines the benefits of fast radiation-induced fault detection with fast restoration of the device functionality and small area and power overheads. Experimental results demonstrate that the novel approach significantly improves the availability in hostile radiation environments of FPGA-based designs. When implemented using a Xilinx XC2V1000 Virtex-II device, the presented technique detects and corrects single bit upsets and double, triple and quadruple multi bit upsets, occupying just 1488 slices and dissipating less than 30 mW at a 50MHz running frequency. Marco Lanuzza, Paolo Zicari, Fabio Frustaci, Stefania Perri, Pasquale Corsonello |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2009 | New performance/power/area efficient, reliable full adder designabstractArithmetic circuits have always played one of the most important roles in the designs of processors, FPGAs, and the rapidly evolving domain of media processing architectures. The full adder cell forms the basic building block of majority of these arithmetic circuits. In this paper we describe a hybrid pseudo static full adder cell designed using Data Driven Dynamic Logic. Simulation results show the adder to out perform its competitors, both static as well as dynamic topologies in terms of performance, while maintaining relatively similar area and power characteristics. This paper presents a complete characterization of the popular adder cells in terms of delay, area, power, noise margin and reliability analysis for both super threshold and sub threshold operating regimes. Sohan Purohit, Martin Margala, Marco Lanuzza, Pasquale Corsonello |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Design and Implementation of a 90nm Low bit-rate Image Compression CoreabstractThis paper presents a low-cost, high throughput discrete wavelet transform-based image compressor. The hardware solution proposed here exploits a modified set partitioning in hierarchical trees (SPIHT) algorithm and ensures that appropriate reconstructed image qualities can be achieved also for compression ratios over 100:1. Obtained results demonstrate that a maximum data rate of about 23 Mpixels/s can be sustained on a 64x64 size tile. In 90 nm technology, the required area is only 1.77 mm2. To obtain higher performance, multiples cores can be used in a parallel implementation. Pasquale Corsonello, Stefania Perri, Giovanni Staino, Marco Lanuzza, Giuseppe Cocorullo |
DSD | 1 |
| 2007 | An efficient and optimized FPGA Feedback M-PSK Symbol Timing Recovery Architecture based on the Gardner Timing Error DetectorabstractThis paper presents an efficient and optimized FPGA implementation of a complete digital Symbol Timing Recovery (STR) architecture based on a digital PLL loop structure. Matlab modelling and then a complete hardware communication system test, reveal that the implemented STR circuit offers the best performances compared with the other implemented works present in literature. When implemented on a Xilinx Virtex-2P XC2VP7 FF672 FPGA chip the proposed STR circuit occupies just 138 slices, uses 2 embedded multipliers and reaches a clock frequency of 106 MHz; a symbol rate of 10 Msymbol/sec can be reached when 10 samples per symbol are employed. The obtained results are promising for its use in software defined radio system applications. Emanuele Sciagura, Paolo Zicari, Stefania Perri, Pasquale Corsonello |
DSD | 4 |
| 2006 | An integrated countermeasure against differential power analysis for secure smart-cardsabstractThis paper presents a new hardware technique for the realization of secure smart-cards. The proposed strategy represents a valid countermeasure against non-invasive attacks, such as power analysis. It is based on a simple sub-circuit (Kocher et al., 1999) that can be easily integrated into the smart-card chip. It has been proven that the new technique decorrelates the power consumed by any digital circuit from the internally elaborated data, thus avoiding extraction of secret information from smart cards during the execution of their internal computations Pasquale Corsonello, Stefania Perri, Martin Margala |
ISCAS | 1 |
| 2006 | Leakage energy reduction techniques in deep submicron cache memories: a comparative studyabstractStatic energy consumption due to subthreshold leakage current is one of the main concern in on-chip level-1 and level-2 cache. In the last few years several techniques have been proposed to limit the subthreshold current in a SRAM cell. Unfortunately, these techniques also increase the dynamic energy during the cell access operation, with respect to the conventional SRAM architecture. In this paper the actual energy saving offered by low leakage approaches is investigated, within the context of a microprocessor memory hierarchy, taking into account their dynamic energy overheads. Simulation based on UMC 0.18mum-1.8V and ST 90nm-1V process models have been performed. Results show that, for both the technologies, the leakage energy saving achieved by the analyzed techniques in the first cache level turns out to be inadequate, owing to the extra dynamic energy dissipation. Only in UL2 they assure a net energy saving due to the smaller number of accesses Fabio Frustaci, Pasquale Corsonello, Stefania Perri, Giuseppe Cocorullo |
ISCAS | 2 |
| 2006 | Low bit rate image compression core for onboard space applicationsabstractThis paper presents low-cost, purpose optimized discrete wavelet transform-based image compressors for future spacecrafts and microsatellites. The hardware solution proposed here exploits a modified set partitioning in hierarchical trees algorithm and ensures that appropriate reconstructed image qualities can be achieved also for compression ratios over 100:1. Several implementations are presented varying the parallelism level and the tile size. Obtained results demonstrate that, using a parallel implementation operating on a 64 /spl times/ 64 size tile, a maximum data rate of about 18 Mpixels/s can be sustained. In this case, only 4500 slices and 24 BlockRAMs of a XILINX Virtex II device are required. Pasquale Corsonello, Stefania Perri, Giovanni Staino, Marco Lanuzza, Giuseppe Cocorullo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Techniques for Leakage Energy Reduction in Deep Submicrometer Cache MemoriesabstractThe techniques known in literature for the design of SRAM structures with low standby leakage typically exploit an additional operation mode, named the sleep mode or the standby mode. In this paper, existing low leakage SRAM structures are analyzed by several SPEC2000 benchmarks. As expected, the examined SRAM architectures have static power consumption lower than the conventional 6-T SRAM cell. However, the additional activities performed to enter and to exit the sleep mode also lead to higher dynamic energy. Our study demonstrates that, due to this, the overall energy consumption achieved by the known low-leakage techniques is greater than the conventional approach. In the second part of this paper, a novel low-leakage SRAM cell is presented. The proposed structure establishes when to enter and to exit the sleep mode, on the basis of the data stored in it, without introducing time and energy penalties with respect to the conventional 6-T cell. The new SRAM structure was realized using the UMC 0.18-mum, 1.8-V, and the ST 90-nm 1-V CMOS technologies. Tests performed with a set of SPEC2000 benchmarks have shown that the proposed approach is actually energy efficient Fabio Frustaci, Pasquale Corsonello, Stefania Perri, Giuseppe Cocorullo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Low-Cost Fully Reconfigurable Data-Path for FPGA-Based Multimedia ProcessorabstractThis paper describes novel data-path architecture for FPGA-based multimedia processors. The proposed circuit can adapt itself at run-time to different operations and data wordlengths avoiding time and power consuming reconfiguration. The new data-path can operate in SIMD fashion and guarantees high parallelism levels when operations on lower precisions are executed. It also supports IEEE-754 compliant single precision floating-point addition and multiplication. The proposed circuit has been characterized using VIRTEXII XILINX devices, but it can be efficiently used also in other FPGA families. Marco Lanuzza, Stefania Perri, Martin Margala, Pasquale Corsonello |
FPL | 4 |
| 2005 | Cost-effective low-power processor-in-memory-based reconfigurable datapath for multimedia applicationsabstractMultimedia applications have become a dominant computing workload for computer systems as well as for wireless-based devices. Due to their repetitive computing and memory intensive nature, they can take effective advantage from Processor-In-Memory (PIM) technology. In this paper, a new low-power PIM-based 32-bit reconfigurable datapath optimized for multimedia applications is presented. The new circuit efficiently performs parallel arithmetic operations on either 8-, 16-, or 32-bit integer data or on 32-bit single precision floating-point data. As a result, high flexibility is provided at a very low hardware cost. When implemented using the UMC 0.18 μm 1.8 V CMOS technology, the proposed datapath exhibits a 285 MHz running frequency, dissipates just 0.12 mW/MHz and occupies a silicon area of only 107,323 μm2. When performing 2D-DCT, proposed architecture consumes 74% less power and is 28% more power efficient compared to top-of-the-line commercial TI DSP Marco Lanuzza, Martin Margala, Pasquale Corsonello |
ISLPED | 3 |
| 2004 | FPGA implementation of Bayesian neural networks for a stand-alone predictor of pollutants concentration in the airabstractWe exploit the potentials of Bayesian neural networks combined with the advantages of a VLSI implementation in order to design a stand-alone predictor system of air pollutants time series. The area under study is Villa San Giovanni, a small town located in front of the Messina Strait (Italy), whose harbor represents the main link to reach Sicily island by cars and trucks. Neural networks are powerful tools to predict air pollutants time series, but almost always they run by software programs on PC or workstations, which make difficult their use when are present constraints such as portability, low power dissipation, limited physical size. In this cases, SRAM based field programmable gate arrays (FPGAs) represent a suitable platform to realize these models, since their reprogrammability offers the possibility to rapidly change the parameters of the network if a new training is needed. The achieved results have highlighted the efficient design of the hardware network, obtained also using a new circuit to compute the activation function of the neurons. Salvatore Marra, Francesco Carlo Morabito, Pasquale Corsonello, Mario Versaci |
IJCNN | 3 |
| 2004 | Variable precision arithmetic circuits for FPGA-based multimedia processorsabstractThis brief describes new efficient variable precision arithmetic circuits for field programmable gate array (FPGA)-based processors. The proposed circuits can adapt themselves to different data word lengths, avoiding time and power consuming reconfiguration. This is made possible thanks to the introduction of on purpose designed auxiliary logic, which enables the new circuits to operate in single instruction multiple data (SIMD) fashion and allows high parallelism levels to be guaranteed when operations on lower precisions are executed. The new SIMD structures have been designed to optimally exploit the resources of a widely used family of SRAM-based FPGAs, but their architectures can be easily adapted to any either SRAM-based or antifuse-based FPGA chips. Stefania Perri, Pasquale Corsonello, Maria Antonia Iachino, Marco Lanuzza, Giuseppe Cocorullo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Variable Precision Multipliers for FPGA-Based Reconfigurable Computing Systems
Pasquale Corsonello, Stefania Perri, Maria Antonia Iachino, Giuseppe Cocorullo |
FPL | 1 |
| 2003 | A high-speed energy-efficient 64-bit reconfigurable binary adderabstractDatapaths for media signal processing are typically built using programmable computational elements such as adders and multipliers, which can be run-time reconfigured to operate on simple integers with 8, 16, or 32 bits of precision. In this brief, a new high-speed energy-efficient reconfigurable adder for media signal processing is presented. The proposed circuit is based on carry-propagation schemes and can be partitioned to perform one 64-, two 32-, four 16-, and eight 8-bit additions. When the Austria Mikro System (AMS) 0.35 /spl mu/m 2-poly 3-metal 3.3 V CMOS (CSD) process is used to produce layout, a worst propagation delay of about 4.9 ns and an average energy dissipation of about 181 /spl mu/W/MHz are obtained. Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | VLSI circuits for low-power high-speed asynchronous additionabstractThis paper presents a new low-power high-speed fully static CMOS variable-time adder. The VLSI implementation proposed here is based on the statistical carry look-ahead addition technique. The new circuit takes advantage of an innovative way of using a composition of propagate signals and of appropriately designed overlapped execution modules to reduce average addition time, layout area, and power dissipation. A 56-bit adder designed as described here and realized using AMS 0.35-/spl mu/m CMOS standard cells at 3.3V supply voltage shows an average addition time of about 4.3 ns and a maximum power dissipation of only 50 mW at 200-MHz repetitive frequency using a silicon area of less than 0.23 mm/sup 2/. Stefania Perri, Pasquale Corsonello, Giuseppe Cocorullo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Area-time-power tradeoff in cellular arrays VLSI implementationsabstractDesigning pipelined cellular arrays for arithmetical purposes, the choice of circuit design style is crucial. Usually, this choice is made by establishing an optimal area-time-power tradeoff. In order to achieve this result, analysis and simulations of the whole designed array have to be repeatedly performed for several design styles. This paper presents a methodology that allows the same result to be obtained avoiding time-consuming simulations of an entire array. The proposed technique is based on an appropriate partitioning of the arrays into small subcircuits. The features of the latter are analytically recomposed to evaluate performances and costs of an array of any size for various design approaches. Pasquale Corsonello, Stefania Perri, G. Cororullo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |