Marco Lanuzza

dblp:03/5549 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-6480-9218ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 4 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SPARCAM: Sparse matrix multiplication accelerator using multi-port dynamic CAM
abstract
Sparse General matrix multiplication (SpGEMM) is a fundamental kernel in many scientific and engineering fields, including Artificial Intelligence (AI). However, its intrinsic computation complexity presents substantial challenges, making efficient hardware implementation particularly difficult. This paper proposes SPARCAM, a novel SpGEMM accelerator, developed and optimized for very energy-efficient AI edge applications. SPARCAM is designed using low-power dense Gain Cell embedded DRAM (GC-eDRAM) technology, a processing near memory paradigm, and a modified outer product matrix multiplication algorithm. Despite its quite limited peak theoretical performance, SPARCAM achieves very high energy efficiency due to its low-power architecture and almost 100% utilization of its computing resources. Designed in a commercial 28 nm FDSOI technology, SPARCAM achieves 13 . 9 × speedup over a high-performance embedded CPU when processing large-scale sparse matrices. When multiplying limited-size sparse matrices, SPARCAM obtains 193 × speedup over high-performance GPU. SPARCAM reaches about 4.3 orders-of-magnitude, on average, higher energy benefits, and 1892 × , 181 × , 2 × , and 3471 × , higher energy efficiency (over CPU) compared with state-of-the-art SpGEMM accelerators SpArch, OuterSPACE, MatRaptor, and high-performance GPU, respectively.
Esteban Garzón, Benjamin Zambrano, David Sheinenzon, Marco Lanuzza, Adam Teman, Leonid Yavits
J. Syst. Archit.4
2026 NV-PCAM: Non-Volatile Precharge-Free Content-Addressable Memory
abstract
Content-addressable memories (CAMs) are a class of associative memories known for their capability to perform massively parallel comparisons between an input query pattern and the entire memory content. In the past decade, the increasing demand for high-performance and energy-efficient computing systems has generated significant interest in non-volatile CAMs (NV-CAMs) based on emerging non-volatile memory devices. In this work, we propose a novel non-volatile, precharge-free CAM (NV-PCAM) scheme based on double-barrier magnetic tunnel junctions (DMTJs). When compared to its counterparts, NV-PCAM presents competitive figures of merit in terms of area, speed, and energy efficiency, while also ensuring low search error rates. We also provide a complete class of voltage-divider-based NV-CAM cells for benchmark comparison. All schemes are designed and laid out using a 65 nm process and evaluated under Monte Carlo and process-voltage-temperature (PVT) simulations. Through Monte Carlo simulations, the proposed NV-PCAM demonstrates up to 81% and 85% lower search energy than NV-NOR and NV-NAND, respectively, as well as a 61% and 16% improvement in terms of search delay with a compact cell area footprint.
Oliver Caisaluisa, Esteban Garzón, Eduardo Holguín, Marco Lanuzza, Luis-Miguel Procel
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Non-Volatile Content-Addressable Memory For Energy-Efficient & High-Performance Search And Update Operations
abstract
This work presents a non-volatile content-addressable memory (NV-CAM) based on double-barrier magnetic tunnel junction technology (DMTJ). Unlike state-of-the-art NV-CAM designs that present low-performance updates, our NV-CAM allows energy-efficient, high-performance search and update operations. This makes it well-suited for applications requiring a high frequency of searches/updates, such as associative processors. The NV-CAM hybrid CMOS/DMTJ was designed using a commercial 65nm CMOS technology and a Verilog-A-based DMTJ compact model. The NV-CAM evaluation was carried out by employing Monte Carlo simulations while accounting for process variations. Simulation results show that our NV-CAM presents competitive figures of merit compared to state-of-the-art design. Our NV-CAM presents energy-efficient operations and reduces the update and search delay by about 71% and 75%, respectively, compared to other NV-CAMs.
Alessandro Bedoya, Benjamin Zambrano, Ramiro Taco, Luis-Miguel Procel, Marco Lanuzza, Esteban Garzón
ISCAS5
2025 Low Matchline Voltage Swing Content-Addressable Memory Cell
abstract
Content-addressable memory (CAM) is a specialized memory architecture designed for fast data searches, allowing a one-clock-cycle comparison between the search input and the entire memory content. In this work, a low matchline voltage swing CAM is proposed to reduce the search power consumption while maintaining high-speed search operations. Low voltage swing in the matchline is enabled by introducing extra circuitry in the conventional CAM cell. By means of comprehensive Monte Carlo and post-layout simulations using a commercial 65nm node, we show that the proposed CAM cell design allows for robustness against process, voltage, and temperature variations without the need for dedicated matchline sense schemes. Compared to conventional precharge high NOR-type CAM, the proposed design achieves 42% higher speed and 29.1% less energy consumption. Post-layout results demonstrate that the proposed CAM operates reliably at 0.6V, maintaining performant and reliable search operations across a wide temperature range.
Cristhopher Mosquera, Ramiro Taco, Benjamin Zambrano, Luis-Miguel Procel, Esteban Garzón, Marco Lanuzza
ISCAS6
2025 Towards Low-Power High-Performance Content-Addressable Memory: A Robust Precharge-Free Approach
abstract
Low-Power high-performance content-addressable memories (CAMs) are important components in modern computing systems. In this work, we present a robust CAM that overcomes the power and performance limitations of conventional precharge-based CAMs. The proposed static transmission gate-based (STAT-TG) CAM design achieves low-power operation comparable to NAND CAMs while maintaining search speeds rivaling those of NOR CAMs. The STAT-TG CAM was designed using a 65nm CMOS technology and comprehensively evaluated under extensive Monte Carlo simulations. Compared to conventional CAMs, the STAT-TG CAM is 14% faster than NAND CAM, while consuming only 25% of the energy per operation relative to NOR CAM. This makes STAT-TG CAM a promising solution for high-performance yet energy-efficient applications.
Ramiro Taco, Esteban Garzón, Adam Teman, Leonid Yavits, Marco Lanuzza
ISCAS5
2025 A Multi-Bit PUF Architecture Using a 2T Sub-Threshold Voltage Divider
abstract
In this paper, a highly reliable multi-bit physically unclonable function (PUF) is proposed. The solution relies on an already tested two-transistor (2T) sub-threshold voltage divider as core circuit along with a multi-bit architecture able to carry out two highly stable bits from three bits generated by a proper entropy quantization. Twenty measured samples of the bitcell core were used to fit a customized Verilog-A model, which was then imported into Cadence Virtuoso environment for the architecture-level analysis. The proposed solution was tested across Monte Carlo simulations at both golden key (GK) and different environmental conditions, while also including the effect of noise. Simulation results prove the effectiveness in generating two highly stable bits for each cell after spatial majority voting and best stability selection. Indeed, no instability was observed in the 0-50 °C temperature range for the two output bits.
Massimo Vatalaro, Raffaele De Rose, Vincenzo Maccaronio, Marco Lanuzza, Felice Crupi
ISCAS4
2025 Highly Stable PUFs Based on Stacked Voltage Divider for Near-Zero BER Native Sensitivity to Voltage Variations
abstract
This paper explores a class of highly stable static monostable physically unclonable functions (PUFs) based on stacked sub-threshold voltage dividers between two nominally identical sub-circuits as bitcell core block. More specifically, compared to our previous works where two-transistor (2T) and four-transistor (4T) voltage divider based PUFs were presented and analyzed, here we propose two novel topological variants based on six-transistor (6T) and eight-transistor (8T) solutions which arise from adopting a proper reverse gate-biasing strategy within the stack with the aim of improving the resilience to on-chip noise and voltage variations, while keeping the area overhead low. These novel solutions, along with those already proposed, were tested in 180-nm CMOS technology. Raw measurements show a nominal (at 1.8 V and 25°C) bit error rate (BER) of 0.15% and 0.08% for the 6T- and 8T-based solutions, respectively, along with a BER variation of 0.016% and 0.002% per 0.1 V. With the implementation of a simple masking technique based on measurements at low supply voltage ($V_{DD} =0.3$V at 25 °C) along with a temporal majority voting (TMV) scheme, a BER of 0.006% and lower than$9.77\times 10^{-5}$%, which is the minimum observable BER for the adopted statistical set, was observed for the 6T-, and 8T-core based implementations, respectively, with a corresponding masking ratio of 8.71% and 7.59%. This is achieved with an area per bit of 5,$174F^{2}$(6T solution) and 6,$994F^{2}$(8T solution).
Massimo Vatalaro, Raffaele De Rose, Vincenzo Maccaronio, Marco Lanuzza, Felice Crupi
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 PUF-Based Authentication-Oriented Architecture for Identification Tags
abstract
Smart tags are compact electronic devices affixed to or embedded into objects to facilitate identification, monitoring, and data exchange. Consequently, secure authentication of these tags is a crucial issue, as objects must reliably verify their identity before sharing sensitive information with other entities. The application of Physical Unclonable Functions (PUF) as a device's “digital fingerprint” has attracted significant attention, yet existing PUF-based authentication methods exhibit security vulnerabilities, either due to the authentication protocol itself or the limited reliability of the PUF technology used. Moreover, there has been a considerable focus on the software aspect, often overlooking the critical role of hardware design, which can become a target for attacks aimed at compromising the device's identity or act as a hindrance in the manufacturing process. In light of these points, this paper introduces an identification tag architecture that leverages PUF technology, focusing on authentication. This architecture features a straightforward but efficient authentication protocol, underpinned by a new and highly stable PUF model. The overall architecture encompasses particular hardware implementation aspects that significantly simplify the tag's enrollment phase and minimize vulnerabilities to attacks. The paper also describes a prototype of this identification tag and provide detailed insights into its application.
Antonino Rullo, Carmelo Felicetti, Massimo Vatalaro, Raffaele De Rose, Marco Lanuzza, Felice Crupi, Domenico Saccà
IEEE Trans. Dependable Secur. Comput.5
2024 Designing Precharge-Free Energy-Efficient Content-Addressable Memories
abstract
Content-addressable memory (CAM) is a specialized type of memory that facilitates massively parallel comparison of a search pattern against its entire content. State-of-the-art (SOTA) CAM solutions are either fast but power-hungry (NOR CAM) or slow while consuming less power (nand CAM). These limitations stem from the dynamic precharge operation, leading to excessive power consumption in NOR CAMs and charge-sharing issues in NAND CAMs. In this work, we propose a precharge-free CAM (PCAM) class for energy-efficient applications. By avoiding precharge operation, PCAM consumes less energy than a NAND CAM, while achieving search speed comparable to a NOR CAM. PCAM was designed using a 65-nm CMOS technology and comprehensively evaluated under extensive Monte Carlo (MC) simulations while taking into account layout parasitics. When benchmarked against conventional NAND CAM, PCAM demonstrates improved search run time (reduced by more than 30%) and 15% less search energy. Moreover, PCAM can cut energy consumption by more than 75% when compared to conventional NOR CAM. We further extend our analysis to the application level, functionally evaluating the CAM designs as a fully associative cache using a CPU simulator running various benchmark workloads. This analysis confirms that PCAMs represent an optimal energy-performance design choice for associative memories and their broad spectrum of applications.
Ramiro Taco, Esteban Garzón, Robert Hanhan, Adam Teman, Leonid Yavits, Marco Lanuzza
IEEE Trans. Very Large Scale Integr. Syst.6
2022 EDAM: edit distance tolerant approximate matching content addressable memory
abstract
We propose a novel edit distance-tolerant content addressable memory (EDAM) for energy-efficient approximate search applications. Unlike state-of-the-art approximate search solutions that tolerate certain Hamming distance between the query pattern and the stored data, EDAM tolerates edit distance, which makes it especially efficient in applications such as text processing and genome analysis. EDAM was designed using a commercial 65 nm 1.2 V CMOS technology and evaluated through extensive Monte Carlo simulations, while considering different process corners. Simulation results show that EDAM can achieve robust approximate search operation with a wide range of edit distance threshold levels. EDAM is functionally evaluated as a pathogen DNA detection and classification accelerator. EDAM achieves up to 1.7× higher F1 score for high-quality DNA reads and up to 19.55× higher F1 score for DNA reads with 15% error rate, compared to state-of-the-art DNA classification tool Kraken2. Simulated at 667 MHz, EDAM provides 1, 214× average speedup over Kraken2. This makes EDAM suitable for hardware acceleration of genomic surveillance of outbreaks, such as the ongoing Covid-19 pandemic.
Robert Hanhan, Esteban Garzón, Zuher Jahshan, Adam Teman, Marco Lanuzza, Leonid Yavits
ISCA5
2022 A RISC-V-based Research Platform for Rapid Design Cycle
abstract
This work proposes a novel platform for bringing a project from the concept to the tapeout stage in a short amount of time. An open-source and extendable RISC-V architecture is exploited to build a small area footprint core. This leads the research platform to be flexible in terms of design integration, while also allowing fast design cycles of research chips.
Esteban Garzón, Roman Golman, Odem Harel, Tzachi Noy, Yehuda Kra, Asaf Pollock, Slava Yuzhaninov, Yonatan Shoshan, Yehuda Rudin, Yoav Weizman, Marco Lanuzza, Adam Teman
ISCAS11
2021 Live Demonstration: A 0.8V, 1.54 pJ / 940 MHz Dual Mode Logic-Based 16x16-Bit Booth Multiplier in 16-nm FinFET
abstract
The Dual Mode Logic (DML) defines run-time adaptive digital architectures that switch to either improved performance or lower energy consumption as a function of actual computational workload. This flexibility is demonstrated for the first time by silicon measurements on a 16×16-bit Booth multiplier fabricated as a part of an ultra-low power digital signal processing (DSP) architecture for 16-nm FinFET technology. When running in the full-speed mode, the DML multiplier can achieve a performance boost of 19.5% as compared to the equivalent standard CMOS design. The same design saves precious energy (-27%, on average) when the energy-efficient mode is enabled, while occupying 13% less silicon area.
Netanel Shavit, Inbal Stanger, Ramiro Taco, Marco Lanuzza, Alexander Fish
ISCAS4
2021 Live Demo: Silicon Evaluation of Multimode Dual Mode Logic for PVT-Aware Datapaths
abstract
This demo demonstrates the unique capabilities of the multimode Dual Mode Logic (DML) design technique to define run-time adaptive datapaths to overcome process and environmental (i.e., temperature and voltage) variations. A proof-of concept benchmark circuit is designed and fabricated in 65 nm technology. Measurements on 10 test chips, while considering supply voltages spanning 0.6V to 1.2V and temperature variations ranging from - 40 ° C to 125 ° C confirmed the effectiveness of the proposed approach to compensate even for severe process, voltage and temperature (PVT) variations.
Inbal Stanger, Netanel Shavit, Ramiro Taco, Marco Lanuzza, Alexander Fish
ISCAS4
2021 Gain-Cell Embedded DRAM Under Cryogenic Operation - A First Study
abstract
Operating circuits under cryogenic conditions is effective for a large spectrum of applications. However, the refrigeration requirement for the cooling of cryogenic systems introduces serious issues in terms of power dissipation. Gain-cell embedded dynamic random access memory (GC-eDRAM) is a low-area, logic-compatible embedded memory alternative to static random access memory (SRAM), which has the potential to provide ultralow-power operation under cryogenic conditions due to the lower leakages at these temperatures. In this article, we present the first comparative design exploration of GC-eDRAM under cryogenic conditions performed with transistor models characterized based on actual silicon measurements under temperatures as low as 77 K. Our study shows that the two-transistor (2T)-based GC-eDRAM configurations turn out to be the best solutions for very low-temperature operation. In particular, the 2T mixed GC-eDRAM configurations allow read sensing margin improvements (up to 99%) within the 2T-based configurations while at the same time excel in terms of data retention time (+44%) and power consumption (-27%) when compared to more complex GC-eDRAM topologies. Moreover, even better improvements in terms of area (-73%), leakage power (-97%), retention power (-76%), and energy (-66%) are observed when compared to conventional 6T-SRAM.
Esteban Garzón, Yosi Greenblatt, Odem Harel, Marco Lanuzza, Adam Teman
IEEE Trans. Very Large Scale Integr. Syst.4
2020 Robust Dual Mode Pass Logic (DMPL) for Energy Efficiency and High Performance
abstract
In the past, Pass Transistor Logic (PTL) was widely used due to benefits in terms of speed and power consumption coming from the reduced number of transistors. However, issues such as threshold drop across the single-channel pass transistors and high sensitivity to process variations have prevented the use of PTL in advanced nanometer technologies. In this paper, we propose a novel logic family named Dual Mode Pass Logic (DMPL), which allows for high speed and low power consumption while maintaining robustness down to the sub-threshold voltage region. The DMPL effectively combines PTL to reduce energy and power consumption along with the flexibility of Dual Mode Logic (DML) to switch to a speed improved operating mode according to the system requirement. Simulation analysis performed on basic NOR/NAND gates implemented in 16 nm Finfet technology demonstrates that DMPL can reduce energy and power by 33% and 42% as compared to logically equivalent static CMOS design. Moreover, running frequency of a DMPL circuit can exceed that of its static CMOS counterpart by 84% when speed is mandatory. Additionally, DMPL gates demonstrate similar robustness as static CMOS implementations under process and temperature variations at lower supply voltages.
Inbal Stanger, Netanel Shavit, Ramiro Taco, Leonid Yavits, Marco Lanuzza, Alexander Fish
ISCAS5
2020 Exploiting Single-Well Design for Energy-Efficient Ultra-Wide Voltage Range Dual Mode Logic-Based Digital Circuits in 28nm FD-SOI Technology
abstract
In this paper we evaluate the implementation options of energy-efficient dual mode logic (DML) circuits in 28nm fully depleted silicon-on-insulator (FD-SOI) technology. The combination of the flexibility of Dual Mode Logic (DML) and the unique characteristics of the FD-SOI technology has enormous potential to design energy-efficient adaptive digital circuits operating on an ultra-wide voltage range. As a main result, we demonstrate that single well option offered by the FD-SOI greatly extends the low-granularity energy-delay (E-D) optimization capability of DML-based designs. By exploiting the above implementation strategy, a 16-bit DML carry skip adder reduces its energy consumption by 41% and increases its speed of about 26% when changing its operation mode (from static to dynamic) at 0.4V as compared to its equivalent standard CMOS design.
Ramiro Taco, Leonid Yavits, Netanel Shavit, Inbal Stanger, Marco Lanuzza, Alexander Fish
ISCAS5
2020 Assessment of STT-MRAMs based on double-barrier MTJs for cache applications by means of a device-to-system level simulation framework
Esteban Garzón, Raffaele De Rose, Felice Crupi, Lionel Trojman, Giovanni Finocchio, Mario Carpentieri, Marco Lanuzza
Integr.7
2019 Live Demo: An 88fJ / 40 MHz [0.4V] - 0.61pJ / 1GHz [0.9V] Dual Mode Logic 8×8-Bit Multiplier Accumulator with a Self-Adjustment Mechanism in 28 nm FD-SOI
abstract
The unique ability of dual mode logic (DML) to self-adapt to computational needs by providing high speed and/or low energy consumption is demonstrated for the first time by silicon measurements in 28nm FD-SOI. At the gate level, the DML design offers the possibility to operate either in the static mode to save energy, or in the dynamic mode to increase speed albeit with higher delay or energy consumption, respectively. In this demonstration, the two operational modes are dynamically managed by a self-adjustment mechanism to increase speed or reduce energy of the design at run-time. As a test case a two-stage pipelined multiply-accumulate (MAC) circuit was selected to assess the advantages of DML in terms of speed, energy and area as compared to a conventional CMOS design. We show that the self-adjusted DML MAC achieves both a performance boost of up to 92% and 16% less energy consumption than the equivalent standard CMOS implementation. The energy saved can be even greater (-35%) when the low-power (fully static) mode is enabled. In addition, the DML MAC occupies 25% less area.
Ramiro Taco, Itamar Levi, Marco Lanuzza, Alexander Fish
ISCAS3
2019 Double-precision Dual Mode Logic carry-save multiplier
Raffaele De Rose, Paul Romero, Marco Lanuzza
Integr.3
2017 A variation-aware simulation framework for hybrid CMOS/spintronic circuits
abstract
In this paper, a variation-aware simulation framework is introduced for hybrid circuits comprising MOS transistors and spintronic devices (e.g., magnetic tunnel junction-MTJ). The simulation framework is based on one-time characterization via micromagnetic multi-domain simulations, as opposed to most of existing frameworks based on single-domain analysis. As further distinctive capability, stochastic variations of the MTJ switching are explicitly incorporated through a Skew Normal distribution, which is adjusted to fit micromagnetic simulations. The framework is implemented in the form of Verilog-A look-up table based model, which assures easy integration with commercial circuit design tools, and very low computational effort. The framework is applied to non-volatile Flip-FIops as case study with 10,000 Monte Carlo runs.
Raffaele De Rose, Marco Lanuzza, Felice Crupi, Giulio Siracusano, Riccardo Tomasello, Giovanni Finocchio, Mario Carpentieri, Massimo Alioto
ISCAS2
2017 Evaluation of Dual Mode Logic in 28nm FD-SOI technology
abstract
For the first time, the Dual Mode Logic (DML) technique is evaluated in 28 nm UTBB FD-SOI technology, with the goal of improving energy efficiency for wide supply voltage operation range. By combining the operating characteristics of the DML and the extended body bias capability of the technology, energy efficient digital circuits that can effectively benefit from adaptive voltage and frequency scaling techniques can be defined. This manuscript reports evaluations of the DML against conventional static and dynamic CMOS logics for two benchmarks in the 0.3V-1V supply voltage range. First, a NAND-NOR chain was considered. Simulation results showed that the DML approach assures roughly the 40% savings in terms of energy consumption with respect to the static CMOS implementation and improves the speed about 20% in comparison to the dynamic CMOS design. Second, a 16-bit Carry Skip Adder was considered. Due to the unique capability of the DML to switch on-the-fly between static and dynamic modes of operation, an improvement of more than 20% in terms of EDP was obtained in comparison to the conventional CMOS adder design.
Ramiro Taco, Itamar Levi, Marco Lanuzza, Alexander Fish
ISCAS3
2016 Extended exploration of low granularity back biasing control in 28nm UTBB FD-SOI technology
abstract
Recently, we proposed a low-granularity back-bias control technique [1] optimized for the ultra-thin body and box (UTBB) fully-depleted silicon-on-insulator (FD-SOI) technology. The technique was preliminary evaluated through the design of a low-voltage 8-bit ripple carry adder (RCA), showing very competitive energy and delay values. In this paper, the characteristics of the low-granularity back-biasing control are explored considering as benchmarks basic logic gates as well as adders with different bit lengths. All the designed circuits were compared to their equivalent dynamic threshold voltage MOSFET (DTMOS) and conventional CMOS designs. The higher efficiency of low granularity body bias control is emphasized by the single well layout strategy, offered by the 28 nm UTBB FD-SOI technology, thus leading our approach to achieve competitive silicon area occupancy along with significant performance and energy improvements. More precisely, post-layout simulations have demonstrated that circuits designed according the suggested strategy, can achieve a delay reduction of 33% compared to conventional CMOS designs, whereas the energy consumption can be reduced down to 46% compared to DTMOS solutions, for a supply voltage of 0.4V. These results were obtained while maintaining robustness against process and temperature variations.
Ramiro Taco, Itamar Levi, Marco Lanuzza, Alexander Fish
ISCAS3
2015 Fast and Wide Range Voltage Conversion in Multisupply Voltage Designs
abstract
Multisupply voltage design technique is widely used in modern system-on-chips to tradeoff energy and speed. Level shifters (LSs) allow different voltage domains to be interfaced. In this brief, a new LS is presented for fast and wide range voltage conversion. Because of a novel architecture combined with the use of multithreshold CMOS technique, the proposed circuit guarantees robust voltage shifting from the deep subthreshold to the above-threshold domain while exhibiting fast response and low energy consumption. When implemented in a 90-nm technology node, considering process-voltage-temperature variations, the proposed design reliably converts 100-mV input signals into 1 V output signals. Post-layout simulation results demonstrate that the new LS shows a propagation delay of 16.6 ns, a static power dissipation of 8.7 nW and a total energy per transition of only 77 fJ for a 0.2 V 1-MHz input pulse.
Marco Lanuzza, Pasquale Corsonello, Stefania Perri
IEEE Trans. Very Large Scale Integr. Syst.1
2010 A new low-power high-speed single-clock-cycle binary comparator
abstract
This paper presents a new ultra-low power high-speed single-clock-cycle binary comparator. It is based on a novel parallel-prefix algorithm which drastically reduces the switching activity of the internal nodes of the circuit. When implemented by using the ST 90nm-1V technology, the proposed 64-bit comparator exhibits an energy dissipation of only 0.77μW/MHz and a delay of 258ps. With respect to a recently published low-power high-speed parallel-prefix adder, the proposed design shows an energy dissipation reduction of 23% and a speed improvement of 7%.
Fabio Frustaci, Stefania Perri, Marco Lanuzza, Pasquale Corsonello
ISCAS3
2010 Exploiting Self-Reconfiguration Capability to Improve SRAM-based FPGA Robustness in Space and Avionics Applications
abstract
This article presents a novel configuration scrubbing core, used for internal detection and correction of radiation-induced configuration single and multiple bit errors, without requiring external scrubbing. The proposed technique combines the benefits of fast radiation-induced fault detection with fast restoration of the device functionality and small area and power overheads. Experimental results demonstrate that the novel approach significantly improves the availability in hostile radiation environments of FPGA-based designs. When implemented using a Xilinx XC2V1000 Virtex-II device, the presented technique detects and corrects single bit upsets and double, triple and quadruple multi bit upsets, occupying just 1488 slices and dissipating less than 30 mW at a 50MHz running frequency.
Marco Lanuzza, Paolo Zicari, Fabio Frustaci, Stefania Perri, Pasquale Corsonello
ACM Trans. Reconfigurable Technol. Syst.1
2009 New performance/power/area efficient, reliable full adder design
abstract
Arithmetic circuits have always played one of the most important roles in the designs of processors, FPGAs, and the rapidly evolving domain of media processing architectures. The full adder cell forms the basic building block of majority of these arithmetic circuits. In this paper we describe a hybrid pseudo static full adder cell designed using Data Driven Dynamic Logic. Simulation results show the adder to out perform its competitors, both static as well as dynamic topologies in terms of performance, while maintaining relatively similar area and power characteristics. This paper presents a complete characterization of the popular adder cells in terms of delay, area, power, noise margin and reliability analysis for both super threshold and sub threshold operating regimes.
Sohan Purohit, Martin Margala, Marco Lanuzza, Pasquale Corsonello
ACM Great Lakes Symposium on VLSI3
2007 Design and Implementation of a 90nm Low bit-rate Image Compression Core
abstract
This paper presents a low-cost, high throughput discrete wavelet transform-based image compressor. The hardware solution proposed here exploits a modified set partitioning in hierarchical trees (SPIHT) algorithm and ensures that appropriate reconstructed image qualities can be achieved also for compression ratios over 100:1. Obtained results demonstrate that a maximum data rate of about 23 Mpixels/s can be sustained on a 64x64 size tile. In 90 nm technology, the required area is only 1.77 mm2. To obtain higher performance, multiples cores can be used in a parallel implementation.
Pasquale Corsonello, Stefania Perri, Giovanni Staino, Marco Lanuzza, Giuseppe Cocorullo
DSD4
2006 Low bit rate image compression core for onboard space applications
abstract
This paper presents low-cost, purpose optimized discrete wavelet transform-based image compressors for future spacecrafts and microsatellites. The hardware solution proposed here exploits a modified set partitioning in hierarchical trees algorithm and ensures that appropriate reconstructed image qualities can be achieved also for compression ratios over 100:1. Several implementations are presented varying the parallelism level and the tile size. Obtained results demonstrate that, using a parallel implementation operating on a 64 /spl times/ 64 size tile, a maximum data rate of about 18 Mpixels/s can be sustained. In this case, only 4500 slices and 24 BlockRAMs of a XILINX Virtex II device are required.
Pasquale Corsonello, Stefania Perri, Giovanni Staino, Marco Lanuzza, Giuseppe Cocorullo
IEEE Trans. Circuits Syst. Video Technol.4
2005 Low-Cost Fully Reconfigurable Data-Path for FPGA-Based Multimedia Processor
abstract
This paper describes novel data-path architecture for FPGA-based multimedia processors. The proposed circuit can adapt itself at run-time to different operations and data wordlengths avoiding time and power consuming reconfiguration. The new data-path can operate in SIMD fashion and guarantees high parallelism levels when operations on lower precisions are executed. It also supports IEEE-754 compliant single precision floating-point addition and multiplication. The proposed circuit has been characterized using VIRTEXII XILINX devices, but it can be efficiently used also in other FPGA families.
Marco Lanuzza, Stefania Perri, Martin Margala, Pasquale Corsonello
FPL1
2005 Cost-effective low-power processor-in-memory-based reconfigurable datapath for multimedia applications
abstract
Multimedia applications have become a dominant computing workload for computer systems as well as for wireless-based devices. Due to their repetitive computing and memory intensive nature, they can take effective advantage from Processor-In-Memory (PIM) technology. In this paper, a new low-power PIM-based 32-bit reconfigurable datapath optimized for multimedia applications is presented. The new circuit efficiently performs parallel arithmetic operations on either 8-, 16-, or 32-bit integer data or on 32-bit single precision floating-point data. As a result, high flexibility is provided at a very low hardware cost. When implemented using the UMC 0.18 μm 1.8 V CMOS technology, the proposed datapath exhibits a 285 MHz running frequency, dissipates just 0.12 mW/MHz and occupies a silicon area of only 107,323 μm2. When performing 2D-DCT, proposed architecture consumes 74% less power and is 28% more power efficient compared to top-of-the-line commercial TI DSP
Marco Lanuzza, Martin Margala, Pasquale Corsonello
ISLPED1
2004 Variable precision arithmetic circuits for FPGA-based multimedia processors
abstract
This brief describes new efficient variable precision arithmetic circuits for field programmable gate array (FPGA)-based processors. The proposed circuits can adapt themselves to different data word lengths, avoiding time and power consuming reconfiguration. This is made possible thanks to the introduction of on purpose designed auxiliary logic, which enables the new circuits to operate in single instruction multiple data (SIMD) fashion and allows high parallelism levels to be guaranteed when operations on lower precisions are executed. The new SIMD structures have been designed to optimally exploit the resources of a widely used family of SRAM-based FPGAs, but their architectures can be easily adapted to any either SRAM-based or antifuse-based FPGA chips.
Stefania Perri, Pasquale Corsonello, Maria Antonia Iachino, Marco Lanuzza, Giuseppe Cocorullo
IEEE Trans. Very Large Scale Integr. Syst.4