Antonio G. M. Strollo

dblp:23/2957 · also Antonio Giuseppe Maria Strollo · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-5737-1783ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 6 first-author · 8 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 A novel approximate multiplier based on improved logarithmic and antilogarithmic conversions
abstract
In this paper, we propose a novel approximate multiplier with error-improved logarithmic and antilogarithmic conversion. In logarithmic multiplication, the forward and backward conversions to the logarithm domain introduce errors in the product. To recover accuracy, we apply an offset when logarithm and antilogarithm are computed, and find suitable values for these offsets in order to make the approximation error with zero mean. The hardware structure of the proposed multiplier is described for signed arithmetic, and a novel hardware-efficient approach, able to compute the sign of the output, is also proposed. Error metrics and synthesis analyses in a 28nm CMOS technology show that our multiplier offers the best accuracy-hardware performance for MRED-2. Furthermore, remarkable results are achieved also in JPEG compression, exhibiting SSIM and PSNR metrics competitive with the state-of-the-art.
Gennaro Di Meo, Davide De Caro, Luca Tegazzini, Ettore Napoli, Antonio G. M. Strollo
ISCAS5
2025 Low-Power High Precision Floating-Point Divider With Bidimensional Linear Approximation
abstract
In this paper we propose a novel approximate floating-point divider based on bidimensional linear approximation. In our approach, the mantissa quotient is seen as a function of the two input mantissas of the divider. The domain of this two-variable function is partitioned into$nx \times ny$subregions, named tiles, where$nx, ny$are chosen as powers of two. In each tile the quotient is approximated with a linear combination of the input mantissas. To achieve fine accuracy, an optimization problem is formulated within each tile to determine the optimal coefficients for the linear combination, which minimize the Mean Relative Error Distance (MRED) of the divider. Furthermore, to make hardware implementation more effective, the minimization problem is appropriately modified to search for optimal quantized coefficients. The hardware structure of the divider only requires a small look-up table to store the linear approximation coefficients, and a carry save adder tree. The proposed architecture is highly tunable at design-time over a wide range of accuracy, depending on the number of tiles chosen for the approximation. The obtained results demonstrate error performance and hardware features superior to the state-of-the-art. The proposed dividers define the Pareto front, considering the trade-off between power-delay-product vs. MRED and area-delay-product vs. MRED, for MRED in the range of$4\times 10^{-3}-2\times 10^{-2}$. Application results for JPEG compression and tone mapping further highlight the strength of our proposal, which exhibits Structural Similarity Index (SSIM) very close to 1 in all cases and Peak Signal-to-Noise Ratio (PSNR) up to 45 dB.
Gennaro Di Meo, Antonio G. M. Strollo, Davide De Caro, Luca Tegazzini, Ettore Napoli
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Comprehensive Analysis of Input Order Invariant Approximate 4-2 Compressors for Binary Multipliers
abstract
Approximate arithmetic circuits sacrifice computing accuracy in exchange for improvements in power, area, and speed. Many approximate binary multipliers that use 4-2 compressors have been proposed but most of the proposals have error performances that depend on the order in which the inputs are connected to the compressors. This complicates the design and prevents a fair comparison among different approximate multipliers. The paper proposes the input order invariant approximate 4-2 compressors, whose behavior remains consistent regardless of the order of input signals. We derive the complete set of such compressors and utilize them to synthesize 8-bit multipliers. Our analysis reveals that only a limited subset of input order invariant approximate 4-2 compressors offers an optimal balance between error and power savings. Furthermore, we demonstrate that this optimal set of 4-2 compressors can be strategically distributed within the columns of the multiplier to further enhance the trade-off between error and power efficiency. The proposed circuits, once implemented in a 14 nm FinFET standard cell technology, favorably compare against the state of the art.
Ettore Napoli, Antonio G. M. Strollo, Efstratios Zacharelos, Gennaro Di Meo
ISCAS2
2023 Approximate squaring circuits exploiting recursive architectures
abstract
Error resilient applications benefit from the use of approximate computing techniques that enhance electrical performances while allowing a deviation from the exact result. Many operations in signal processing require the square of a signal. Despite the fact that the squaring operation can be regarded as a special multiplication case, it is often preferable to develop independent squaring circuits to exploit possible architectural symmetries. This paper proposes novel approximate binary squarers, obtained by recursively exploiting 4-bit approximate multipliers and squarers. The final designs cover a wide range of computing precision, providing the user with multiple choices of different cost vs. accuracy trade-offs. The proposed circuits, as well as competitive designs, are synthesized targeting a 14 nm FinFET technology to determine the electrical characteristics. It is demonstrated that the proposed squarers outperform the state-of-the-art in terms of power vs. precision. Compared to the exact 8-bit squarer, the least dissipative proposed design reduces silicon area by 76%, power consumption by 71%, and critical delay by 72%. The same circuit dissipates 2.4% less power than the least dissipative design found in literature, while providing 34% more accurate results. The behavior of the considered designs is also tested in common error resilient applications, like signal demodulation and image processing.
Efstratios Zacharelos, Italo Nunziata, Gerardo Saggese, Antonio G. M. Strollo, Ettore Napoli
Integr.4
2023 Novel Low-Power Floating-Point Divider With Linear Approximation and Minimum Mean Relative Error
abstract
Floating-point division involves the computation of the ratio (1$+$Mx)/(1$+$My), whereMxandMyrepresents the mantissas of the input values. In this paper, we propose a new method for approximating this operation using a linear function ofMx, with coefficients that depend onMy. The coefficients are calculated to minimize the Mean Relative Error Distance (MRED) of the approximation. To this end, the range of My is partitioned in N sub-intervals where the minimization ofMREDis formulated as a linear programming problem, whose solution gives optimal coefficient values. The hardware implementation requires a small lookup table, two multipliers and an adder. An aggressive coefficients quantization is exploited to further optimize the design. ObtainedMREDimproves by increasing$N$, ranging from 1.4% to 0.33%. Implementation results in a 28nm CMOS technology show that the proposed design outperforms the state-of-the-art, offering the best trade-off between hardware complexity and accuracy. Results for two image processing applications, change detection and JPEG compression, demonstrate remarkable performance, with SSIM very close to 1 and PSNR values exceeding 50dB.
Gennaro Di Meo, Antonio G. M. Strollo, Davide De Caro
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 Approximate Recursive Multipliers Using Low Power Building Blocks
Efstratios Zacharelos, Italo Nunziata, Gerardo Saggese, Antonio G. M. Strollo, Ettore Napoli
ARITH4
2022 A Novel Module-Sign Low-Power Implementation for the DLMS Adaptive Filter With Low Steady-State Error
abstract
In this paper, a novel implementation is proposed for the Delayed LMS (DLMS) filter, able to reduce the power dissipation while preserving regime performances. The approach relies on the observation that the error signal is small in magnitude and oscillates around zero when the circuit is close to the convergence point. Therefore, the most significant bits of the error signal continuously toggle from positive to negative values causing high switching activity in the multipliers of the feedback section. This paper proposes to employ a sign-modulus representation of the error signal, to substantially reduce the switching activity of the feedback path of the filter. Additional approximation techniques are also devised to further reduce power dissipation. Comparisons with the state-of-the-art show that the proposed filter is the only one able to approach the MSE of the exact implementation with a remarkable reduction of power dissipation. A test-chip in TSMC 28nm CMOS technology has been realized to experimentally verify the validity of our technique. The experimental results show the possibility of saving up to 45.4% of power consumption with respect to the exact implementation of the filter.
Gennaro Di Meo, Davide De Caro, Giacinto Paolo Saggese, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 Approximate Multipliers Using Static Segmentation: Error Analysis and Improvements
abstract
Approximate multipliers are used in error-tolerant applications, sacrificing the accuracy of results to minimize power or delay. In this paper we investigate approximate multipliers using static segmentation. In these circuits a set of$m$contiguous bits (a segment of$m$bits) is extracted from each of the two$n$-bits operand, the two segments are in input to a small$m\times m$internal multiplier whose output is suitably shifted to obtain the result. We investigate both signed and unsigned multipliers, and for the latter we propose a new segmentation approach. We also present simple and effective correction techniques that can significantly reduce the approximation error with reduced hardware costs. We perform a detailed comparison with previously proposed approximate multipliers, considering a hardware implementation in 28 nm technology. The comparison shows that static segmented multipliers with the proposed correction technique have the desirable characteristic of being on (or close to) the Pareto-optimal frontier for both power vs normalized mean error distance and power vs mean relative error distance trade-off plots. These multipliers, therefore, are promising candidates for applications where their error performance is acceptable. This is confirmed by the results obtained for image processing and image classification applications.
Antonio G. M. Strollo, Ettore Napoli, Davide De Caro, Nicola Petra, Giacinto Paolo Saggese, Gennaro Di Meo
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Real-Time Downsampling in Digital Storage Oscilloscopes With Multichannel Architectures
abstract
Digital Storage Oscilloscopes (DSOs) conjugate high performance with large number of features and flexibility. The basic structure, based on fast Analog to Digital Converter (ADC) and memory, is augmented with several components for channel matching and bandwidth improvement, and processors that provide visualization, frequency processing, jitter and stability measurement, etc. Unfortunately, fine resolution in sample rate selection is not available, such that for several applications the user must run complex measurement procedures that require data download and offline processing. The paper proposes a dedicated digital circuit that offers fine control of the time-base by real time downsampling the input stream at an almost arbitrary sampling rate. The proposed circuit implements a series of operations involving real-time data filtering, defragmentation and packing, which are not considered by alternative approaches, like those based on polyphase filters, that offer very limited choices for the sampling rate. The circuit is designed to work in conjunction with the highest performance DSOs that use a multichannel architecture. Design rules for circuit design are provided together with implementation results in 14nm FinFET technology. When designed for an architecture with 64 channels with 8bit input samples the circuit works in real time with a sampling rate of 220 GSps running at 3.42 GHz, with a silicon footprint of 0.17 mm2and a power dissipation of 0.85W.
Ettore Napoli, Efstratios Zacharelos, Mauro D'Arco, Antonio G. M. Strollo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2020 A Binary Line Buffer Circuit Featuring Lossy Data Compression at Fixed Maximum Data Rate
abstract
Video processing requires an increasing amount of buffered data. The paper proposes a multi-line buffer circuit that stores compressed data thus saving logic and power. The lossy compression algorithm provides the output stream with a fixed, decided by the user, delay from the input stream. Further, the amount of memory of the compressed buffer can be designed to trade off the correctness of the output with the logic resources footprint. The circuit is tested in a computer vision processing chain as a binary line buffer for the stream of foreground pixels. The paper also proposes a multi-line buffer circuit that integrates a morphological operation. The latter, when compared against the state of the art and against the proposed multi-line lossy buffer, allows larger saving of logic resources (up to 75%) and power (up to 85%), with reduced penalty on video quality.
Ettore Napoli, Davide De Caro, Nicola Petra, Antonio G. M. Strollo
ISCAS4
2020 Low-Power Approximate Multiplier with Error Recovery using a New Approximate 4-2 Compressor
abstract
In this paper we propose an energy-efficient approximate multiplier which uses a new approximate 4-2 compressor. The proposed compressor has a low error probability and its error conditions can be easily detected. This, as previously shown in the literature, makes it possible to implement error recovery, when the compressor is used in the partial product reduction phase of a multiplier. Simulation results show that proposed approximate multipliers exhibit a sensible reduction in Mean Error Distance and in maximum Error Distance, compared to previous art. Application to an image processing task shows an improvement of about 8dB in peak signal-to-noise ratio. Implementation results in 28nm CMOS show that the electrical performance of multipliers designed with the novel circuit are close to the one obtained with previously proposed approximate compressors.
Antonio G. M. Strollo, Davide De Caro, Ettore Napoli, Nicola Petra, Gennaro Di Meo
ISCAS1
2018 On the Use of Approximate Multipliers in LMS Adaptive Filters
abstract
Approximate computing relaxes algorithm precision constraints to improve digital circuit performance. Adaptive filters based on least-mean-square (LMS) algorithm constitute a standard in many DSP applications. The LMS algorithm, being an approximation of the Wiener filter, is inherently imprecise, and constitutes a fertile ground to employ approximate hardware techniques with the additional challenge related to the presence of a feedback path for coefficients update. In this paper, approximate LMS adaptive filters are explored for the first time, by employing approximate multipliers. A system identification scenario is adopted to assess the algorithm behavior. The analysis reveals that the choice of the approximate multiplier topology should be carefully examined, otherwise the stability and convergence performance of the algorithm can be compromised. We propose a novel approximate multiplier able to reduce the power dissipation in adaptive LMS filters up to 29% with tolerable convergence error degradation.
Darjn Esposito, Gennaro Di Meo, Davide De Caro, Nicola Petra, Ettore Napoli, Antonio G. M. Strollo
ISCAS6
2017 On the use of approximate adders in carry-save multiplier-accumulators
abstract
Approximate computing improves digital circuit performance by relaxing the requirement of performing exact calculations. In this paper, we investigate the use of approximate adders in the final stage of a carry save multiplier-accumulator (MAC), designed for image filtering application. We propose a design flow based on synthesis tools, starting from HDL description. After a first step in which an exact carry-propagate adder is used, the synthesized netlist is simulated to extract the statistics of the terms summed in the carry-propagate adder. We then design the approximate adder, to meet the required error characteristics given the inputs statistics. The netlist is finally modified by substituting the exact adder with the approximate one, and a final synthesis and optimization is performed. The presented design example in 28nm CMOS shows that a 14% power gain can be obtained, with a limited image quality degradation.
Darjn Esposito, Davide De Caro, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
ISCAS5
2017 Power-precision scalable latch memories
abstract
Approximate computing leverages the inherent error resiliency present in many applications to improve circuits performance. Precision-scalable systems dynamically introduce approximations to trade off power and quality, based on the application under execution and the incoming dataset. In this paper, this principle is explored for the first time in the context of latch memories by introducing the ability to scale their precision, while retaining the ability to synthesize them in an automated manner. This offers additional opportunities to reduce energy, compared to the well-known suitability for aggressive voltage scaling of latch memories. A case study based on image processing applications is presented to evaluate the quality-power trade-off in 40nm CMOS. The analysis shows that the total power is reduced by up to 56% when the precision requirement is relaxed.
Darjn Esposito, Antonio G. M. Strollo, Massimo Alioto
ISCAS2
2017 A SISO Register Circuit Tailored for Input Data with Low Transition Probability
abstract
The paper proposes a SISO register circuit, functionally equivalent to a Shift Register, that is the optimal design choice when the input data have a reduced transition probability. The proposed circuit obtains improved performances by only storing the transitions of the input data, thus saving logic and power.
Ettore Napoli, Gerardo Castellano, Davide De Caro, Darjn Esposito, Nicola Petra, Antonio G. M. Strollo
IEEE Trans. Computers6
2016 Approximate adder with output correction for error tolerant applications and Gaussian distributed inputs
abstract
Approximate computing is emerging as a new paradigm to improve digital circuit performance by relaxing the requirement of performing exact calculations. Approximate adders rely on the idea that for uniformly distributed inputs, long carry-propagation chains are rarely activated. Unfortunately, however, the above assumption on input signal statistics is not always verified; in this paper we focus on the case (often encountered in practical signal processing applications) when the inputs have a Gaussian distribution. We show that for Gaussian inputs the error probability of previously proposed approximate adders approaches 25% for low sigma values, which is much larger than the uniform case. On the basis of this analysis, we propose an approximate adder with a correction circuit that drastically reduces the error rate for Gaussian distributed operand s. In order to investigate the performance of our approach in a real application, simulated results for a simple audio processing system are reported. Implementation results in 65nm technology are also presented.
Darjn Esposito, Gerardo Castellano, Davide De Caro, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
ISCAS6
2015 An FPGA processor for real-time, fixed-point refinement of CDVS keypoints
abstract
Computer Vision is a more and more pervasive technology in nowadays image and video processing applications: examples include image driven search, stereoscopical matching, panorama stitching and industrial automation. Compact Descriptors for Visual Search (CDVS) is an algorithm for Computer Vision recently proposed as part of the MPEG-7 standard: it has the ability to select points of interest in the image (also referred to as keypoints) that exhibit robustness, in a certain degree, with respect to changes like homogeneous variations in luminance, changes in point of view, rotations, rescaling and geometrical distortion of the image. Keypoint Refinement is a phase of the CDVS algorithm which is aimed at discarding candidate keypoints that are likely to be unstable for their algebraic properties. This paper presents an FPGA circuit design that implements this phase on fixed point data with real time compatible throughput. Implementation results show a negligible impact on resources allocation even on mid-sized FPGAs.
Giorgio Lopez, Ettore Napoli, Domenico Meglio, Antonio G. M. Strollo
ISCAS4
2015 Hardware implementation of a spatio-temporal average filter for real-time denoising of fluoroscopic images
Mariangela Genovese, Paolo Bifulco, Davide De Caro, Ettore Napoli, Nicola Petra, Maria Romano, Mario Cesarelli, Antonio G. M. Strollo
Integr.8
2014 FPGA based system for the generation of noise with programmable power spectrum
abstract
Noise sources are needed for test and validation of noise sensitive electronic systems but only wide band white noise sources are directly available on the market. In this paper a programmable colored noise generator is proposed. The system allows to configure the spectral features of the noise and is implemented with a Field Programmable Gate Array that produces the digital samples of the noise and a Digital to Analog Converter that produces the analogue output. The proposed generator overcomes the state of the art in terms of bandwidth and flexibility and produces a noise sequence whose length is unlimited for practical purposes. Experimental results show that the bandwidth of the generated noise can be selected up to a maximum of 120 MHz while defining the power spectral density with a frequency resolution equal to 0.2 % of the selected bandwidth.
Ettore Napoli, Mauro D'Arco, Pasquale Di Cosmo, Mariangela Genovese, Antonio G. M. Strollo
ISCAS5
2014 Analysis and comparison of Direct Digital Frequency Synthesizers implemented on FPGA
Mariangela Genovese, Ettore Napoli, Davide De Caro, Nicola Petra, Antonio G. M. Strollo
Integr.5
2011 Elementary Functions Hardware Implementation Using Constrained Piecewise-Polynomial Approximations
abstract
A novel technique for designing piecewise-polynomial interpolators for hardware implementation of elementary functions is investigated in this paper. In the proposed approach, the interval where the function is approximated is subdivided in equal length segments and two adjacent segments are grouped in a segment pair. Suitable constraints are then imposed between the coefficients of the two interpolating polynomials in each segment pair. This allows reducing the total number of stored coefficients. It is found that the increase in the approximation error due to constraints between polynomial coefficients can easily be overcome by increasing the fractional bits of the coefficients. Overall, compared with standard unconstrained piecewise-polynomial approximation having the same accuracy, the proposed method results in a considerable advantage in terms of the size of the lookup table needed to store polynomial coefficients. The calculus of the coefficients of constrained polynomials and the optimization of coefficients bit width is also investigated in this paper. Results for several elementary functions and target precision ranging from 12 to 42 bits are presented. The paper also presents VLSI implementation results, targeting a 90 nm CMOS technology, and using both direct and Horner architectures for constrained degree-1, degree-2, and degree-3 approximations.
Antonio G. M. Strollo, Davide De Caro, Nicola Petra
IEEE Trans. Computers1
2010 High-speed differential resistor ladder for A/D converters
abstract
This paper describes the implementation of a novel high-speed differential resistor ladder suited for A/D converters. The novel ladder yields, theoretically, up to a sixteen-fold reduction of the propagation delay with respect to the conventional differential ladder. Simulation results, for a BiCMOS 0.25μm technology, show that the novel ladder results in a fivefold increase of the maximum sampling frequency when employed to design a 8-bit Flash converter. A 70% higher speed is also highlighted when the ladder is employed in a Folding and Interpolating 8-bit converter.
Davide De Caro, Marino Coppola, Nicola Petra, Ettore Napoli, Antonio G. M. Strollo, Valeria Garofalo
ISCAS5
2010 A novel truncated squarer with linear compensation function
abstract
A truncated binary squarer is a squarer with a n bit input that produces a n bit output. The proposed design minimizes the mean square error of the squarer and results in a very simple and fast circuital implementation. The squarer, compared against state of the art circuits, provides a reduction of the mean square error ranging from 20% to 5%. At the same time, the proposed squarer is able to reduce the power dissipation, reduce the silicon area occupation, and increase the maximum working frequency. Implementations results are provided for a 0.18μm technology.
Valeria Garofalo, Marino Coppola, Davide De Caro, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
ISCAS6
2010 Fixed-width CSD multipliers with minimum mean square error
abstract
Many multimedia and DSP applications require fixed-width multipliers, in which input data and output results have the same bit width. In this paper we investigate fixed-width multipliers where one of the input operand is a constant, encoded using canonic signed digit (CSD) representation. This is a very important case in many practical applications such as the calculation of Fast Fourier Transform. In the paper we derive in closed form the expression of the compensation function giving the minimum mean square error for CSD fixed-width multiplier. On the basis of this analytical result, we propose a hardware efficient implementation of the multiplier. Fixed width CSD multipliers implemented with the approach presented in this paper are accurate and can be implemented by using a simple partial-product reduction tree followed by a fast adder, without requiring additional look-up tables. The proposed approach is general and is well suited for implementation in circuit synthesizers. Implementation results in 90 nm technology are presented, to demonstrate the effectiveness of the proposed technique.
Nicola Petra, Davide De Caro, Antonio G. M. Strollo, Valeria Garofalo, Ettore Napoli, Marino Coppola, Pietro Todisco
ISCAS3
2008 A high performance floating-point special function unit using constrained piecewise quadratic approximation
abstract
A special function unit, able to compute square root, reciprocal square root, logarithm and exponential functions is presented in this paper. The system supports single precision IEEE-754 floating-point standard and uses a novel constrained piecewise quadratic interpolation technique to approximate the implemented functions. The proposed approach allows to reduce look-up table size of 40% with respect to previously proposed techniques. The SFU has been implemented in a test chip in 0.18 mum CMOS. A maximum clock frequency of 420 MHz and a power dissipation of 160 mW@420 MHz have been measured.
Davide De Caro, Nicola Petra, Antonio G. M. Strollo
ISCAS3
2007 A Novel Architecture for Galois Fields GF(2^m) Multipliers Based on Mastrovito Scheme
abstract
In the paper a new GF(2^m) multiplier for standard basis representation is developed. Proposed multiplier implements the Mastrovito multiplication scheme and can be designed for every field GF(2^m). A minimum area implementation of the first block of Mastrovito multiplier and a high-speed delay-driven tree architecture for the second block of Mastrovito multiplier are employed in the new circuit. Multiplier complexity and delay are analytically evaluated for many polynomial classes. Timing and area occupation performances of the proposed multiplier are also calculated for many fields used in Reed-Solomon codes applications and compared with those of previously proposed solutions. The comparison shows that the proposed multiplier outperforms previous architectures for every considered GF(2^m) field. The effectiveness of the proposed solution in a real application is verified by implementing in a 0.25ìm CMOS technology the key equation solving block of a (255,239) Reed-Solomon decoder. The use of the proposed multiplier in this application results in a substantial speed improvement without any penalty in silicon area occupation.
Nicola Petra, Davide De Caro, Antonio G. M. Strollo
IEEE Trans. Computers3
2005 A novel high-speed sense-amplifier-based flip-flop
abstract
A new sense-amplifier-based flip-flop is presented. The output latch of the proposed circuit can be considered as an hybrid solution between the standard NAND-based set/reset latch and the NC-/sup 2/MOS approach. The proposed flip-flop provides ratioless design, reduced short-circuit power dissipation, and glitch-free operation. The simulation results, obtained for a 0.25-/spl mu/m technology, show improvements in the clock-to-output delay and the power dissipation with respect to the recently proposed high-speed flip-flops. The new circuit has been successfully employed in a high-speed direct digital frequency synthesizer chip, highlighting the effectiveness of the proposed flip-flop in high-speed standard cell-based applications.
Antonio G. M. Strollo, Davide De Caro, Ettore Napoli, Nicola Petra
IEEE Trans. Very Large Scale Integr. Syst.1
2003 An FPGA-Based Performance Analysis of the Unrolling, Tiling, and Pipelining of the AES Algorithm
Giacinto Paolo Saggese, Antonino Mazzeo, Nicola Mazzocca, Antonio G. M. Strollo
FPL4
2002 A Technique for FPGA Synthesis Driven by Automatic Source Code Analysis and Transformations
Beniamino Di Martino, Nicola Mazzocca, Giacinto Paolo Saggese, Antonio G. M. Strollo
FPL4
2000 New clock-gating techniques for low-power flip-flops
abstract
Two novel low power flip-flops are presented in the paper. Proposed flip-flops use new gating techniques that reduce power dissipation deactivating the clock signal. Presented circuits overcome the clock duty-cycle limitation of previously reported gated flip-flops. Circuit simulations with the inclusion of parasitics show that sensible power dissipation reduction is possible if input signal has reduced switching activity. A 16-bit counter is presented as a simple low power application.
Antonio G. M. Strollo, Ettore Napoli, Davide De Caro
ISLPED1
2000 Analysis of power dissipation in double edge-triggered flip-flops
abstract
A comprehensive analysis of double edge triggered (DET) flip-flops' power dissipation, taking into account input signal statistics, is presented in this paper. It is shown that using DET instead of a single edge-triggered flip-flop may result in significant energy savings if the input signal has reduced activity. On the other hand, the high switching rate of DET internal nodes may result in larger power dissipation if the input signal has a high transition probability or significant glitching.
Antonio G. M. Strollo, Ettore Napoli, Carlo Cimino
IEEE Trans. Very Large Scale Integr. Syst.1