Davide De Caro

dblp:36/4231 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-0204-0949ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 A novel approximate multiplier based on improved logarithmic and antilogarithmic conversions
abstract
In this paper, we propose a novel approximate multiplier with error-improved logarithmic and antilogarithmic conversion. In logarithmic multiplication, the forward and backward conversions to the logarithm domain introduce errors in the product. To recover accuracy, we apply an offset when logarithm and antilogarithm are computed, and find suitable values for these offsets in order to make the approximation error with zero mean. The hardware structure of the proposed multiplier is described for signed arithmetic, and a novel hardware-efficient approach, able to compute the sign of the output, is also proposed. Error metrics and synthesis analyses in a 28nm CMOS technology show that our multiplier offers the best accuracy-hardware performance for MRED-2. Furthermore, remarkable results are achieved also in JPEG compression, exhibiting SSIM and PSNR metrics competitive with the state-of-the-art.
Gennaro Di Meo, Davide De Caro, Luca Tegazzini, Ettore Napoli, Antonio G. M. Strollo
ISCAS2
2025 Low-Power High Precision Floating-Point Divider With Bidimensional Linear Approximation
abstract
In this paper we propose a novel approximate floating-point divider based on bidimensional linear approximation. In our approach, the mantissa quotient is seen as a function of the two input mantissas of the divider. The domain of this two-variable function is partitioned into$nx \times ny$subregions, named tiles, where$nx, ny$are chosen as powers of two. In each tile the quotient is approximated with a linear combination of the input mantissas. To achieve fine accuracy, an optimization problem is formulated within each tile to determine the optimal coefficients for the linear combination, which minimize the Mean Relative Error Distance (MRED) of the divider. Furthermore, to make hardware implementation more effective, the minimization problem is appropriately modified to search for optimal quantized coefficients. The hardware structure of the divider only requires a small look-up table to store the linear approximation coefficients, and a carry save adder tree. The proposed architecture is highly tunable at design-time over a wide range of accuracy, depending on the number of tiles chosen for the approximation. The obtained results demonstrate error performance and hardware features superior to the state-of-the-art. The proposed dividers define the Pareto front, considering the trade-off between power-delay-product vs. MRED and area-delay-product vs. MRED, for MRED in the range of$4\times 10^{-3}-2\times 10^{-2}$. Application results for JPEG compression and tone mapping further highlight the strength of our proposal, which exhibits Structural Similarity Index (SSIM) very close to 1 in all cases and Peak Signal-to-Noise Ratio (PSNR) up to 45 dB.
Gennaro Di Meo, Antonio G. M. Strollo, Davide De Caro, Luca Tegazzini, Ettore Napoli
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 Novel Low-Power Floating-Point Divider With Linear Approximation and Minimum Mean Relative Error
abstract
Floating-point division involves the computation of the ratio (1$+$Mx)/(1$+$My), whereMxandMyrepresents the mantissas of the input values. In this paper, we propose a new method for approximating this operation using a linear function ofMx, with coefficients that depend onMy. The coefficients are calculated to minimize the Mean Relative Error Distance (MRED) of the approximation. To this end, the range of My is partitioned in N sub-intervals where the minimization ofMREDis formulated as a linear programming problem, whose solution gives optimal coefficient values. The hardware implementation requires a small lookup table, two multipliers and an adder. An aggressive coefficients quantization is exploited to further optimize the design. ObtainedMREDimproves by increasing$N$, ranging from 1.4% to 0.33%. Implementation results in a 28nm CMOS technology show that the proposed design outperforms the state-of-the-art, offering the best trade-off between hardware complexity and accuracy. Results for two image processing applications, change detection and JPEG compression, demonstrate remarkable performance, with SSIM very close to 1 and PSNR values exceeding 50dB.
Gennaro Di Meo, Antonio G. M. Strollo, Davide De Caro
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 A Novel Module-Sign Low-Power Implementation for the DLMS Adaptive Filter With Low Steady-State Error
abstract
In this paper, a novel implementation is proposed for the Delayed LMS (DLMS) filter, able to reduce the power dissipation while preserving regime performances. The approach relies on the observation that the error signal is small in magnitude and oscillates around zero when the circuit is close to the convergence point. Therefore, the most significant bits of the error signal continuously toggle from positive to negative values causing high switching activity in the multipliers of the feedback section. This paper proposes to employ a sign-modulus representation of the error signal, to substantially reduce the switching activity of the feedback path of the filter. Additional approximation techniques are also devised to further reduce power dissipation. Comparisons with the state-of-the-art show that the proposed filter is the only one able to approach the MSE of the exact implementation with a remarkable reduction of power dissipation. A test-chip in TSMC 28nm CMOS technology has been realized to experimentally verify the validity of our technique. The experimental results show the possibility of saving up to 45.4% of power consumption with respect to the exact implementation of the filter.
Gennaro Di Meo, Davide De Caro, Giacinto Paolo Saggese, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 Approximate Multipliers Using Static Segmentation: Error Analysis and Improvements
abstract
Approximate multipliers are used in error-tolerant applications, sacrificing the accuracy of results to minimize power or delay. In this paper we investigate approximate multipliers using static segmentation. In these circuits a set of$m$contiguous bits (a segment of$m$bits) is extracted from each of the two$n$-bits operand, the two segments are in input to a small$m\times m$internal multiplier whose output is suitably shifted to obtain the result. We investigate both signed and unsigned multipliers, and for the latter we propose a new segmentation approach. We also present simple and effective correction techniques that can significantly reduce the approximation error with reduced hardware costs. We perform a detailed comparison with previously proposed approximate multipliers, considering a hardware implementation in 28 nm technology. The comparison shows that static segmented multipliers with the proposed correction technique have the desirable characteristic of being on (or close to) the Pareto-optimal frontier for both power vs normalized mean error distance and power vs mean relative error distance trade-off plots. These multipliers, therefore, are promising candidates for applications where their error performance is acceptable. This is confirmed by the results obtained for image processing and image classification applications.
Antonio G. M. Strollo, Ettore Napoli, Davide De Caro, Nicola Petra, Giacinto Paolo Saggese, Gennaro Di Meo
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 A Binary Line Buffer Circuit Featuring Lossy Data Compression at Fixed Maximum Data Rate
abstract
Video processing requires an increasing amount of buffered data. The paper proposes a multi-line buffer circuit that stores compressed data thus saving logic and power. The lossy compression algorithm provides the output stream with a fixed, decided by the user, delay from the input stream. Further, the amount of memory of the compressed buffer can be designed to trade off the correctness of the output with the logic resources footprint. The circuit is tested in a computer vision processing chain as a binary line buffer for the stream of foreground pixels. The paper also proposes a multi-line buffer circuit that integrates a morphological operation. The latter, when compared against the state of the art and against the proposed multi-line lossy buffer, allows larger saving of logic resources (up to 75%) and power (up to 85%), with reduced penalty on video quality.
Ettore Napoli, Davide De Caro, Nicola Petra, Antonio G. M. Strollo
ISCAS2
2020 Low-Power Approximate Multiplier with Error Recovery using a New Approximate 4-2 Compressor
abstract
In this paper we propose an energy-efficient approximate multiplier which uses a new approximate 4-2 compressor. The proposed compressor has a low error probability and its error conditions can be easily detected. This, as previously shown in the literature, makes it possible to implement error recovery, when the compressor is used in the partial product reduction phase of a multiplier. Simulation results show that proposed approximate multipliers exhibit a sensible reduction in Mean Error Distance and in maximum Error Distance, compared to previous art. Application to an image processing task shows an improvement of about 8dB in peak signal-to-noise ratio. Implementation results in 28nm CMOS show that the electrical performance of multipliers designed with the novel circuit are close to the one obtained with previously proposed approximate compressors.
Antonio G. M. Strollo, Davide De Caro, Ettore Napoli, Nicola Petra, Gennaro Di Meo
ISCAS2
2018 On the Use of Approximate Multipliers in LMS Adaptive Filters
abstract
Approximate computing relaxes algorithm precision constraints to improve digital circuit performance. Adaptive filters based on least-mean-square (LMS) algorithm constitute a standard in many DSP applications. The LMS algorithm, being an approximation of the Wiener filter, is inherently imprecise, and constitutes a fertile ground to employ approximate hardware techniques with the additional challenge related to the presence of a feedback path for coefficients update. In this paper, approximate LMS adaptive filters are explored for the first time, by employing approximate multipliers. A system identification scenario is adopted to assess the algorithm behavior. The analysis reveals that the choice of the approximate multiplier topology should be carefully examined, otherwise the stability and convergence performance of the algorithm can be compromised. We propose a novel approximate multiplier able to reduce the power dissipation in adaptive LMS filters up to 29% with tolerable convergence error degradation.
Darjn Esposito, Gennaro Di Meo, Davide De Caro, Nicola Petra, Ettore Napoli, Antonio G. M. Strollo
ISCAS3
2017 On the use of approximate adders in carry-save multiplier-accumulators
abstract
Approximate computing improves digital circuit performance by relaxing the requirement of performing exact calculations. In this paper, we investigate the use of approximate adders in the final stage of a carry save multiplier-accumulator (MAC), designed for image filtering application. We propose a design flow based on synthesis tools, starting from HDL description. After a first step in which an exact carry-propagate adder is used, the synthesized netlist is simulated to extract the statistics of the terms summed in the carry-propagate adder. We then design the approximate adder, to meet the required error characteristics given the inputs statistics. The netlist is finally modified by substituting the exact adder with the approximate one, and a final synthesis and optimization is performed. The presented design example in 28nm CMOS shows that a 14% power gain can be obtained, with a limited image quality degradation.
Darjn Esposito, Davide De Caro, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
ISCAS2
2017 A SISO Register Circuit Tailored for Input Data with Low Transition Probability
abstract
The paper proposes a SISO register circuit, functionally equivalent to a Shift Register, that is the optimal design choice when the input data have a reduced transition probability. The proposed circuit obtains improved performances by only storing the transitions of the input data, thus saving logic and power.
Ettore Napoli, Gerardo Castellano, Davide De Caro, Darjn Esposito, Nicola Petra, Antonio G. M. Strollo
IEEE Trans. Computers3
2016 Approximate adder with output correction for error tolerant applications and Gaussian distributed inputs
abstract
Approximate computing is emerging as a new paradigm to improve digital circuit performance by relaxing the requirement of performing exact calculations. Approximate adders rely on the idea that for uniformly distributed inputs, long carry-propagation chains are rarely activated. Unfortunately, however, the above assumption on input signal statistics is not always verified; in this paper we focus on the case (often encountered in practical signal processing applications) when the inputs have a Gaussian distribution. We show that for Gaussian inputs the error probability of previously proposed approximate adders approaches 25% for low sigma values, which is much larger than the uniform case. On the basis of this analysis, we propose an approximate adder with a correction circuit that drastically reduces the error rate for Gaussian distributed operand s. In order to investigate the performance of our approach in a real application, simulated results for a simple audio processing system are reported. Implementation results in 65nm technology are also presented.
Darjn Esposito, Gerardo Castellano, Davide De Caro, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
ISCAS3
2015 Hardware implementation of a spatio-temporal average filter for real-time denoising of fluoroscopic images
Mariangela Genovese, Paolo Bifulco, Davide De Caro, Ettore Napoli, Nicola Petra, Maria Romano, Mario Cesarelli, Antonio G. M. Strollo
Integr.3
2014 Analysis and comparison of Direct Digital Frequency Synthesizers implemented on FPGA
Mariangela Genovese, Ettore Napoli, Davide De Caro, Nicola Petra, Antonio G. M. Strollo
Integr.3
2013 Glitch-Free NAND-Based Digitally Controlled Delay-Lines
abstract
The recently proposed NAND-based digitally controlled delay-lines (DCDL) present a glitching problem which may limit their employ in many applications. This paper presents a glitch-free NAND-based DCDL which overcame this limitation by opening the employ of NAND-based DCDLs in a wide range of applications. The proposed NAND-based DCDL maintains the same resolution and minimum delay of previously proposed NAND-based DCDL. The theoretical demonstration of the glitch-free operation of proposed DCDL is also derived in the paper. Following this analysis, three driving circuits for the delay control-bits are also proposed. Proposed DCDLs have been designed in a 90-nm CMOS technology and compared, in this technology, to the state-of-the-art. Simulation results show that novel circuits result in the lowest resolution, with a little worsening of the minimum delay with respect to the previously proposed DCDL with the lowest delay. Simulations also confirm the correctness of developed glitching model and sizing strategy. As example application, proposed DCDL is used to realize an All-digital spread-spectrum clock generator (SSCG). The employ of proposed DCDL in this circuit allows to reduce the peak-to-peak absolute output jitter of more than the 40% with respect to a SSCG using three-state inverter based DCDLs.
Davide De Caro
IEEE Trans. Very Large Scale Integr. Syst.1
2012 An Experimental Power-Lines Model for Digital ASICs Based on Transmission Lines
abstract
In this paper, we present a transmission-line-based model developed to accurately describe the power and ground-line interconnections of modern digital ASICs. The proposed model employs transmission lines as the core component to properly describe both the capacitive and inductive behavior of the metal lines. In addition, the nonlinear frequency dependence of the line resistance, due to the skin-effect, is modeled with an additional lumped model. The model is completely derived from measurement data and allows describing both in-house and third-party ASICs. High-frequency$S$-parameter measured data are used to benchmark the model. Finally, on-board voltage measurements of a Numonyx 64 Mbit flash memory are performed and compared with transistor-level simulations.
Maurizio Costagliola, Davide De Caro, Antonio Girardi, Roberto Izzi, Niccolò Rinaldi, Marco Spirito, Paolo Spirito
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Elementary Functions Hardware Implementation Using Constrained Piecewise-Polynomial Approximations
abstract
A novel technique for designing piecewise-polynomial interpolators for hardware implementation of elementary functions is investigated in this paper. In the proposed approach, the interval where the function is approximated is subdivided in equal length segments and two adjacent segments are grouped in a segment pair. Suitable constraints are then imposed between the coefficients of the two interpolating polynomials in each segment pair. This allows reducing the total number of stored coefficients. It is found that the increase in the approximation error due to constraints between polynomial coefficients can easily be overcome by increasing the fractional bits of the coefficients. Overall, compared with standard unconstrained piecewise-polynomial approximation having the same accuracy, the proposed method results in a considerable advantage in terms of the size of the lookup table needed to store polynomial coefficients. The calculus of the coefficients of constrained polynomials and the optimization of coefficients bit width is also investigated in this paper. Results for several elementary functions and target precision ranging from 12 to 42 bits are presented. The paper also presents VLSI implementation results, targeting a 90 nm CMOS technology, and using both direct and Horner architectures for constrained degree-1, degree-2, and degree-3 approximations.
Antonio G. M. Strollo, Davide De Caro, Nicola Petra
IEEE Trans. Computers2
2010 High-speed differential resistor ladder for A/D converters
abstract
This paper describes the implementation of a novel high-speed differential resistor ladder suited for A/D converters. The novel ladder yields, theoretically, up to a sixteen-fold reduction of the propagation delay with respect to the conventional differential ladder. Simulation results, for a BiCMOS 0.25μm technology, show that the novel ladder results in a fivefold increase of the maximum sampling frequency when employed to design a 8-bit Flash converter. A 70% higher speed is also highlighted when the ladder is employed in a Folding and Interpolating 8-bit converter.
Davide De Caro, Marino Coppola, Nicola Petra, Ettore Napoli, Antonio G. M. Strollo, Valeria Garofalo
ISCAS1
2010 A novel truncated squarer with linear compensation function
abstract
A truncated binary squarer is a squarer with a n bit input that produces a n bit output. The proposed design minimizes the mean square error of the squarer and results in a very simple and fast circuital implementation. The squarer, compared against state of the art circuits, provides a reduction of the mean square error ranging from 20% to 5%. At the same time, the proposed squarer is able to reduce the power dissipation, reduce the silicon area occupation, and increase the maximum working frequency. Implementations results are provided for a 0.18μm technology.
Valeria Garofalo, Marino Coppola, Davide De Caro, Ettore Napoli, Nicola Petra, Antonio G. M. Strollo
ISCAS3
2010 Fixed-width CSD multipliers with minimum mean square error
abstract
Many multimedia and DSP applications require fixed-width multipliers, in which input data and output results have the same bit width. In this paper we investigate fixed-width multipliers where one of the input operand is a constant, encoded using canonic signed digit (CSD) representation. This is a very important case in many practical applications such as the calculation of Fast Fourier Transform. In the paper we derive in closed form the expression of the compensation function giving the minimum mean square error for CSD fixed-width multiplier. On the basis of this analytical result, we propose a hardware efficient implementation of the multiplier. Fixed width CSD multipliers implemented with the approach presented in this paper are accurate and can be implemented by using a simple partial-product reduction tree followed by a fast adder, without requiring additional look-up tables. The proposed approach is general and is well suited for implementation in circuit synthesizers. Implementation results in 90 nm technology are presented, to demonstrate the effectiveness of the proposed technique.
Nicola Petra, Davide De Caro, Antonio G. M. Strollo, Valeria Garofalo, Ettore Napoli, Marino Coppola, Pietro Todisco
ISCAS2
2008 A high performance floating-point special function unit using constrained piecewise quadratic approximation
abstract
A special function unit, able to compute square root, reciprocal square root, logarithm and exponential functions is presented in this paper. The system supports single precision IEEE-754 floating-point standard and uses a novel constrained piecewise quadratic interpolation technique to approximate the implemented functions. The proposed approach allows to reduce look-up table size of 40% with respect to previously proposed techniques. The SFU has been implemented in a test chip in 0.18 mum CMOS. A maximum clock frequency of 420 MHz and a power dissipation of 160 mW@420 MHz have been measured.
Davide De Caro, Nicola Petra, Antonio G. M. Strollo
ISCAS1
2007 A Novel Architecture for Galois Fields GF(2^m) Multipliers Based on Mastrovito Scheme
abstract
In the paper a new GF(2^m) multiplier for standard basis representation is developed. Proposed multiplier implements the Mastrovito multiplication scheme and can be designed for every field GF(2^m). A minimum area implementation of the first block of Mastrovito multiplier and a high-speed delay-driven tree architecture for the second block of Mastrovito multiplier are employed in the new circuit. Multiplier complexity and delay are analytically evaluated for many polynomial classes. Timing and area occupation performances of the proposed multiplier are also calculated for many fields used in Reed-Solomon codes applications and compared with those of previously proposed solutions. The comparison shows that the proposed multiplier outperforms previous architectures for every considered GF(2^m) field. The effectiveness of the proposed solution in a real application is verified by implementing in a 0.25ìm CMOS technology the key equation solving block of a (255,239) Reed-Solomon decoder. The use of the proposed multiplier in this application results in a substantial speed improvement without any penalty in silicon area occupation.
Nicola Petra, Davide De Caro, Antonio G. M. Strollo
IEEE Trans. Computers2
2005 A novel high-speed sense-amplifier-based flip-flop
abstract
A new sense-amplifier-based flip-flop is presented. The output latch of the proposed circuit can be considered as an hybrid solution between the standard NAND-based set/reset latch and the NC-/sup 2/MOS approach. The proposed flip-flop provides ratioless design, reduced short-circuit power dissipation, and glitch-free operation. The simulation results, obtained for a 0.25-/spl mu/m technology, show improvements in the clock-to-output delay and the power dissipation with respect to the recently proposed high-speed flip-flops. The new circuit has been successfully employed in a high-speed direct digital frequency synthesizer chip, highlighting the effectiveness of the proposed flip-flop in high-speed standard cell-based applications.
Antonio G. M. Strollo, Davide De Caro, Ettore Napoli, Nicola Petra
IEEE Trans. Very Large Scale Integr. Syst.2
2000 New clock-gating techniques for low-power flip-flops
abstract
Two novel low power flip-flops are presented in the paper. Proposed flip-flops use new gating techniques that reduce power dissipation deactivating the clock signal. Presented circuits overcome the clock duty-cycle limitation of previously reported gated flip-flops. Circuit simulations with the inclusion of parasitics show that sensible power dissipation reduction is possible if input signal has reduced switching activity. A 16-bit counter is presented as a simple low power application.
Antonio G. M. Strollo, Ettore Napoli, Davide De Caro
ISLPED3