VLDB 2026 Research / reviewers in the wild / expert
Eduardo A. C. da Costa
dblp:39/451 · also Eduardo Antonio Cesar da Costa, Eduardo Antônio César da Costa, Eduardo Costa 0001
· DBLP profile ↗
40ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-0521-5898ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-efficient discrete Haar Wavelet Transform architectures exploring approximate adders for high-quality image compression and reconstruction
Carlos Eduardo Reis Urban, Morgana Macedo Azevedo da Rosa, Eduardo A. C. da Costa |
Future Gener. Comput. Syst. | 3 |
| 2025 | Low-Energy NTT and INTT Architectures for Image Encryption and DecryptionabstractThe number theoretic transform (NTT) and its inverse (INTT) are efficient mathematical tools for polynomial multiplication, making them highly suitable for cryptographic applications such as image encryption and decryption. This work presents novel hardware architectures for NTT and INTT, designed explicitly for energy-efficient image encryption. Our primary contributions include the introduction of an approximate radix-2mlogarithm (AxRLL-16) based on a leading one detector (LOD) and Radix-16 encoder, which optimizes modular reduction operations by significantly reducing computational complexity. The NTT and INTT proposals are synthesized under a 65nm technology and achieve substantial improvements in power, area, and energy efficiency. Compared to state-of-the-art designs, our NTT and INTT implementations exhibit over 96.13% areasavings and 96.88% power-savings, with energy consumption reduced by up to 207 times. Eloisa Barros, Leonardo Antonietti, Rodrigo Lopes, Morgana Macedo Azevedo da Rosa, Eduardo A. C. da Costa, Rafael Soares |
ISCAS | 5 |
| 2025 | Dynamically Reconfigurable Approximate Multiplier for Precision ControlabstractThis paper presents a novel architecture for an approximate multiplier (AxM) based on Leading One-Bit Approximation (LoBA), aimed at enhancing error-resilient applications through a quality-configurable multipliers (QCMs) approach. The proposed DR-LoBA design is a statically and dynamically reconfigurable LoBA multiplier that offers flexibility and adaptability to varying application requirements. Featuring 16 approximation levels, it achieves significant power savings of 13.2% to 72% compared to precise multipliers with the same bit-width when tested with random inputs. On average, DR-LoBA delivers 27% greater precision when compared to state-of-art truncation-based reconfigurable multiplier across all precision levels. When performed in an actual application using the Filtered-x Least Mean Square (FXLMS) filter in active noise cancellation, our DR-LoBA multiplier reduces power consumption by 7.6% to 42.7% by varying the precision during the process while maintaining a noise reduction level of just 0.11 to 1.94dB lower than the full-precision system. João M. Bedin, Pedro Tauã Lopes Pereira, Eduardo A. C. da Costa, Sergio Bampi |
ISCAS | 3 |
| 2025 | FALSAx: An Integrated Framework for Accuracy and Logic Synthesis Estimation of Approximate AddersabstractThis work proposes an integrated framework for accuracy and logic synthesis (LS) estimation of approximate adders (FALSAx). It represents a versatile and robust framework designed to estimate the accuracy, power, and area of various approximate adders (AxAs) for any input width (W) and K bits of approximation using machine learning (ML) models. FALSAx facilitates performance predictions and optimization for different AxAs configurations through meticulously curated datasets and ML-driven analysis. The framework’s capability to automatically generate Pareto fronts from estimated values aids in identifying optimal trade-offs among crucial metrics, providing essential insights for circuit design and optimization. The FALSAx includes four internal frameworks: FrAQ, PILSE, and FELSE, which estimates dynamic power, total leakage power, and area, with frequency variations automatically, and the FALED dataset of the FALSAx. As a case study, this work analyzed 16 types of AxAs on FALSAx: AMA-V, AxPPA, COPY, TRUNC, ETA, LOA, HOERAA, LDCA, LZTA, HEAA, M-HEAA, HERLOA, M-HERLOA, HOAANED, OLOCA, and SETA. The rigorous analysis provided by FALSAx revealed that HERLOA, M-HERLOA, M-HEAA, and AxPPA demonstrated superior accuracy metrics such as SSIM, NCC, MAE, and MRE. Furthermore, power analysis showed that AxPPA exhibited the best power efficiency for lower approximation bits ($K \leq 3$). At the same time, gate-free adders like COPY, TRUNC, AMA-V, LDCA, and LZTA were more power-efficient for higher approximation bits ($K \gt 3$). Area estimations indicated that AxPPA maintained competitive efficiency for lower approximation bits ($K \leq 5$), while TRUNC and LDCA were more efficient for higher bits ($K \gt 5$). Morgana Macedo Azevedo da Rosa, Leonardo Antonietti, Rodrigo Lopes, Eloisa Barros, Eduardo A. C. da Costa, Rafael Soares |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | RCU- 2m: A VLSI Radix- 2m Cubic UnitabstractCubic operations are among the most used arithmetic operations in many applications that demand higher order simultaneous operand computation, such as cryptography and bicubic polynomial interpolation. This article proposes a novel VLSI radix-$2^{m}$cubic unit (RCU-$2^{m}$) capable of processing cubic operations at m bits simultaneously, with m values of 2 (RCU-4), 3 (RCU-8), and 4 (RCU-16). RCU-16 emerges as the most area-efficient configuration, surpassing RCU-8 and notably outperforming RCU-4. In the 8-bit scenario, RCU-16 achieves remarkable area savings, surpassing the literature’s proposed cubic unit by$11.58\times $. Across all configurations, RCU-$2^{m}$consistently outperforms the automatically selected cube unit, with energy savings ranging from$1.04\times $to$2\times $. In application specific integrated circuit (ASIC) and field-programmable gate array (FPGA)-based analyses, RCU-16 consistently exhibits superior performance in both area and energy savings compared with RCU-4, RCU-8, and solutions from the literature. These findings emphasize the importance of adopting radix-$2^{m}$configurations, particularly RCU-16, for optimal energy-constrained VLSI applications. Eduardo A. C. da Costa, Morgana Macedo Azevedo da Rosa |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | ReAdapt-II: Energy-Quality Optimizations for VLSI Adaptive Filters Through Automatic Reconfiguration and Built-In Iterative DividersabstractAdaptive filters using least mean square (LMS) algorithms offer high precision, low complexity, and fast convergence, but choosing the correct algorithm can be difficult and time-consuming. In this brief, we present ReAdapt-II, a VLSI circuit that enhances energy efficiency in adaptive filters through automatic reconfiguration and built-in iterative dividers, optimizing the energy-quality (EQ) tradeoff. This design features a self-selecting, reconfigurable hardware system with four adaptive algorithms, integrating iterative-based dividers and reusing arithmetic operators. Our results show a minimum energy consumption reduction of 39.75%, a 66.61% reduction in the circuit area, and a maximum accuracy increase of 17.07% compared with the previous ReAdapt architecture. Pedro Tauã Lopes Pereira, Patrícia Ücker, Eduardo A. C. da Costa, Paulo F. Flores, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | VLSI Architectures of Approximate Arithmetic Units Applied to Parallel Sensors CalibrationabstractApproximate computing maximizes area and energy savings for a trade-off between quality and efficiency. Approximate arithmetic operators have emerged as an efficient alternative to design low-power VLSI circuits. This paper investigates the design of approximate arithmetic operator units used in the calibration procedure for radio astronomy light sensors — the so-called StEFCal (statistically efficient and fast calibration) method. The StEFCal algorithm comprises arithmetic operations like a divider, square-accumulate (SAC), and multiply-accumulate (MAC) units. The StEFCal circuit of this work explores the following arithmetic operators: i) two approximate squarer units from the literature, i.e., radix-4 (AxRSU) and SquASH, ii) two approximate iterative-based Newton-Raphson (NR) and Goldschmidt (GLD) dividers, iii) one approximate parallel prefix adder (AxPPA), and iv) a new approximate radix-4 multiplier (AxRMU), proposed in this work, explored in the StEFCal multiply-accumulate circuit design. The AxRSU utilizes the parameters$K1$and$K2$to represent the number of exact encoders for squarer- and conventional-partial products, respectively, subsequently replaced with approximate encoders. The same principle applies to AxRMU, where the parameter$K$indicates the number of exact encoders for conventional-partial products, subsequently exchanged with approximate encoders. We demonstrate the efficiency of StEFCal using the approximate arithmetic operators from the Pareto-optimal front that expresses the area- and power-quality trade-off. The results show that using the AxRSU with$K1=4$and$K2=6$, AxRMU, and AxPPA with$K=16$and NR with one iteration has an MSE equal to 89.98dB and offers up to$158\times $energy-savings compared to the exact StEFCal, and up to$25\times $more energy-savings and$3.33\times $area-savings compared with our previous work,$440\times $energy-savings compared to the accurate state-of-the-art, and$258\times $compared with the approximate state-of-the-art. Morgana Macedo Azevedo da Rosa, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | ReAdapt: A Reconfigurable Datapath for Runtime Energy-Quality Scalable Adaptive FiltersabstractThis paper proposes ReAdapt–a reconfigurable datapath architecture for scaling the energy-quality trade-off of adaptive filtering at runtime. The ReAdapt can dynamically select four adaptive filtering algorithms for gradating complexity levels during runtime by reconfiguring the processing flow in its datapath and by blocking the switching activity (e.g., reducing the CMOS dynamic power) of unused modules with data-gating. The ReAdapt proposal can scale the energy-quality trade-off by choosing the following four different levels of filter algorithms complexity: 1) least mean square (LMS); 2) partial update normalized LMS (PU-NLMS); 3) set-membership normalized LMS (SM-NLMS); 4) normalized LMS (NLMS). The ReAdapt architecture reuses common modules of each adaptive filter, resulting in a compact VLSI hardware implementation. The ReAdapt architecture operation is implemented in a case-study for interference mitigation for electroencephalogram (EEG) signal processing. The hardware synthesis results show an increase of 6.80 times in throughput and at least a reduction of 2.84 times in energy per operation compared with the state-of-the-art adaptive filters. This paper also investigates the benefits of dynamically reconfiguring the four ReAdapt operating modes at runtime for different levels of signal-to-noise ratio (SNR) for the processed signals. We also demonstrate that dynamically reconfiguring the ReAdapt operating modes during runtime results in an optimal energy-quality trade-off which is advantageous over the conventional single static mode. Pedro Tauã Lopes Pereira, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | AxPPA: Approximate Parallel Prefix AddersabstractAddition units are widely used in many computational kernels of several error-tolerant applications such as machine learning and signal, image, and video processing. Besides their use as stand-alone, additions are essential building blocks for other math operations such as subtraction, comparison, multiplication, squaring, and division. The parallel prefix adders (PPAs) is among the fastest adders. It represents a parallel prefix graph consisting of the carry operator nodes, called prefix operators (POs). The PPAs, in particular, are among the fastest adders because they optimize the parallelization of the carry generation ($G$) and propagation ($P$). In this work, we introduce approximate PPAs (AxPPAs) by exploiting approximations in the POs. To evaluate our proposal for approximate POs (AxPOs), we generate the following AxPPAs, consisting of a set of four PPAs: approximate Brent–Kung (AxPPA-BK), approximate Kogge–Stone (AxPPA-KS), Ladner-Fischer (AxPPA-LF), and Sklansky (AxPPA-SK). We compare four AxPPA architectures with energy-efficient approximate adders (AxAs) [i.e., Copy, error-tolerant adder I (ETAI), lower-part OR adder (LOA), and Truncation (trunc)]. We tested them generically in stand-alone cases and embedded them in two important signal processing application kernels: a sum of squared differences (SSDs) video accelerator and a finite impulse response (FIR) filter kernel. The AxPPA-LF provides a new Pareto front in both energy-quality and area-quality results compared to state-of-the-art energy-efficient AxAs. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | AxRSU: Approximate Radix-4 Squarer UnitabstractApproximate computing emerged as a design alternative to boost design efficiency by leveraging the intrinsic error resiliency of many applications. Several error-resilient and compute-intensive applications such as signal, image, and video processing, computer vision, and supervised machine learning perform mean squared error (MSE) estimation during the runtime demanding dedicated squarer logic units in their hardware accelerators. This work proposes an approximate Radix-4 squarer unit architecture (AxRSU). Our AxRSU proposal reduces the encoder complexity and the number of required partial products, which considerably boosts energy and circuit area savings. We demonstrate the AxRSU error-quality trade-off in an SSD (Sum Squared Difference) hardware accelerator as a case study targeting a video processing application. We offer a new Pareto front with eighth optimal AxRSU solutions ranging 52-97% of cross-correlation (i.e., accuracy) for savings of 15-47% in energy consumption and 12-32% in circuit area. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Jorge Castro-Godínez, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
ISCAS | 4 |
| 2022 | A Framework for Crossing Temperature-Induced Timing Errors Underlying Hardware Accelerators to the Algorithm and Application LayersabstractTemperature rising is an unavoidable effect on VLSI and has always been a critical issue in any system-on-chip – especially when targeting compute-intensive applications. This effect increases the delay in hardware accelerators, resulting in timing errors due to unsustainable clock frequency, whose impact must be carefully evaluated on design time to measure the performance degradation of the hardware accelerator. Further, a hardware operating at a higher temperature accelerates device aging, which incurs in more timing errors. This issue is usually addressed with the inclusion of timing guardbands that compensate for the deleterious effects of temperature, ensuring the hardware accelerator works within a reliable zone, i.e., without any timing errors caused by temperature effects at runtime. However, guardbands directly result in considerable performance and efficiency losses because the circuit will be clocked at a frequency lower than its full potential. Accelerators on edge devices often dismiss such guardbands to explore the full potential of the designed circuits, posing an enormous design challenge as this approach requires a careful evaluation of the impact of timing errors on the quality of the target applications. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating hardware errors. Yet, these algorithms have a dynamic behavior (i.e., closed-loop) where a timing error can be propagated, affecting subsequent steps. Measuring the degradation-induced errors in these applications is very challenging given that an accurate gate-level simulation to investigate degradation-induced timing errors needs to be coupled dynamically with a system-level simulator to unveil how induced errors in the underlying hardware ultimately impact the algorithm execution in the hardware accelerator.This is the first work to achieve this goal. State-of-the-art works have studied accelerators under timing-errors when removing (or narrowing) guardbands. However, their approach was suitableonly for open-loop hardware accelerators which are entirely agnostic of complex interactions of the algorithms. Unlike prior work, this paper investigates temperature- and aging-induced timing-errors in the joint accelerator-algorithm interactions and their runtime impacts. Our framework investigates aging effects across the different layers starting from transistor physics all the way up to the algorithm layer. The hardware accelerator employed as a case study in this work is the sum of absolute differences (SAD), which is the most compute-intensive accelerator on commercial video encoder for mobile applications. Our results demonstrate the runtime behavior impacts of three advanced block-matching algorithms of the video encoder in a joint operation by a SAD accelerator under timing-errors induced by temperature and aging effects considering a 14nm FinFET technology. Guilherme Paim, Hussam Amrouch, Leandro M. G. Rocha, Brunno Abreu, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Computers | 5 |
| 2022 | Energy-Quality Scalable Design Space Exploration of Approximate FFT Hardware ArchitecturesabstractThis paper presents a comprehensive design space exploration for boosting energy efficiency of a fast Fourier transform (FFT) VLSI accelerator, exploiting several approximate multipliers (AxM) combined with approximate adder (AxA) circuits. The FFT hardware herein presented consists of a fixed-point sequential architecture using a radix-2 butterfly with decimation in time. We explore a set of AxMs – namely Dynamic Range Unbiased (DRUM), Rounding-based Approximate (RoBA), leading one Bit-based Approximate (LoBA), and Truncated approach – jointly with the LOA, ETA-I, CopyA, CopyB, Trunc0, Trunc1 approximate adders. The approximate arithmetic operators are used in the butterfly kernel with exploration of the approximation levels (for the${L}$and${K}$least-significant bits, respectively, for the AxM and AxA), aiming at discovering the most energy-efficient configuration under a design-time QoR constraint. The mean square error and peak signal-to-noise ratio metrics define which approximate levels combining${L}$and${K}$variations will enable the FFT to process signals to generate spectrograms without significant losses. Our results show that the LoBA multiplier with$L$=8 together with the LOA, Trunc1 and Trunc0, at different approximation levels, provide most energy savings with controllable quality degradation, presenting a minimum decrease of 20.2% in power dissipation without degrading the spectrogram generation quality. Pedro Tauã Lopes Pereira, Patrícia Ücker, Guilherme da Costa Ferreira, Brunno Abreu, Guilherme Paim, Eduardo A. C. da Costa, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Bridging the Gap Between Voltage Over-Scaling and Joint Hardware Accelerator-Algorithm Closed-LoopabstractVoltage over-scaling (VOS) optimizes energy while causing timing errors due to an unsustainable clock frequency. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating such errors. VOS has never been investigated in hardware accelerators running closed-loop algorithms. As the errors impact most decisions and actions in the subsequent steps, closed-loops dynamically change the execution flow. Timing errors should be evaluated by an accurate gate-level simulation, but a large gap still remains: how these timing errors propagate from the underlying hardware all the way up to the entire algorithm run, where they just may degrade the performance and quality of service of the application at stake? This paper tackles this issue showing a framework for VOS investigation, embracing any kind of application. Our framework simulates the VOS-induced timing errors at gate-level, dynamically linking the hardware result with the algorithm and vice versa during the evolution of the runtime of the application. The state-of-the-art VOS literature for video encoding application fails to assess the ultimate impacts of VOS-induced timing errors, as current works open the encoding loops. Unlike those, our work investigates the ultimate impact of a hardware accelerator dynamically carrying through to the video encoder all VOS-induced timing errors and preserving the full compliance to the standard. We employ a parallel sum of absolute differences (SAD) hardware accelerator as a case study. We assess the performance of the overall encoder under varying timing guardbands. Next, it is demonstrated that, under VOS, the ultimate impact in compression efficiency is related to the video’s motion intensity. Additionally, the advantages of timing guardband controlled reduction are clearly quantified in our results by virtue of the framework. Reducing at maximum 9.5% the clock frequency, energy savings (up to 16.5% in energy/operation) are achieved in SAD for video compression. Guilherme Paim, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | On the Resiliency of NCFET Circuits Against Voltage Over-ScalingabstractApproximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds. Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Approximate Pruned and Truncated Haar Discrete Wavelet Transform VLSI Hardware for Energy-Efficient ECG Signal ProcessingabstractThe approximate computing paradigm emerged as a key alternative for trading off accuracy and energy efficiency. Error-tolerant applications, such as multimedia and signal processing, can process the information with lower-than-standard accuracy at the circuit level while still fulfilling a good and acceptable service quality at the application level. The automatic detection of R-peaks in an electrocardiogram (ECG) signal is the essential step preceding ECG processing and analysis. The Haar discrete wavelet transform (HDWT) is a low-complexity pre-processing filter suitable to detect ECG R-peaks in embedded systems like wearable devices, which are incredibly energy-constrained. This work presents an approximate HDWT hardware architecture for ECG processing at very high energy efficiency. Our best-proposal employing pruning within the approximate HDWT hardware architecture requires just seven additions. The use of a truncation technique to improve energy efficiency is also investigated herein by observing the evolution of the signal-to-noise ratio and the ultimate impact in the ECG peak-detection application. This research finds that our HDWT approximate hardware architecture proposal accepts higher truncation levels than the original HDWT. In summary: Our results show about 9 times energy reduction when combining our HDWT matrix approximation proposal with the pruning and the highest acceptable level of truncation while still maintaining the R-peak detection performance accuracy of 99.68% on average. Henrique Seidel, Morgana Macedo Azevedo da Rosa, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Architectural Exploration for Energy-Efficient Fixed-Point Kalman Filter VLSI DesignabstractEfficient Kalman filter (KF) designs for real-time mobile applications, such as nano-drones navigation, robots localization, spacecraft orbit control, GPS positioning, image recognition, and multisensor data fusion for wearable systems, are key technology goals. The KF is a compute-intensive kernel composed of consecutive complex matrix operations, like multiplications and matrix inversions. The most complex block in the KF is the Kalman gain (KG) function, which involves matrices inversion at each iteration, applying the determinant matrix calculation and division operations. In this article, we combine architectural solutions of different types, for which balancing conflicting low-power and high-performance requirements aiming at real-time KF processing is a key design issue. The key finding in our architectural exploration herein presented is that the KF architectures in semiparallel and sequential forms offer the best balance of circuit area size, power dissipation, and processing speed. Compared to the state-of-the-art solutions, our KF architecture is more efficient, with 2.8 times fewer arithmetic operators, requiring 3.3 times fewer clock cycles. The usefulness of the developed KF in digital signal processing (DSP) is shown herein by simulations of system identification, noise elimination, and state estimation applications. These figures highlight the results of the KF architecture: the speed of adaptation for the system identification applications with root mean square error (RMSE) of 0.01 after 12 samples, precision level in noise elimination applications with RMSE of 0.13, and reliability in state estimation processes with RMSE less than 10% of system peak response. Pedro Tauã Lopes Pereira, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | A Cross-Layer Gate-Level-to-Application Co-Simulation for Design Space Exploration of Approximate Circuits in HEVC Video EncodersabstractA cross-layer design space exploration (DSE) method based on a proposed co-simulation technique is presented herein. The proposed method is demonstrated evaluating the impacts on both coding efficiency and power dissipation of applying distinct approximate logic operators in a sum of absolute differences (SAD) kernel that accelerates an H.265/HEVC (high-efficiency video coding) encoder. The proposed method simulates the gate-level circuit dynamically inside the application, with realistic results of the impact of the adder-tree approximate logic implementation on both quality and encoder bit-rate results. A comprehensive DSE is shown herein, with 13 types of 6 classes of approximate adders in the SAD accelerator hardware blocks. Over 3,000 logic variants of approximations at gate-level were developed. Actual video sequences as inputs to the x265 software encoder are co-simulated, to dynamically capture the video motion-estimation (ME) behavior in the presence of logic approximations. While the prior art that only estimates the impact of the approximate logic on power, area, and quality on static designs with statistical assumptions, which are agnostic to the actual algorithm data-dependent behavior in the application, our method explores accurately the trade-off between power dissipation and coding efficiency dynamically over the entire HEVC encoding. Our approach shows that the lower-part-or and error-tolerant adder I approximate adders, as well as truncation-to-zero deliver better compression-power trade-offs, with substantial differences from the static analysis. Guilherme Paim, Leandro M. G. Rocha, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Maximizing Side Channel Attack-Resistance and Energy-Efficiency of the STTL Combining Multi-Vt Transistors with Current and Capacitance BalancingabstractSecure triple track logic (STTL) is a circuit-level countermeasure to differential power analysis (DPA) attacks based on dual-rail precharge logic (DPL). STTL is robust to attacks due to the delay insensitive topology characteristic that avoids the glitches generated by the different path delays, before the logic gate inputs stabilize. However, the main STTL drawbacks are the validation of timing-robustness and the unbalanced and asymmetric transistors arrangement that result in variable internal capacitances and different internal paths to the current flow behaviors. The main contribution of this work is a new STTL-based topology called MT-BSTTL that combines multi-threshold with a set of circuit balancing improvements on capacitance, current paths, and fan-in, aiming to maximize the energy-efficiency while still preserving the side-channel attack-resistance. Three basic logic gates were implemented using the proposed strategy and other secure transistor topologies, all using the TSMC 40 nm technology. Results show that MT-BSTTL outperforms all state-of-the-art logic styles in terms of robustness against DPA attacks. Comparing to the baseline STTL, the proposed MT-BSTTL is, at least, 50% faster, has 53.5% higher energy-efficient, and it is 44% more robust, incurring in a 40% circuit area penalty. Vitor G. Lima, Guilherme Paim, Leandro M. G. Rocha, Leomar S. da Rosa Jr., Felipe S. Marques 0001, Eduardo A. C. da Costa, Vinícius V. Camargo, Rafael Soares, Sergio Bampi |
ISCAS | 6 |
| 2014 | Implementation of power efficient multicore FFT datapaths by reordering the twiddle factorsabstractThis paper addresses the reordering of coefficients, i.e., twiddle factors in multicore FFT in order to obtain power efficient datapaths. The coefficients are divided in smaller ones into the different cores and they are reordered through the Improved Anedma heuristic-based algorithm. According to the characteristics of the FFT algorithms, which involve multiplications of input data with appropriate coefficients, the best ordering of these operations, into each core, can contribute for the reduction of the switching activity, what leads to the minimization of power consumption in the FFTs. Therefore, the appropriate ordering of coefficients in the different cores allows finding the best architecture in terms of both performance and power consumption. The FFT architectures were synthesized using SYNOPSYS Design Compiler for the XFAB 180 nm technology. The results show that it is possible to achieve power reduction in the FFTs close to 9%, on average, after reordering the twiddle factors in the different cores. Sidinei Ghissoni, Eduardo A. C. da Costa, Angelo Goncalves da Luz |
VLSI-SoC | 2 |
| 2013 | Design of Digit-Serial FIR Filters: Algorithms, Architectures, and a CAD ToolabstractIn the last two decades, many efficient algorithms and architectures have been introduced for the design of low-complexity bit-parallel multiple constant multiplications (MCM) operation which dominates the complexity of many digital signal processing systems. On the other hand, little attention has been given to the digit-serial MCM design that offers alternative low-complexity MCM operations albeit at the cost of an increased delay. In this paper, we address the problem of optimizing the gate-level area in digit-serial MCM designs and introduce high-level synthesis algorithms, design architectures, and a computer-aided design tool. Experimental results show the efficiency of the proposed optimization algorithms and of the digit-serial MCM architectures in the design of digit-serial MCM operations and finite impulse response filters. Levent Aksoy, Cristiano Lazzari, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Design of low-complexity digital finite impulse response filters on FPGAsabstractThe multiple constant multiplications (MCM) operation, which realizes the multiplication of a set of constants by a variable, has a significant impact on the complexity and performance of the digital finite impulse response (FIR) filters. Over the years, many high-level algorithms and design methods have been proposed for the efficient implementation of the MCM operation using only addition, subtraction, and shift operations. The main contribution of this paper is the introduction of a high-level synthesis algorithm that optimizes the area of the MCM operation and, consequently, of the FIR filter design, on field programmable gate arrays (FPGAs) by taking into account the implementation cost of each addition and subtraction operation in terms of the number of fundamental building blocks of FPGAs. It is observed from the experimental results that the solutions of the proposed algorithm yield less complex FIR filters on FPGAs with respect to those whose MCM part is implemented using prominent MCM algorithms and design methods. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
DATE | 2 |
| 2012 | Multiple tunable constant multiplications: Algorithms and applicationsabstractThe multiple constant multiplications (MCM) problem, that is defined as finding the minimum number of addition and subtraction operations required for the multiplication of multiple constants by an input variable, has been the subject of great interest since the complexity of many digital signal processing (DSP) systems is dominated by an MCM operation. This paper introduces a variant of the MCM problem, called multiple tunable constant multiplications (MTCM) problem, where each constant is not fixed as in the MCM problem, but can be selected from a set of possible constants. We present an exact algorithm that formalizes the MTCM problem as a 0--1 integer linear programming (ILP) problem when constants are defined under a number representation. We also introduce a local search method for the MTCM problem that includes an efficient MCM algorithm. Furthermore, we show that these techniques can be used to solve various optimization problems in finite impulse response (FIR) filter design and we apply them to one of these problems. Experimental results clearly show the efficiency of the proposed methods when compared to prominent algorithms designed for the MCM problem. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
ICCAD | 2 |
| 2012 | High-level algorithms for the optimization of gate-level area in digit-serial multiple constant multiplications
Levent Aksoy, Cristiano Lazzari, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
Integr. | 3 |
| 2012 | Optimization Algorithms for the Multiplierless Realization of Linear TransformsabstractThis article addresses the problem of finding the fewest numbers of addition and subtraction operations in the multiplication of a constant matrix with an input vector---a fundamental operation in many linear digital signal processing transforms. We first introduce an exact common subexpression elimination (CSE) algorithm that formalizes the minimization of the number of operations as a 0-1 integer linear programming problem. Since there are still instances that the proposed exact algorithm cannot handle due to the NP-completeness of the problem, we also introduce a CSE heuristic algorithm that iteratively finds the most common 2-term subexpressions with the minimum conflicts among the expressions. Furthermore, since the main drawback of CSE algorithms is their dependency on a particular number representation, we propose a hybrid algorithm that initially finds promising realizations of linear transforms using a numerical difference method, and then applies the proposed CSE algorithm to utilize the common subexpressions iteratively. The experimental results on a comprehensive set of instances indicate that the proposed approximate algorithms find competitive results with those of the exact CSE algorithm and obtain better solutions than the prominent, previously proposed, heuristics. It is also observed that our solutions yield significant area reductions in the design of linear transforms after circuit synthesis, compared to direct realizations of linear transforms. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | Design of low-power multiple constant multiplications using low-complexity minimum depth operationsabstractExisting optimization algorithms for the multiplierless realization of multiple constant multiplications (MCM) typically target the minimization of the number of addition and subtraction operations. Since power dissipation is directly related to the amount of hardware, some power reduction is indirectly achieved by these algorithms. However, in many cases, glitching plays an equally important role in defining the power consumption. This is specially true for arithmetic circuits, and in particular to MCM due to high logic depth and large number of re-convergent paths. This paper introduces exact algorithms that search the optimal area of an MCM design at gate-level where each constant multiplication is implemented in its minimum depth. Experimental results show that the proposed algorithms lead to MCM designs consuming significantly less power with respect to those obtained by the MCM algorithms. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Efficient shift-adds design of digit-serial multiple constant multiplicationsabstractBit-parallel realization of the multiplication of a variable by a set of constants using only addition, subtraction, and shift operations has been explored extensively over the years as large number of constant multiplications dominate the complexity of many digital signal processing systems. On the other hand, digit-serial architectures offer alternative low-complexity designs since digit-serial operators occupy less area and are independent of the data wordlength. This paper introduces an approximate algorithm that targets the optimization of gate-level area in digit-serial constant multiplications under the shift-adds architecture. Experimental results indicate that our approximate algorithm gives better solutions than the previously proposed algorithms in terms of area at gate-level and yields alternative low-complexity designs relatively to the bit-parallel design. It is also observed on digit-serial filter designs that the use of shift-adds architecture yields area reduction up to 43.6% with respect to designs that use generic digit-serial constant multipliers. Levent Aksoy, Cristiano Lazzari, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2011 | Optimization of area in digit-serial Multiple Constant Multiplications at gate-levelabstractThe last two decades have seen many efficient algorithms and architectures for the design of low-complexity bit-parallel Multiple Constant Multiplications (MCM) operation, that dominates the complexity of Digital Signal Processing (DSP) systems. On the other hand, digit-serial architectures offer alternative low-complexity designs, since digit-serial operators occupy less area and are independent of the data wordlength. This paper introduces the problem of designing a digit-serial MCM operation with minimal area at gate-level and presents the exact formalization of the area optimization problem as a 0-1 Integer Linear Programming (ILP) problem. Experimental results show the efficiency of the proposed algorithm and digit- serial MCM designs in terms of area at gate-level. Levent Aksoy, Cristiano Lazzari, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
ISCAS | 3 |
| 2011 | A hybrid algorithm for the optimization of area and delay in linear DSP transformsabstractThis paper addresses the problem of multiplierless realization of linear transforms using the fewest number of addition and subtraction operations and introduces a hybrid algorithm that incorporates a graph-based technique, called the difference method, and a Common Subexpression Elimination (CSE) algorithm. In the proposed algorithm, while the difference method extracts the most promising realizations of linear transforms in each iteration, the CSE algorithm achieves the most common minimum conflicting subexpressions in each solution of the difference method. This paper also describes how the hybrid algorithm can be modified in order to find a solution with the fewest number of operations under a delay constraint. The experimental results on a comprehensive set of instances show the efficiency of the hybrid algorithms, at both high-level and gate-level, in comparison to previously proposed algorithms. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
VLSI-SoC | 2 |
| 2010 | Optimization of Area and Delay at Gate-Level in Multiple Constant MultiplicationsabstractAlthough many efficient high-level algorithms have been proposed for the realization of Multiple Constant Multiplications (MCM) using the fewest number of addition and subtraction operations, they do not consider the low-level implementation issues that directly affect the area, delay, and power dissipation of the MCM design. In this paper, we initially present area efficient addition and subtraction architectures used in the design of the MCM operation. Then, we propose an algorithm that searches an MCM design with the smallest area taking into account the cost of each operation at gate-level. To address the area and delay tradeoff in MCM design, the proposed algorithm is improved to find the smallest area solution under a delay constraint. The experimental results show that the proposed algorithms yield low-complexity and high-speed MCM designs with respect to those obtained by the prominent algorithms designed for the optimization of the number of operations and the optimization of area at gate-level. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
DSD | 2 |
| 2010 | Fast forward and inverse transforms for the H.264/AVC standard using hierarchical adder compressorsabstractThis paper presents fast architectures for the forward and inverse transforms of the H.264/AVC video compression standard. These transforms can be computed exactly as in integer arithmetic, thus avoiding mismatch problems between the encoder and decoder. They are inserted into the T and T-1block of the H.264/AVC and they can be computed by using only additions and shifts. Since the transforms algorithms are composed by a large number of addition/subtraction, fast architectures for the 4×4 Discrete Cosine transforms and 4×4 and 2×2 Hadamard transforms are proposed using efficient hierarchical adder compressors. The designs were described in VHDL and mapped to TSMC 0.18μm CMOS standard cells. Experimental results show that the architectures using 8-2 adder compressor can reach high frequency operation, high throughput and they are more efficient than the solutions of the literature. João S. Altermann, Eduardo A. C. da Costa, Sergio Bampi |
VLSI-SoC | 2 |
| 2010 | Design of low-complexity and high-speed digital Finite Impulse Response filtersabstractIn this paper, we introduce a design methodology to implement low-complexity and high-speed digital Finite Impulse Response (FIR) filters. Since FIR filters suffer from a large number of constant multiplications, in the proposed method the constant multiplications are replaced by addition/subtraction and shift operations. Also, based on the design objective, i.e., low-complexity or high-speed, the addition/subtraction operations are implemented using Ripple Carry Adder (RCA) or Carry-Save Adder (CSA) architectures respectively. Furthermore, high-level algorithms designed for the optimization of the number of RCA and CSA blocks are used to reduce the complexity of the FIR filter. Thus, a Computer-Aided Design (CAD) tool that synthesizes low-complexity and high-speed FIR filters in a shift-adds architecture is developed. It is observed from the experimental results on FIR filter instances that the developed CAD tool can find better FIR filter designs in terms of area and delay than those obtained using efficient general multipliers. Diego Jaccottet, Eduardo A. C. da Costa, Levent Aksoy, Paulo F. Flores, José Monteiro 0001 |
VLSI-SoC | 2 |
| 2008 | Exact and Approximate Algorithms for the Optimization of Area and Delay in Multiple Constant MultiplicationsabstractThe main contribution of this paper is an exact common subexpression elimination algorithm for the optimum sharing of partial terms in multiple constant multiplications (MCMs). We model this problem as a Boolean network that covers all possible partial terms that may be used to generate the set of coefficients in the MCM instance. We cast this problem into a 0–1 integer linear programming (ILP) problem by requiring that the single output of this network is asserted while minimizing the number of gates representing operations in the MCM implementation that evaluate to one. A satisfiability (SAT)-based 0–1 ILP solver is used to obtain the exact solution. We argue that for many real problems, the size of the problem is within the capabilities of current SAT solvers. Because performance is often a primary design parameter, we describe how this algorithm can be modified to target the minimum area solution under a user-specified delay constraint. Additionally, we propose an approximate algorithm based on the exact approach with extremely competitive results. We have applied these algorithms on the design of digital filters and present a comprehensive set of results that evaluate ours and existing approximation schemes against exact solutions under different number representations and using different SAT solvers. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Optimization of Area in Digital FIR Filters using Gate-Level MetricsabstractIn the paper, we propose a new metric for the minimization of area in the generic problem of multiple constant multiplications, and demonstrate its effectiveness for digital FIR filters. Previous methods use the number of required additions or subtractions as a cost function. We make the observation that not all of these operations have the same design cost. In the proposed algorithm, a minimum area solution is obtained by considering area estimates for each operation. To this end, we introduce accurate hardware models for addition and subtraction operations in terms of gate-level metrics, under both signed and unsigned representations. Our algorithm not only computes the best design solution among those that have the same number of operations, but is also able to find better area solutions using a non-minimum number of operations. The results obtained by the proposed exact algorithm are compared with the results of the exact algorithm designed for the minimum number of operations on FIR filters and it is shown that the area of the design can be reduced by up to 18%. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
DAC | 2 |
| 2007 | A new array architecture for signed multiplication using Gray encoded radix-2m operands
Eduardo A. C. da Costa, José Monteiro 0001, Sergio Bampi |
Integr. | 1 |
| 2006 | Optimization of area under a delay constraint in digital filter synthesis using SAT-based integer linear programmingabstractIn this paper, we propose an exact algorithm for the problem of area optimization under a delay constraint in the synthesis of multiplierless FIR filters. To the best of our knowledge, the method presented in this paper is the only exact algorithm designed for this problem. We present the results of the algorithm on real-sized filter instances and compare with an improved version of a recently proposed exact algorithm designed for the minimization of area. We show that in many cases delay can be minimized without any area penalty. Additionally, we describe two approximate algorithms that can be applied to instances which cannot be solved, or take too long, with the exact algorithm. We show that these algorithms find similar solutions to the exact algorithm in less CPU time. Levent Aksoy, Eduardo A. C. da Costa, Paulo F. Flores, José Monteiro 0001 |
DAC | 2 |
| 2006 | A VHDL Generation Tool for Optimized Parallel FIR FiltersabstractThis paper presents generation tool and performance results on a method to minimize the amount of hardware needed to implement a parallel digital finite impulse response (FIR) filters for hardwired (fixed coefficients) implementation targeted for high performance. The generation tool employ a combination of two approaches: first, the reduction of the coefficients to n-power-of-two (NPT) terms, using cannonical signed digit (CSD) as an option, followed by common subexpression elimination (CSE) among multipliers. Synthesis results for a range of different filter specifications, using Quartus II FPGA synthesis tool and Cadence PKS standard cell synthesis tool are presented Vagner Santos Da Rosa, Eduardo A. C. da Costa, Sergio Bampi |
VLSI-SoC | 2 |
| 2005 | An exact algorithm for the maximal sharing of partial terms in multiple constant multiplicationsabstractIn this paper, we propose an exact algorithm that maximizes the sharing of partial terms in multiple constant multiplication (MCM) operations. We model this problem as a Boolean network that covers all possible partial terms which may be used to generate the set of coefficients in the MCM instance. The PIs to this network are shifted versions of the MCM input. An AND gate represents an adder or a subtracter, i.e., an AND gate generates a new partial term. All partial terms that have the same numerical value are ORed together. There is a single output which is an /spl and/ over all the coefficients in the MCM. We cast this problem into a 0-1 integer linear programming (ILP) problem by requiring that the output is asserted while minimizing the total number of AND gates that evaluate to one. A SAT-based solver is used to obtain the exact solution. We argue that for many real problems the size of the problem is within the capabilities of current SAT solvers. We present results using binary, CSD and MSD representations. Two main conclusions can be drawn from the results. One is that, in many cases, existing heuristics perform well, computing the best solution, or one close to it. The other is that the flexibility of the MSD representation does not have a significant impact in the solution obtained. Paulo F. Flores, José Monteiro 0001, Eduardo A. C. da Costa |
ICCAD | 3 |
| 2005 | A Comparison of Layout Implementations of Pipelined and Non-Pipelined Signed Radix-4 Array Multiplier and Modified Booth Multiplier Architectures
Leonardo Londero de Oliveira, Cristiano Santos, Daniel Lima Ferrão, Eduardo A. C. da Costa, José Monteiro 0001, João Baptista dos Santos Martins, Sergio Bampi, Ricardo Augusto da Luz Reis |
VLSI-SoC | 4 |
| 2003 | Gray Encoded Arithmetic Operators Applied to FFT and FIR Dedicated Datapaths
Eduardo A. C. da Costa, José Monteiro 0001, Sergio Bampi |
VLSI-SOC | 1 |
| 2002 | A New Architecture for Signed Radix-2m Pure Array MultipliersabstractWe present a new architecture for signed multiplication which maintains the pure form of an array multiplier, exhibiting a much lower overhead than the Booth architecture. This architecture is extended for radix-2/sup m/ encoding, which leads to a reduction of the number of partial lines, enabling a significant improvement in performance and power consumption. The flexibility of our architecture allows for the easy construction of multipliers for different values of m, as opposed to the Booth architecture for which implementations for m > 2 are complex. The results we present show that the proposed architecture with radix-4 compares favorably in performance and power with the Modified Booth multiplier. We have experimented our architecture with different values of m and concluded that m = 4 minimizes both delay and power. Eduardo A. C. da Costa, Sergio Bampi, José Monteiro 0001 |
ICCD | 1 |