Patrícia Ücker

dblp:257/5308 · also Patrícia U. L. da Costa, Patrícia Ücker Leleu da Costa · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-5121-7101ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021
YearPublicationVenuePosition
2025 ReAdapt-II: Energy-Quality Optimizations for VLSI Adaptive Filters Through Automatic Reconfiguration and Built-In Iterative Dividers
abstract
Adaptive filters using least mean square (LMS) algorithms offer high precision, low complexity, and fast convergence, but choosing the correct algorithm can be difficult and time-consuming. In this brief, we present ReAdapt-II, a VLSI circuit that enhances energy efficiency in adaptive filters through automatic reconfiguration and built-in iterative dividers, optimizing the energy-quality (EQ) tradeoff. This design features a self-selecting, reconfigurable hardware system with four adaptive algorithms, integrating iterative-based dividers and reusing arithmetic operators. Our results show a minimum energy consumption reduction of 39.75%, a 66.61% reduction in the circuit area, and a maximum accuracy increase of 17.07% compared with the previous ReAdapt architecture.
Pedro Tauã Lopes Pereira, Patrícia Ücker, Eduardo A. C. da Costa, Paulo F. Flores, Sergio Bampi
IEEE Trans. Very Large Scale Integr. Syst.2
2024 VLSI Architectures of Approximate Arithmetic Units Applied to Parallel Sensors Calibration
abstract
Approximate computing maximizes area and energy savings for a trade-off between quality and efficiency. Approximate arithmetic operators have emerged as an efficient alternative to design low-power VLSI circuits. This paper investigates the design of approximate arithmetic operator units used in the calibration procedure for radio astronomy light sensors — the so-called StEFCal (statistically efficient and fast calibration) method. The StEFCal algorithm comprises arithmetic operations like a divider, square-accumulate (SAC), and multiply-accumulate (MAC) units. The StEFCal circuit of this work explores the following arithmetic operators: i) two approximate squarer units from the literature, i.e., radix-4 (AxRSU) and SquASH, ii) two approximate iterative-based Newton-Raphson (NR) and Goldschmidt (GLD) dividers, iii) one approximate parallel prefix adder (AxPPA), and iv) a new approximate radix-4 multiplier (AxRMU), proposed in this work, explored in the StEFCal multiply-accumulate circuit design. The AxRSU utilizes the parameters$K1$and$K2$to represent the number of exact encoders for squarer- and conventional-partial products, respectively, subsequently replaced with approximate encoders. The same principle applies to AxRMU, where the parameter$K$indicates the number of exact encoders for conventional-partial products, subsequently exchanged with approximate encoders. We demonstrate the efficiency of StEFCal using the approximate arithmetic operators from the Pareto-optimal front that expresses the area- and power-quality trade-off. The results show that using the AxRSU with$K1=4$and$K2=6$, AxRMU, and AxPPA with$K=16$and NR with one iteration has an MSE equal to 89.98dB and offers up to$158\times $energy-savings compared to the exact StEFCal, and up to$25\times $more energy-savings and$3.33\times $area-savings compared with our previous work,$440\times $energy-savings compared to the accurate state-of-the-art, and$258\times $compared with the approximate state-of-the-art.
Morgana Macedo Azevedo da Rosa, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 AxPPA: Approximate Parallel Prefix Adders
abstract
Addition units are widely used in many computational kernels of several error-tolerant applications such as machine learning and signal, image, and video processing. Besides their use as stand-alone, additions are essential building blocks for other math operations such as subtraction, comparison, multiplication, squaring, and division. The parallel prefix adders (PPAs) is among the fastest adders. It represents a parallel prefix graph consisting of the carry operator nodes, called prefix operators (POs). The PPAs, in particular, are among the fastest adders because they optimize the parallelization of the carry generation ($G$) and propagation ($P$). In this work, we introduce approximate PPAs (AxPPAs) by exploiting approximations in the POs. To evaluate our proposal for approximate POs (AxPOs), we generate the following AxPPAs, consisting of a set of four PPAs: approximate Brent–Kung (AxPPA-BK), approximate Kogge–Stone (AxPPA-KS), Ladner-Fischer (AxPPA-LF), and Sklansky (AxPPA-SK). We compare four AxPPA architectures with energy-efficient approximate adders (AxAs) [i.e., Copy, error-tolerant adder I (ETAI), lower-part OR adder (LOA), and Truncation (trunc)]. We tested them generically in stand-alone cases and embedded them in two important signal processing application kernels: a sum of squared differences (SSDs) video accelerator and a finite impulse response (FIR) filter kernel. The AxPPA-LF provides a new Pareto front in both energy-quality and area-quality results compared to state-of-the-art energy-efficient AxAs.
Morgana Macedo Azevedo da Rosa, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi
IEEE Trans. Very Large Scale Integr. Syst.3
2022 Energy-Quality Scalable Design Space Exploration of Approximate FFT Hardware Architectures
abstract
This paper presents a comprehensive design space exploration for boosting energy efficiency of a fast Fourier transform (FFT) VLSI accelerator, exploiting several approximate multipliers (AxM) combined with approximate adder (AxA) circuits. The FFT hardware herein presented consists of a fixed-point sequential architecture using a radix-2 butterfly with decimation in time. We explore a set of AxMs – namely Dynamic Range Unbiased (DRUM), Rounding-based Approximate (RoBA), leading one Bit-based Approximate (LoBA), and Truncated approach – jointly with the LOA, ETA-I, CopyA, CopyB, Trunc0, Trunc1 approximate adders. The approximate arithmetic operators are used in the butterfly kernel with exploration of the approximation levels (for the${L}$and${K}$least-significant bits, respectively, for the AxM and AxA), aiming at discovering the most energy-efficient configuration under a design-time QoR constraint. The mean square error and peak signal-to-noise ratio metrics define which approximate levels combining${L}$and${K}$variations will enable the FFT to process signals to generate spectrograms without significant losses. Our results show that the LoBA multiplier with$L$=8 together with the LOA, Trunc1 and Trunc0, at different approximation levels, provide most energy savings with controllable quality degradation, presenting a minimum decrease of 20.2% in power dissipation without degrading the spectrogram generation quality.
Pedro Tauã Lopes Pereira, Patrícia Ücker, Guilherme da Costa Ferreira, Brunno Abreu, Guilherme Paim, Eduardo A. C. da Costa, Sergio Bampi
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Architectural Exploration for Energy-Efficient Fixed-Point Kalman Filter VLSI Design
abstract
Efficient Kalman filter (KF) designs for real-time mobile applications, such as nano-drones navigation, robots localization, spacecraft orbit control, GPS positioning, image recognition, and multisensor data fusion for wearable systems, are key technology goals. The KF is a compute-intensive kernel composed of consecutive complex matrix operations, like multiplications and matrix inversions. The most complex block in the KF is the Kalman gain (KG) function, which involves matrices inversion at each iteration, applying the determinant matrix calculation and division operations. In this article, we combine architectural solutions of different types, for which balancing conflicting low-power and high-performance requirements aiming at real-time KF processing is a key design issue. The key finding in our architectural exploration herein presented is that the KF architectures in semiparallel and sequential forms offer the best balance of circuit area size, power dissipation, and processing speed. Compared to the state-of-the-art solutions, our KF architecture is more efficient, with 2.8 times fewer arithmetic operators, requiring 3.3 times fewer clock cycles. The usefulness of the developed KF in digital signal processing (DSP) is shown herein by simulations of system identification, noise elimination, and state estimation applications. These figures highlight the results of the KF architecture: the speed of adaptation for the system identification applications with root mean square error (RMSE) of 0.01 after 12 samples, precision level in noise elimination applications with RMSE of 0.13, and reliability in state estimation processes with RMSE less than 10% of system peak response.
Pedro Tauã Lopes Pereira, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi
IEEE Trans. Very Large Scale Integr. Syst.3