EDBT 2026 Demo / reviewers in the wild / expert
Pedro Tauã Lopes Pereira
dblp:257/5293 · also Pedro T. L. Pereira
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-5231-3963ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamically Reconfigurable Approximate Multiplier for Precision ControlabstractThis paper presents a novel architecture for an approximate multiplier (AxM) based on Leading One-Bit Approximation (LoBA), aimed at enhancing error-resilient applications through a quality-configurable multipliers (QCMs) approach. The proposed DR-LoBA design is a statically and dynamically reconfigurable LoBA multiplier that offers flexibility and adaptability to varying application requirements. Featuring 16 approximation levels, it achieves significant power savings of 13.2% to 72% compared to precise multipliers with the same bit-width when tested with random inputs. On average, DR-LoBA delivers 27% greater precision when compared to state-of-art truncation-based reconfigurable multiplier across all precision levels. When performed in an actual application using the Filtered-x Least Mean Square (FXLMS) filter in active noise cancellation, our DR-LoBA multiplier reduces power consumption by 7.6% to 42.7% by varying the precision during the process while maintaining a noise reduction level of just 0.11 to 1.94dB lower than the full-precision system. João M. Bedin, Pedro Tauã Lopes Pereira, Eduardo A. C. da Costa, Sergio Bampi |
ISCAS | 2 |
| 2025 | ReAdapt-II: Energy-Quality Optimizations for VLSI Adaptive Filters Through Automatic Reconfiguration and Built-In Iterative DividersabstractAdaptive filters using least mean square (LMS) algorithms offer high precision, low complexity, and fast convergence, but choosing the correct algorithm can be difficult and time-consuming. In this brief, we present ReAdapt-II, a VLSI circuit that enhances energy efficiency in adaptive filters through automatic reconfiguration and built-in iterative dividers, optimizing the energy-quality (EQ) tradeoff. This design features a self-selecting, reconfigurable hardware system with four adaptive algorithms, integrating iterative-based dividers and reusing arithmetic operators. Our results show a minimum energy consumption reduction of 39.75%, a 66.61% reduction in the circuit area, and a maximum accuracy increase of 17.07% compared with the previous ReAdapt architecture. Pedro Tauã Lopes Pereira, Patrícia Ücker, Eduardo A. C. da Costa, Paulo F. Flores, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Cross-Layer Approximate Design of Low-Power Fractional Motion Estimation Accelerators for VVCabstractThe versatile video coding (VVC) standard introduces several innovative tools designed to enhance coding efficiency compared with its predecessors. One example is the adoption of an alternative filter for fractional motion estimation (FME) that is part of the adaptive motion vector resolution (AMVR) extension. While this allows a more precise motion representation, it also incurs more complexity for hardware implementations that aim at supporting most of VVC features. This work introduces a low-power hardware architecture accelerator specifically designed for FME with support for the AMVR extension of VVC. The proposed solution enables a systematic exploration of the design space of cross-layer approximate computing by combining approximations at both the operator and algorithm levels. This is achieved through the design of two novel architectures (2TAxA/4Tand2TAxA/2TAxA) alongside a newly proposed approximate filter applicable to both regular and alternative interpolation modes. Furthermore, we evaluate eight different approximate adder (AA) topologies to optimize power–quality tradeoffs. Experimental results demonstrate that for a complete FME multifilter interpolation unit (MIU) and maintaining an image quality threshold of$\text {SSIM} \geq 0.88$, our method achieves up to 72% power savings and 59.64% area savings. Rafael da Silva, Pedro Tauã Lopes Pereira, Mateus Grellert, Ricardo Augusto da Luz Reis |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | ReAdapt: A Reconfigurable Datapath for Runtime Energy-Quality Scalable Adaptive FiltersabstractThis paper proposes ReAdapt–a reconfigurable datapath architecture for scaling the energy-quality trade-off of adaptive filtering at runtime. The ReAdapt can dynamically select four adaptive filtering algorithms for gradating complexity levels during runtime by reconfiguring the processing flow in its datapath and by blocking the switching activity (e.g., reducing the CMOS dynamic power) of unused modules with data-gating. The ReAdapt proposal can scale the energy-quality trade-off by choosing the following four different levels of filter algorithms complexity: 1) least mean square (LMS); 2) partial update normalized LMS (PU-NLMS); 3) set-membership normalized LMS (SM-NLMS); 4) normalized LMS (NLMS). The ReAdapt architecture reuses common modules of each adaptive filter, resulting in a compact VLSI hardware implementation. The ReAdapt architecture operation is implemented in a case-study for interference mitigation for electroencephalogram (EEG) signal processing. The hardware synthesis results show an increase of 6.80 times in throughput and at least a reduction of 2.84 times in energy per operation compared with the state-of-the-art adaptive filters. This paper also investigates the benefits of dynamically reconfiguring the four ReAdapt operating modes at runtime for different levels of signal-to-noise ratio (SNR) for the processed signals. We also demonstrate that dynamically reconfiguring the ReAdapt operating modes during runtime results in an optimal energy-quality trade-off which is advantageous over the conventional single static mode. Pedro Tauã Lopes Pereira, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Energy-Quality Scalable Design Space Exploration of Approximate FFT Hardware ArchitecturesabstractThis paper presents a comprehensive design space exploration for boosting energy efficiency of a fast Fourier transform (FFT) VLSI accelerator, exploiting several approximate multipliers (AxM) combined with approximate adder (AxA) circuits. The FFT hardware herein presented consists of a fixed-point sequential architecture using a radix-2 butterfly with decimation in time. We explore a set of AxMs – namely Dynamic Range Unbiased (DRUM), Rounding-based Approximate (RoBA), leading one Bit-based Approximate (LoBA), and Truncated approach – jointly with the LOA, ETA-I, CopyA, CopyB, Trunc0, Trunc1 approximate adders. The approximate arithmetic operators are used in the butterfly kernel with exploration of the approximation levels (for the${L}$and${K}$least-significant bits, respectively, for the AxM and AxA), aiming at discovering the most energy-efficient configuration under a design-time QoR constraint. The mean square error and peak signal-to-noise ratio metrics define which approximate levels combining${L}$and${K}$variations will enable the FFT to process signals to generate spectrograms without significant losses. Our results show that the LoBA multiplier with$L$=8 together with the LOA, Trunc1 and Trunc0, at different approximation levels, provide most energy savings with controllable quality degradation, presenting a minimum decrease of 20.2% in power dissipation without degrading the spectrogram generation quality. Pedro Tauã Lopes Pereira, Patrícia Ücker, Guilherme da Costa Ferreira, Brunno Abreu, Guilherme Paim, Eduardo A. C. da Costa, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Architectural Exploration for Energy-Efficient Fixed-Point Kalman Filter VLSI DesignabstractEfficient Kalman filter (KF) designs for real-time mobile applications, such as nano-drones navigation, robots localization, spacecraft orbit control, GPS positioning, image recognition, and multisensor data fusion for wearable systems, are key technology goals. The KF is a compute-intensive kernel composed of consecutive complex matrix operations, like multiplications and matrix inversions. The most complex block in the KF is the Kalman gain (KG) function, which involves matrices inversion at each iteration, applying the determinant matrix calculation and division operations. In this article, we combine architectural solutions of different types, for which balancing conflicting low-power and high-performance requirements aiming at real-time KF processing is a key design issue. The key finding in our architectural exploration herein presented is that the KF architectures in semiparallel and sequential forms offer the best balance of circuit area size, power dissipation, and processing speed. Compared to the state-of-the-art solutions, our KF architecture is more efficient, with 2.8 times fewer arithmetic operators, requiring 3.3 times fewer clock cycles. The usefulness of the developed KF in digital signal processing (DSP) is shown herein by simulations of system identification, noise elimination, and state estimation applications. These figures highlight the results of the KF architecture: the speed of adaptation for the system identification applications with root mean square error (RMSE) of 0.01 after 12 samples, precision level in noise elimination applications with RMSE of 0.13, and reliability in state estimation processes with RMSE less than 10% of system peak response. Pedro Tauã Lopes Pereira, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |