EDBT 2026 Demo / reviewers in the wild / expert
Morgana Macedo Azevedo da Rosa
dblp:222/9334 · also Morgana M. A. da Rosa, Morgana Macedo, Morgana da Rosa
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-9582-9011ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-efficient discrete Haar Wavelet Transform architectures exploring approximate adders for high-quality image compression and reconstruction
Carlos Eduardo Reis Urban, Morgana Macedo Azevedo da Rosa, Eduardo A. C. da Costa |
Future Gener. Comput. Syst. | 2 |
| 2025 | Low-Energy NTT and INTT Architectures for Image Encryption and DecryptionabstractThe number theoretic transform (NTT) and its inverse (INTT) are efficient mathematical tools for polynomial multiplication, making them highly suitable for cryptographic applications such as image encryption and decryption. This work presents novel hardware architectures for NTT and INTT, designed explicitly for energy-efficient image encryption. Our primary contributions include the introduction of an approximate radix-2mlogarithm (AxRLL-16) based on a leading one detector (LOD) and Radix-16 encoder, which optimizes modular reduction operations by significantly reducing computational complexity. The NTT and INTT proposals are synthesized under a 65nm technology and achieve substantial improvements in power, area, and energy efficiency. Compared to state-of-the-art designs, our NTT and INTT implementations exhibit over 96.13% areasavings and 96.88% power-savings, with energy consumption reduced by up to 207 times. Eloisa Barros, Leonardo Antonietti, Rodrigo Lopes, Morgana Macedo Azevedo da Rosa, Eduardo A. C. da Costa, Rafael Soares |
ISCAS | 4 |
| 2025 | FALSAx: An Integrated Framework for Accuracy and Logic Synthesis Estimation of Approximate AddersabstractThis work proposes an integrated framework for accuracy and logic synthesis (LS) estimation of approximate adders (FALSAx). It represents a versatile and robust framework designed to estimate the accuracy, power, and area of various approximate adders (AxAs) for any input width (W) and K bits of approximation using machine learning (ML) models. FALSAx facilitates performance predictions and optimization for different AxAs configurations through meticulously curated datasets and ML-driven analysis. The framework’s capability to automatically generate Pareto fronts from estimated values aids in identifying optimal trade-offs among crucial metrics, providing essential insights for circuit design and optimization. The FALSAx includes four internal frameworks: FrAQ, PILSE, and FELSE, which estimates dynamic power, total leakage power, and area, with frequency variations automatically, and the FALED dataset of the FALSAx. As a case study, this work analyzed 16 types of AxAs on FALSAx: AMA-V, AxPPA, COPY, TRUNC, ETA, LOA, HOERAA, LDCA, LZTA, HEAA, M-HEAA, HERLOA, M-HERLOA, HOAANED, OLOCA, and SETA. The rigorous analysis provided by FALSAx revealed that HERLOA, M-HERLOA, M-HEAA, and AxPPA demonstrated superior accuracy metrics such as SSIM, NCC, MAE, and MRE. Furthermore, power analysis showed that AxPPA exhibited the best power efficiency for lower approximation bits ($K \leq 3$). At the same time, gate-free adders like COPY, TRUNC, AMA-V, LDCA, and LZTA were more power-efficient for higher approximation bits ($K \gt 3$). Area estimations indicated that AxPPA maintained competitive efficiency for lower approximation bits ($K \leq 5$), while TRUNC and LDCA were more efficient for higher bits ($K \gt 5$). Morgana Macedo Azevedo da Rosa, Leonardo Antonietti, Rodrigo Lopes, Eloisa Barros, Eduardo A. C. da Costa, Rafael Soares |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | RCU- 2m: A VLSI Radix- 2m Cubic UnitabstractCubic operations are among the most used arithmetic operations in many applications that demand higher order simultaneous operand computation, such as cryptography and bicubic polynomial interpolation. This article proposes a novel VLSI radix-$2^{m}$cubic unit (RCU-$2^{m}$) capable of processing cubic operations at m bits simultaneously, with m values of 2 (RCU-4), 3 (RCU-8), and 4 (RCU-16). RCU-16 emerges as the most area-efficient configuration, surpassing RCU-8 and notably outperforming RCU-4. In the 8-bit scenario, RCU-16 achieves remarkable area savings, surpassing the literature’s proposed cubic unit by$11.58\times $. Across all configurations, RCU-$2^{m}$consistently outperforms the automatically selected cube unit, with energy savings ranging from$1.04\times $to$2\times $. In application specific integrated circuit (ASIC) and field-programmable gate array (FPGA)-based analyses, RCU-16 consistently exhibits superior performance in both area and energy savings compared with RCU-4, RCU-8, and solutions from the literature. These findings emphasize the importance of adopting radix-$2^{m}$configurations, particularly RCU-16, for optimal energy-constrained VLSI applications. Eduardo A. C. da Costa, Morgana Macedo Azevedo da Rosa |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | VLSI Architectures of Approximate Arithmetic Units Applied to Parallel Sensors CalibrationabstractApproximate computing maximizes area and energy savings for a trade-off between quality and efficiency. Approximate arithmetic operators have emerged as an efficient alternative to design low-power VLSI circuits. This paper investigates the design of approximate arithmetic operator units used in the calibration procedure for radio astronomy light sensors — the so-called StEFCal (statistically efficient and fast calibration) method. The StEFCal algorithm comprises arithmetic operations like a divider, square-accumulate (SAC), and multiply-accumulate (MAC) units. The StEFCal circuit of this work explores the following arithmetic operators: i) two approximate squarer units from the literature, i.e., radix-4 (AxRSU) and SquASH, ii) two approximate iterative-based Newton-Raphson (NR) and Goldschmidt (GLD) dividers, iii) one approximate parallel prefix adder (AxPPA), and iv) a new approximate radix-4 multiplier (AxRMU), proposed in this work, explored in the StEFCal multiply-accumulate circuit design. The AxRSU utilizes the parameters$K1$and$K2$to represent the number of exact encoders for squarer- and conventional-partial products, respectively, subsequently replaced with approximate encoders. The same principle applies to AxRMU, where the parameter$K$indicates the number of exact encoders for conventional-partial products, subsequently exchanged with approximate encoders. We demonstrate the efficiency of StEFCal using the approximate arithmetic operators from the Pareto-optimal front that expresses the area- and power-quality trade-off. The results show that using the AxRSU with$K1=4$and$K2=6$, AxRMU, and AxPPA with$K=16$and NR with one iteration has an MSE equal to 89.98dB and offers up to$158\times $energy-savings compared to the exact StEFCal, and up to$25\times $more energy-savings and$3.33\times $area-savings compared with our previous work,$440\times $energy-savings compared to the accurate state-of-the-art, and$258\times $compared with the approximate state-of-the-art. Morgana Macedo Azevedo da Rosa, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | AxPPA: Approximate Parallel Prefix AddersabstractAddition units are widely used in many computational kernels of several error-tolerant applications such as machine learning and signal, image, and video processing. Besides their use as stand-alone, additions are essential building blocks for other math operations such as subtraction, comparison, multiplication, squaring, and division. The parallel prefix adders (PPAs) is among the fastest adders. It represents a parallel prefix graph consisting of the carry operator nodes, called prefix operators (POs). The PPAs, in particular, are among the fastest adders because they optimize the parallelization of the carry generation ($G$) and propagation ($P$). In this work, we introduce approximate PPAs (AxPPAs) by exploiting approximations in the POs. To evaluate our proposal for approximate POs (AxPOs), we generate the following AxPPAs, consisting of a set of four PPAs: approximate Brent–Kung (AxPPA-BK), approximate Kogge–Stone (AxPPA-KS), Ladner-Fischer (AxPPA-LF), and Sklansky (AxPPA-SK). We compare four AxPPA architectures with energy-efficient approximate adders (AxAs) [i.e., Copy, error-tolerant adder I (ETAI), lower-part OR adder (LOA), and Truncation (trunc)]. We tested them generically in stand-alone cases and embedded them in two important signal processing application kernels: a sum of squared differences (SSDs) video accelerator and a finite impulse response (FIR) filter kernel. The AxPPA-LF provides a new Pareto front in both energy-quality and area-quality results compared to state-of-the-art energy-efficient AxAs. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | AxRSU: Approximate Radix-4 Squarer UnitabstractApproximate computing emerged as a design alternative to boost design efficiency by leveraging the intrinsic error resiliency of many applications. Several error-resilient and compute-intensive applications such as signal, image, and video processing, computer vision, and supervised machine learning perform mean squared error (MSE) estimation during the runtime demanding dedicated squarer logic units in their hardware accelerators. This work proposes an approximate Radix-4 squarer unit architecture (AxRSU). Our AxRSU proposal reduces the encoder complexity and the number of required partial products, which considerably boosts energy and circuit area savings. We demonstrate the AxRSU error-quality trade-off in an SSD (Sum Squared Difference) hardware accelerator as a case study targeting a video processing application. We offer a new Pareto front with eighth optimal AxRSU solutions ranging 52-97% of cross-correlation (i.e., accuracy) for savings of 15-47% in energy consumption and 12-32% in circuit area. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Jorge Castro-Godínez, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
ISCAS | 1 |
| 2021 | Approximate Pruned and Truncated Haar Discrete Wavelet Transform VLSI Hardware for Energy-Efficient ECG Signal ProcessingabstractThe approximate computing paradigm emerged as a key alternative for trading off accuracy and energy efficiency. Error-tolerant applications, such as multimedia and signal processing, can process the information with lower-than-standard accuracy at the circuit level while still fulfilling a good and acceptable service quality at the application level. The automatic detection of R-peaks in an electrocardiogram (ECG) signal is the essential step preceding ECG processing and analysis. The Haar discrete wavelet transform (HDWT) is a low-complexity pre-processing filter suitable to detect ECG R-peaks in embedded systems like wearable devices, which are incredibly energy-constrained. This work presents an approximate HDWT hardware architecture for ECG processing at very high energy efficiency. Our best-proposal employing pruning within the approximate HDWT hardware architecture requires just seven additions. The use of a truncation technique to improve energy efficiency is also investigated herein by observing the evolution of the signal-to-noise ratio and the ultimate impact in the ECG peak-detection application. This research finds that our HDWT approximate hardware architecture proposal accepts higher truncation levels than the original HDWT. In summary: Our results show about 9 times energy reduction when combining our HDWT matrix approximation proposal with the pruning and the highest acceptable level of truncation while still maintaining the R-peak detection performance accuracy of 99.68% on average. Henrique Seidel, Morgana Macedo Azevedo da Rosa, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |