EDBT 2026 Demo / reviewers in the wild / expert
Guilherme Paim
dblp:170/1424 · also Guilherme Pereira Paim
· DBLP profile ↗
22ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0001-7809-9563ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OpenGeMM: A Highly-Efficient GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory CouplingabstractDeep neural networks (DNNs) face significant challenges when deployed on resource-constrained extreme edge devices due to their computational and data-intensive nature. While standalone accelerators tailored for specific application scenarios suffer from inflexible control and limited programmability, generic hardware acceleration platforms coupled with RISC-V CPUs can enable high reusability and flexibility, yet typically at the expense of system-level efficiency and low utilization. Xiaoling Yi, Ryan Antonio, Joren Dumoulin, Jiacong Sun, Josse Van Delm, Guilherme Paim, Marian Verhelst |
ASP-DAC | 6 |
| 2025 | DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow AcceleratorsabstractDeep Neural Networks (DNNs) have achieved remarkable success across various intelligent tasks but encounter performance and energy challenges in inference execution due to data movement bottlenecks. We introduce DataMaestro, a versatile and efficient data streaming unit that brings the decoupled access/execute architecture to DNN dataflow accelerators to address this issue. DataMaestro supports flexible and programmable access patterns to accommodate diverse workload types and dataflows, incorporates fine-grained prefetch and addressing mode switching to mitigate bank conflicts, and enables customizable on-the-fly data manipulation to reduce memory footprints and access counts. We integrate five DataMaestros with a Tensor Core-like GeMM accelerator and a Quantization accelerator into a RISC-V host system for evaluation. The FPGA prototype and VLSI synthesis results demonstrate that DataMaestro helps the GeMM core achieve nearly 100% utilization, which is 1.05 $21.39 \times$ better than state-of-the-art solutions, while minimizing area and energy consumption to merely 6.43% and 15.06% of the total system. Xiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong, Guilherme Paim, Marian Verhelst |
DAC | 5 |
| 2025 | An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator SystemsabstractHeterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movement efficiency and encounter compatibility issues in hardware and software. This prevents a unified approach that balances performance and ease of use. To this end, we present SNAX, an open-source integrated HW-SW framework enabling efficient multi-accelerator platforms through a novel hybrid-coupling scheme, consisting of loosely coupled asynchronous control and tightly coupled data access. SNAX brings reusable hardware modules designed to enhance compute accelerator utilization, and its customizable MLIR-based compiler to automate key system management tasks, jointly enabling rapid development and deployment of customized multi-accelerator compute clusters. Through extensive experimentation, we demonstrate SNAX’s efficiency and flexibility in a low-power heterogeneous SoC. Accelerators can be easily integrated and programmed to achieve >10× improvement in neural network performance compared to other accelerator systems while maintaining accelerator utilization of >90% in full system operation. Ryan Antonio, Joren Dumoulin, Xiaoling Yi, Josse Van Delm, Yunhao Deng, Guilherme Paim, Marian Verhelst |
ISLPED | 6 |
| 2025 | On the Efficacy and Vulnerabilities of Logic Locking in Tree-Based Machine LearningabstractThe popularity and widespread usage of machine learning (ML) hardware have created challenges for its intellectual property (IP) protection. Logic locking is a widely used technique for IP protection but has received little attention in error-resilient applications such as ML hardware modules. This work investigates the effectiveness of logic locking when applied to tree-based ML circuits and reveals a critical vulnerability that undermines its effectiveness for single-label ML classifiers. We propose a logic locking scheme to eliminate the vulnerabilities in decision trees (DTs) and random forests (RFs) circuits. In our extensive simulation involving 16 DTs and 16 RFs, our solution consistently thwarts the vulnerability. We further evaluated the security of our approach by considering different obfuscation percentages and launching state-of-the-art oracle-less attacks on logic locking. Our method proves resilient, indicating that by fixing the identified vulnerability, we did not introduce new attack vectors. Further, our investigation indicates that DT/RF accelerators are significantly less vulnerable to oracle-less attacks compared to exact circuits. Overall, our work lays the foundation for future investigations into the effectiveness of logic locking for ML circuits. Brunno Abreu, Guilherme Paim, Lilas Alrahis, Paulo F. Flores, Ozgur Sinanoglu, Sergio Bampi, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | TreeGRNG: Binary Tree Gaussian Random Number Generator for Efficient Probabilistic AI HardwareabstractBayesian Neural Networks (BNNs) offer opportunities for greatly enhancing the trustworthiness of conventional neural networks by monitoring the uncertainties in decision-making. A significant drawback for BNN inference at the extreme edge, however, is the imperative need to incorporate Gaussian Random Number Generators (GRNG) within each neuron. State-of-the-art GRNG algorithms heavily depend on multiple arithmetic operations and the use of extensive look-up tables, posing significant implementation challenges for ultra-low power hardware implementations. To overcome this, this paper presents an innovative binary tree random number generator (TreeGRNG) allowing the use of ultra-low-cost constant comparators instead of arithmetic units. We further enhance the TreeGRNG proposal with a set of hardware-aware optimizations exploiting the Gaussian properties. The optimized TreeGRNG surpasses the State-of-the-Art (SoTA) in terms of distribution accuracy while achieving a 3.7 × reduction in energy per sample and boosting the throughput per unit area by 5.8×. Moreover, our TreeGRNG proposal possesses a distinct advantage over the current SoTA in terms of flexibility, as it easily enables designers to adjust the shape of the sampled probability distribution, extending beyond the capabilities of traditional GRNGs, opening the horizon towards future probabilistic AI designs. The TreeGRNG design is available open-source in the link11https://github.com/KULeuven-MICAS/TreeGRNG. Jonas Crols, Guilherme Paim, Shirui Zhao, Marian Verhelst |
DATE | 2 |
| 2023 | ReAdapt: A Reconfigurable Datapath for Runtime Energy-Quality Scalable Adaptive FiltersabstractThis paper proposes ReAdapt–a reconfigurable datapath architecture for scaling the energy-quality trade-off of adaptive filtering at runtime. The ReAdapt can dynamically select four adaptive filtering algorithms for gradating complexity levels during runtime by reconfiguring the processing flow in its datapath and by blocking the switching activity (e.g., reducing the CMOS dynamic power) of unused modules with data-gating. The ReAdapt proposal can scale the energy-quality trade-off by choosing the following four different levels of filter algorithms complexity: 1) least mean square (LMS); 2) partial update normalized LMS (PU-NLMS); 3) set-membership normalized LMS (SM-NLMS); 4) normalized LMS (NLMS). The ReAdapt architecture reuses common modules of each adaptive filter, resulting in a compact VLSI hardware implementation. The ReAdapt architecture operation is implemented in a case-study for interference mitigation for electroencephalogram (EEG) signal processing. The hardware synthesis results show an increase of 6.80 times in throughput and at least a reduction of 2.84 times in energy per operation compared with the state-of-the-art adaptive filters. This paper also investigates the benefits of dynamically reconfiguring the four ReAdapt operating modes at runtime for different levels of signal-to-noise ratio (SNR) for the processed signals. We also demonstrate that dynamically reconfiguring the ReAdapt operating modes during runtime results in an optimal energy-quality trade-off which is advantageous over the conventional single static mode. Pedro Tauã Lopes Pereira, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | AxPPA: Approximate Parallel Prefix AddersabstractAddition units are widely used in many computational kernels of several error-tolerant applications such as machine learning and signal, image, and video processing. Besides their use as stand-alone, additions are essential building blocks for other math operations such as subtraction, comparison, multiplication, squaring, and division. The parallel prefix adders (PPAs) is among the fastest adders. It represents a parallel prefix graph consisting of the carry operator nodes, called prefix operators (POs). The PPAs, in particular, are among the fastest adders because they optimize the parallelization of the carry generation ($G$) and propagation ($P$). In this work, we introduce approximate PPAs (AxPPAs) by exploiting approximations in the POs. To evaluate our proposal for approximate POs (AxPOs), we generate the following AxPPAs, consisting of a set of four PPAs: approximate Brent–Kung (AxPPA-BK), approximate Kogge–Stone (AxPPA-KS), Ladner-Fischer (AxPPA-LF), and Sklansky (AxPPA-SK). We compare four AxPPA architectures with energy-efficient approximate adders (AxAs) [i.e., Copy, error-tolerant adder I (ETAI), lower-part OR adder (LOA), and Truncation (trunc)]. We tested them generically in stand-alone cases and embedded them in two important signal processing application kernels: a sum of squared differences (SSDs) video accelerator and a finite impulse response (FIR) filter kernel. The AxPPA-LF provides a new Pareto front in both energy-quality and area-quality results compared to state-of-the-art energy-efficient AxAs. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | AppGNN: Approximation-Aware Functional Reverse Engineering Using Graph Neural NetworksabstractThe globalization of the Integrated Circuit (IC) market is attracting an ever-growing number of partners, while remarkably lengthening the supply chain. Thereby, security concerns, such as those imposed by functional Reverse Engineering (RE), have become quintessential. RE leads to disclosure of confidential information to competitors, potentially enabling the theft of intellectual property. Traditional functional RE methods analyze a given gate-level netlist through employing pattern matching towards reconstructing the underlying basic blocks, and hence, reverse engineer the circuit's function. Tim Bücher, Lilas Alrahis, Guilherme Paim, Sergio Bampi, Ozgur Sinanoglu, Hussam Amrouch |
ICCAD | 3 |
| 2022 | AxRSU: Approximate Radix-4 Squarer UnitabstractApproximate computing emerged as a design alternative to boost design efficiency by leveraging the intrinsic error resiliency of many applications. Several error-resilient and compute-intensive applications such as signal, image, and video processing, computer vision, and supervised machine learning perform mean squared error (MSE) estimation during the runtime demanding dedicated squarer logic units in their hardware accelerators. This work proposes an approximate Radix-4 squarer unit architecture (AxRSU). Our AxRSU proposal reduces the encoder complexity and the number of required partial products, which considerably boosts energy and circuit area savings. We demonstrate the AxRSU error-quality trade-off in an SSD (Sum Squared Difference) hardware accelerator as a case study targeting a video processing application. We offer a new Pareto front with eighth optimal AxRSU solutions ranging 52-97% of cross-correlation (i.e., accuracy) for savings of 15-47% in energy consumption and 12-32% in circuit area. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Jorge Castro-Godínez, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
ISCAS | 2 |
| 2022 | A Framework for Crossing Temperature-Induced Timing Errors Underlying Hardware Accelerators to the Algorithm and Application LayersabstractTemperature rising is an unavoidable effect on VLSI and has always been a critical issue in any system-on-chip – especially when targeting compute-intensive applications. This effect increases the delay in hardware accelerators, resulting in timing errors due to unsustainable clock frequency, whose impact must be carefully evaluated on design time to measure the performance degradation of the hardware accelerator. Further, a hardware operating at a higher temperature accelerates device aging, which incurs in more timing errors. This issue is usually addressed with the inclusion of timing guardbands that compensate for the deleterious effects of temperature, ensuring the hardware accelerator works within a reliable zone, i.e., without any timing errors caused by temperature effects at runtime. However, guardbands directly result in considerable performance and efficiency losses because the circuit will be clocked at a frequency lower than its full potential. Accelerators on edge devices often dismiss such guardbands to explore the full potential of the designed circuits, posing an enormous design challenge as this approach requires a careful evaluation of the impact of timing errors on the quality of the target applications. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating hardware errors. Yet, these algorithms have a dynamic behavior (i.e., closed-loop) where a timing error can be propagated, affecting subsequent steps. Measuring the degradation-induced errors in these applications is very challenging given that an accurate gate-level simulation to investigate degradation-induced timing errors needs to be coupled dynamically with a system-level simulator to unveil how induced errors in the underlying hardware ultimately impact the algorithm execution in the hardware accelerator.This is the first work to achieve this goal. State-of-the-art works have studied accelerators under timing-errors when removing (or narrowing) guardbands. However, their approach was suitableonly for open-loop hardware accelerators which are entirely agnostic of complex interactions of the algorithms. Unlike prior work, this paper investigates temperature- and aging-induced timing-errors in the joint accelerator-algorithm interactions and their runtime impacts. Our framework investigates aging effects across the different layers starting from transistor physics all the way up to the algorithm layer. The hardware accelerator employed as a case study in this work is the sum of absolute differences (SAD), which is the most compute-intensive accelerator on commercial video encoder for mobile applications. Our results demonstrate the runtime behavior impacts of three advanced block-matching algorithms of the video encoder in a joint operation by a SAD accelerator under timing-errors induced by temperature and aging effects considering a 14nm FinFET technology. Guilherme Paim, Hussam Amrouch, Leandro M. G. Rocha, Brunno Abreu, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2022 | C2PAx: Complexity-Aware Constant Parameter Approximation for Energy-Efficient Tree-Based Machine Learning AcceleratorsabstractTree-based machine learning models, like random forests and decision trees, are low-complexity solutions that provide an efficient prediction for a wide range of applications. These models are particularly interesting for energy-constrained platforms since they can be implemented with simple logical operations. Tree-based accelerators are also intrinsically resilient to errors, and this can be leveraged to boost energy efficiency with approximate computing techniques. The key operations in these models are comparisons to constants, making comparators excellent candidates for approximation. This paper presents a technique to approximate comparisons to constants called C2PAx, which is capable of reducing the area and energy of tree-based accelerators. The method consists in finding alternative constants that reduce circuit area while keeping an efficient prediction performance. It is also shown that the selection of the constant parameters directly influences both hardware complexity and model performance, demanding cross-layer optimization. For that, we extend an existing framework that generates VLSI tree-based accelerators, inserting our approximation proposal that allows selecting the constant parameters that maximize energy efficiency at the cost of minor accuracy drops. Simulation results demonstrate that C2PAx outperforms the Don’t Care logic approximation technique when accuracy and energy are jointly considered. C2PAx trades accuracy for significant reductions in the VLSI area, power, delay, and energy consumption compared to precise models. Brunno Abreu, Guilherme Paim, Mateus Grellert, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Energy-Quality Scalable Design Space Exploration of Approximate FFT Hardware ArchitecturesabstractThis paper presents a comprehensive design space exploration for boosting energy efficiency of a fast Fourier transform (FFT) VLSI accelerator, exploiting several approximate multipliers (AxM) combined with approximate adder (AxA) circuits. The FFT hardware herein presented consists of a fixed-point sequential architecture using a radix-2 butterfly with decimation in time. We explore a set of AxMs – namely Dynamic Range Unbiased (DRUM), Rounding-based Approximate (RoBA), leading one Bit-based Approximate (LoBA), and Truncated approach – jointly with the LOA, ETA-I, CopyA, CopyB, Trunc0, Trunc1 approximate adders. The approximate arithmetic operators are used in the butterfly kernel with exploration of the approximation levels (for the${L}$and${K}$least-significant bits, respectively, for the AxM and AxA), aiming at discovering the most energy-efficient configuration under a design-time QoR constraint. The mean square error and peak signal-to-noise ratio metrics define which approximate levels combining${L}$and${K}$variations will enable the FFT to process signals to generate spectrograms without significant losses. Our results show that the LoBA multiplier with$L$=8 together with the LOA, Trunc1 and Trunc0, at different approximation levels, provide most energy savings with controllable quality degradation, presenting a minimum decrease of 20.2% in power dissipation without degrading the spectrogram generation quality. Pedro Tauã Lopes Pereira, Patrícia Ücker, Guilherme da Costa Ferreira, Brunno Abreu, Guilherme Paim, Eduardo A. C. da Costa, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Bridging the Gap Between Voltage Over-Scaling and Joint Hardware Accelerator-Algorithm Closed-LoopabstractVoltage over-scaling (VOS) optimizes energy while causing timing errors due to an unsustainable clock frequency. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating such errors. VOS has never been investigated in hardware accelerators running closed-loop algorithms. As the errors impact most decisions and actions in the subsequent steps, closed-loops dynamically change the execution flow. Timing errors should be evaluated by an accurate gate-level simulation, but a large gap still remains: how these timing errors propagate from the underlying hardware all the way up to the entire algorithm run, where they just may degrade the performance and quality of service of the application at stake? This paper tackles this issue showing a framework for VOS investigation, embracing any kind of application. Our framework simulates the VOS-induced timing errors at gate-level, dynamically linking the hardware result with the algorithm and vice versa during the evolution of the runtime of the application. The state-of-the-art VOS literature for video encoding application fails to assess the ultimate impacts of VOS-induced timing errors, as current works open the encoding loops. Unlike those, our work investigates the ultimate impact of a hardware accelerator dynamically carrying through to the video encoder all VOS-induced timing errors and preserving the full compliance to the standard. We employ a parallel sum of absolute differences (SAD) hardware accelerator as a case study. We assess the performance of the overall encoder under varying timing guardbands. Next, it is demonstrated that, under VOS, the ultimate impact in compression efficiency is related to the video’s motion intensity. Additionally, the advantages of timing guardband controlled reduction are clearly quantified in our results by virtue of the framework. Reducing at maximum 9.5% the clock frequency, energy savings (up to 16.5% in energy/operation) are achieved in SAD for video compression. Guilherme Paim, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | On the Resiliency of NCFET Circuits Against Voltage Over-ScalingabstractApproximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds. Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Approximate Pruned and Truncated Haar Discrete Wavelet Transform VLSI Hardware for Energy-Efficient ECG Signal ProcessingabstractThe approximate computing paradigm emerged as a key alternative for trading off accuracy and energy efficiency. Error-tolerant applications, such as multimedia and signal processing, can process the information with lower-than-standard accuracy at the circuit level while still fulfilling a good and acceptable service quality at the application level. The automatic detection of R-peaks in an electrocardiogram (ECG) signal is the essential step preceding ECG processing and analysis. The Haar discrete wavelet transform (HDWT) is a low-complexity pre-processing filter suitable to detect ECG R-peaks in embedded systems like wearable devices, which are incredibly energy-constrained. This work presents an approximate HDWT hardware architecture for ECG processing at very high energy efficiency. Our best-proposal employing pruning within the approximate HDWT hardware architecture requires just seven additions. The use of a truncation technique to improve energy efficiency is also investigated herein by observing the evolution of the signal-to-noise ratio and the ultimate impact in the ECG peak-detection application. This research finds that our HDWT approximate hardware architecture proposal accepts higher truncation levels than the original HDWT. In summary: Our results show about 9 times energy reduction when combining our HDWT matrix approximation proposal with the pruning and the highest acceptable level of truncation while still maintaining the R-peak detection performance accuracy of 99.68% on average. Henrique Seidel, Morgana Macedo Azevedo da Rosa, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Architectural Exploration for Energy-Efficient Fixed-Point Kalman Filter VLSI DesignabstractEfficient Kalman filter (KF) designs for real-time mobile applications, such as nano-drones navigation, robots localization, spacecraft orbit control, GPS positioning, image recognition, and multisensor data fusion for wearable systems, are key technology goals. The KF is a compute-intensive kernel composed of consecutive complex matrix operations, like multiplications and matrix inversions. The most complex block in the KF is the Kalman gain (KG) function, which involves matrices inversion at each iteration, applying the determinant matrix calculation and division operations. In this article, we combine architectural solutions of different types, for which balancing conflicting low-power and high-performance requirements aiming at real-time KF processing is a key design issue. The key finding in our architectural exploration herein presented is that the KF architectures in semiparallel and sequential forms offer the best balance of circuit area size, power dissipation, and processing speed. Compared to the state-of-the-art solutions, our KF architecture is more efficient, with 2.8 times fewer arithmetic operators, requiring 3.3 times fewer clock cycles. The usefulness of the developed KF in digital signal processing (DSP) is shown herein by simulations of system identification, noise elimination, and state estimation applications. These figures highlight the results of the KF architecture: the speed of adaptation for the system identification applications with root mean square error (RMSE) of 0.01 after 12 samples, precision level in noise elimination applications with RMSE of 0.13, and reliability in state estimation processes with RMSE less than 10% of system peak response. Pedro Tauã Lopes Pereira, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | A Cross-Layer Gate-Level-to-Application Co-Simulation for Design Space Exploration of Approximate Circuits in HEVC Video EncodersabstractA cross-layer design space exploration (DSE) method based on a proposed co-simulation technique is presented herein. The proposed method is demonstrated evaluating the impacts on both coding efficiency and power dissipation of applying distinct approximate logic operators in a sum of absolute differences (SAD) kernel that accelerates an H.265/HEVC (high-efficiency video coding) encoder. The proposed method simulates the gate-level circuit dynamically inside the application, with realistic results of the impact of the adder-tree approximate logic implementation on both quality and encoder bit-rate results. A comprehensive DSE is shown herein, with 13 types of 6 classes of approximate adders in the SAD accelerator hardware blocks. Over 3,000 logic variants of approximations at gate-level were developed. Actual video sequences as inputs to the x265 software encoder are co-simulated, to dynamically capture the video motion-estimation (ME) behavior in the presence of logic approximations. While the prior art that only estimates the impact of the approximate logic on power, area, and quality on static designs with statistical assumptions, which are agnostic to the actual algorithm data-dependent behavior in the application, our method explores accurately the trade-off between power dissipation and coding efficiency dynamically over the entire HEVC encoding. Our approach shows that the lower-part-or and error-tolerant adder I approximate adders, as well as truncation-to-zero deliver better compression-power trade-offs, with substantial differences from the static analysis. Guilherme Paim, Leandro M. G. Rocha, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Maximizing Side Channel Attack-Resistance and Energy-Efficiency of the STTL Combining Multi-Vt Transistors with Current and Capacitance BalancingabstractSecure triple track logic (STTL) is a circuit-level countermeasure to differential power analysis (DPA) attacks based on dual-rail precharge logic (DPL). STTL is robust to attacks due to the delay insensitive topology characteristic that avoids the glitches generated by the different path delays, before the logic gate inputs stabilize. However, the main STTL drawbacks are the validation of timing-robustness and the unbalanced and asymmetric transistors arrangement that result in variable internal capacitances and different internal paths to the current flow behaviors. The main contribution of this work is a new STTL-based topology called MT-BSTTL that combines multi-threshold with a set of circuit balancing improvements on capacitance, current paths, and fan-in, aiming to maximize the energy-efficiency while still preserving the side-channel attack-resistance. Three basic logic gates were implemented using the proposed strategy and other secure transistor topologies, all using the TSMC 40 nm technology. Results show that MT-BSTTL outperforms all state-of-the-art logic styles in terms of robustness against DPA attacks. Comparing to the baseline STTL, the proposed MT-BSTTL is, at least, 50% faster, has 53.5% higher energy-efficient, and it is 44% more robust, incurring in a 40% circuit area penalty. Vitor G. Lima, Guilherme Paim, Leandro M. G. Rocha, Leomar S. da Rosa Jr., Felipe S. Marques 0001, Eduardo A. C. da Costa, Vinícius V. Camargo, Rafael Soares, Sergio Bampi |
ISCAS | 2 |
| 2016 | High-throughput and memory-aware hardware of a sub-pixel interpolator for multiple video coding standardsabstractReal-time operation and low-power dissipation in video coding systems have become important research challenges, especially in mobile devices with limited battery and computational resources. There are many video coding standards coexisting in the market nowadays, so it is important for current devices to support different video coding standards. This paper presents a multi-standard luminance sub-samples interpolator hardware design for the Motion Compensation (MC) and Fractional Motion Estimation (FME), with support to MPEG-2/4, H.264/AVC, HEVC, and AVS/2 video coding standards. Our design is able to save hardware resources through an optimized filter organization, totally compliant with the focused standards and capable to interpolate samples for UHD 4320p@60fps at real time. The 45nm standard-cell library implementation dissipates 10mW, when processing according MPEG-2 standard, up to 46.4mW when processing AVS2. Guilherme Paim, Jones Goebel, Wagner Penny, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 1 |
| 2016 | An efficient sub-sample interpolator hardware for VP9-10 standardsabstractThis paper presents a hardware design for the sub-sample interpolator used in FME (Fractional Motion Estimation) and MC (Motion Compensation) stages according to the VP9 and VP10 video-coding standards. The proposed architecture is able to save hardware resources through an optimized-filter organization whereas reaching high-throughput and low-power dissipation. The hardware design was described in Verilog and synthesized for ASIC technology. The synthesis results were generated for 45nm Nangate standard cells and demonstrate that the developed architecture is able to process 2160p@60fps videos with a power dissipation of 2.34mW focusing on a VP9-10 decoder. Guilherme Paim, Wagner Penny, Jones Goebel, Vladimir Afonso, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 1 |
| 2016 | An HEVC multi-size DCT hardware with constant throughput and supporting heterogeneous CUsabstractThis paper presents an efficient hardware design for the Discrete Cosine Transform (DCT) of High Efficiency Video Coding standard (HEVC). This hardware supports all HEVC transform sizes: 4×4, 8×8, 16×16, and 32×32 including any combination of the Transform Unit (TU) sizes. The proposed DCT architecture has a constant throughput of 32 coefficients per cycle, independently of the transform sizes combination. The architecture was synthesized for a Nangate 45nm standard-cell library and the power analysis was made considering real input vectors. The synthesis results show a very good tradeoff between area, power dissipation and processing rates. The architecture is able to process 1.6G coeff/s when running at 50MHz dissipating 24.2 mW. These results allow a processing rate of 30 HD 1080p frames per second when evaluating 17 HEVC prediction modes. Jones Goebel, Guilherme Paim, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 2 |
| 2015 | A multi-standard interpolation hardware solution for H.264 and HEVCabstractAttending real-time constraints in video coding systems represents a big challenge for nowadays systems, especially for high definition videos at mobile systems. The Fractional Motion Estimation (FME) and Motion Compensation (MC) are responsible for a large share of processing effort in both state-of-the-art video coding standards, the High Efficiency Video Coding (HEVC), and its predecessor, the H.264. This work proposes a multi-standard hardware solution for the fractional sample interpolation used in FME/MC processing of the HEVC and H.264 standards. The hardware design is composed of four IP (Intellectual Property) cores able to process 1080p@60fps videos independently. The whole architecture can process 2160p@60fps with 80.69mW, considering bi-prediction. Henrique Maich, Guilherme Paim, Vladimir Afonso, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ICIP | 2 |