VLDB 2026 Research / reviewers in the wild / expert
Sergio Bampi
dblp:84/5693
· DBLP profile ↗
123ranked-venue papers
2as first author
28since 2021 · last 2025
0000-0002-9018-6309ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 81 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 since 2021Software engineering, systems software and programming languages · 10 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamically Reconfigurable Approximate Multiplier for Precision ControlabstractThis paper presents a novel architecture for an approximate multiplier (AxM) based on Leading One-Bit Approximation (LoBA), aimed at enhancing error-resilient applications through a quality-configurable multipliers (QCMs) approach. The proposed DR-LoBA design is a statically and dynamically reconfigurable LoBA multiplier that offers flexibility and adaptability to varying application requirements. Featuring 16 approximation levels, it achieves significant power savings of 13.2% to 72% compared to precise multipliers with the same bit-width when tested with random inputs. On average, DR-LoBA delivers 27% greater precision when compared to state-of-art truncation-based reconfigurable multiplier across all precision levels. When performed in an actual application using the Filtered-x Least Mean Square (FXLMS) filter in active noise cancellation, our DR-LoBA multiplier reduces power consumption by 7.6% to 42.7% by varying the precision during the process while maintaining a noise reduction level of just 0.11 to 1.94dB lower than the full-precision system. João M. Bedin, Pedro Tauã Lopes Pereira, Eduardo A. C. da Costa, Sergio Bampi |
ISCAS | 4 |
| 2025 | A 365-µW 915-MHz & 2.45-GHz Flexible Mixer-First Discrete-Time-Receiver Front-EndabstractThis paper presents a highly flexible low-power mixer-first discrete-time (DT) receiver (RX). The combination of "mixer-first" and "discrete-time" designs realizes a new versatile architecture that can support multiple bands defined mainly by the clock generation. Here, the RX is tested for two ISM bands for IoT applications: 915 MHz and 2.45 GHz. The RX was fabricated in TSMC 28-nm LP CMOS and achieves competitive performance in both bands. Although consuming merely 365 µW, it occupies only 0.415 mm2despite the on-chip LC network. The DTRX achieves IIP3 of -4 dBm in-band, +10 dBm out-of-band, and noise figure (NF) below 15 dB without requiring any external components. These features make the architecture very appealing for IoT applications. Sandro Binsfeld Ferreira, Amir Bozorg, Filipe D. Baumgratz, Sergio Bampi |
ISCAS | 4 |
| 2025 | On the Efficacy and Vulnerabilities of Logic Locking in Tree-Based Machine LearningabstractThe popularity and widespread usage of machine learning (ML) hardware have created challenges for its intellectual property (IP) protection. Logic locking is a widely used technique for IP protection but has received little attention in error-resilient applications such as ML hardware modules. This work investigates the effectiveness of logic locking when applied to tree-based ML circuits and reveals a critical vulnerability that undermines its effectiveness for single-label ML classifiers. We propose a logic locking scheme to eliminate the vulnerabilities in decision trees (DTs) and random forests (RFs) circuits. In our extensive simulation involving 16 DTs and 16 RFs, our solution consistently thwarts the vulnerability. We further evaluated the security of our approach by considering different obfuscation percentages and launching state-of-the-art oracle-less attacks on logic locking. Our method proves resilient, indicating that by fixing the identified vulnerability, we did not introduce new attack vectors. Further, our investigation indicates that DT/RF accelerators are significantly less vulnerable to oracle-less attacks compared to exact circuits. Overall, our work lays the foundation for future investigations into the effectiveness of logic locking for ML circuits. Brunno Abreu, Guilherme Paim, Lilas Alrahis, Paulo F. Flores, Ozgur Sinanoglu, Sergio Bampi, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Improving Coding Efficiency of Massive Parallel Intra Prediction Using Alternative ReferencesabstractExploring massive parallelism is a common strategy to mitigate the processing time of modern video encoding standards. Nonetheless, data dependencies challenge parallelism exploitation, especially during intra prediction, where the reconstructed adjacent blocks are used as references. Some works use the original frame samples as references to decouple adjacent blocks and allow parallelism. Still, the original samples are static and cannot model the nuances of different bitrates. In this context, this work seeks to improve the coding efficiency of parallel intra prediction implementations by using alternative reference samples based on low-pass filters that better represent the nuances of different bitrates for any partitioning structure. Variations in multiple aspects of the filters are considered, such as their dimension and also the precision and distribution of their coefficients. Experimental evaluations assessed the similarity of such alternative samples when compared to the regular ones, in addition to their impacts on coding efficiency and the processing overhead required to obtain such samples. The results from such experiments demonstrate that the alternative references improve coding efficiency when compared to the original samples, especially at lower bitrates. Furthermore, the additional filtering stage poses negligible timing overhead in most computing systems. Iago Storch, Nuno Roma, Daniel Palomino 0001, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | ReAdapt-II: Energy-Quality Optimizations for VLSI Adaptive Filters Through Automatic Reconfiguration and Built-In Iterative DividersabstractAdaptive filters using least mean square (LMS) algorithms offer high precision, low complexity, and fast convergence, but choosing the correct algorithm can be difficult and time-consuming. In this brief, we present ReAdapt-II, a VLSI circuit that enhances energy efficiency in adaptive filters through automatic reconfiguration and built-in iterative dividers, optimizing the energy-quality (EQ) tradeoff. This design features a self-selecting, reconfigurable hardware system with four adaptive algorithms, integrating iterative-based dividers and reusing arithmetic operators. Our results show a minimum energy consumption reduction of 39.75%, a 66.61% reduction in the circuit area, and a maximum accuracy increase of 17.07% compared with the previous ReAdapt architecture. Pedro Tauã Lopes Pereira, Patrícia Ücker, Eduardo A. C. da Costa, Paulo F. Flores, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | VLSI Architectures of Approximate Arithmetic Units Applied to Parallel Sensors CalibrationabstractApproximate computing maximizes area and energy savings for a trade-off between quality and efficiency. Approximate arithmetic operators have emerged as an efficient alternative to design low-power VLSI circuits. This paper investigates the design of approximate arithmetic operator units used in the calibration procedure for radio astronomy light sensors — the so-called StEFCal (statistically efficient and fast calibration) method. The StEFCal algorithm comprises arithmetic operations like a divider, square-accumulate (SAC), and multiply-accumulate (MAC) units. The StEFCal circuit of this work explores the following arithmetic operators: i) two approximate squarer units from the literature, i.e., radix-4 (AxRSU) and SquASH, ii) two approximate iterative-based Newton-Raphson (NR) and Goldschmidt (GLD) dividers, iii) one approximate parallel prefix adder (AxPPA), and iv) a new approximate radix-4 multiplier (AxRMU), proposed in this work, explored in the StEFCal multiply-accumulate circuit design. The AxRSU utilizes the parameters$K1$and$K2$to represent the number of exact encoders for squarer- and conventional-partial products, respectively, subsequently replaced with approximate encoders. The same principle applies to AxRMU, where the parameter$K$indicates the number of exact encoders for conventional-partial products, subsequently exchanged with approximate encoders. We demonstrate the efficiency of StEFCal using the approximate arithmetic operators from the Pareto-optimal front that expresses the area- and power-quality trade-off. The results show that using the AxRSU with$K1=4$and$K2=6$, AxRMU, and AxPPA with$K=16$and NR with one iteration has an MSE equal to 89.98dB and offers up to$158\times $energy-savings compared to the exact StEFCal, and up to$25\times $more energy-savings and$3.33\times $area-savings compared with our previous work,$440\times $energy-savings compared to the accurate state-of-the-art, and$258\times $compared with the approximate state-of-the-art. Morgana Macedo Azevedo da Rosa, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | ReAdapt: A Reconfigurable Datapath for Runtime Energy-Quality Scalable Adaptive FiltersabstractThis paper proposes ReAdapt–a reconfigurable datapath architecture for scaling the energy-quality trade-off of adaptive filtering at runtime. The ReAdapt can dynamically select four adaptive filtering algorithms for gradating complexity levels during runtime by reconfiguring the processing flow in its datapath and by blocking the switching activity (e.g., reducing the CMOS dynamic power) of unused modules with data-gating. The ReAdapt proposal can scale the energy-quality trade-off by choosing the following four different levels of filter algorithms complexity: 1) least mean square (LMS); 2) partial update normalized LMS (PU-NLMS); 3) set-membership normalized LMS (SM-NLMS); 4) normalized LMS (NLMS). The ReAdapt architecture reuses common modules of each adaptive filter, resulting in a compact VLSI hardware implementation. The ReAdapt architecture operation is implemented in a case-study for interference mitigation for electroencephalogram (EEG) signal processing. The hardware synthesis results show an increase of 6.80 times in throughput and at least a reduction of 2.84 times in energy per operation compared with the state-of-the-art adaptive filters. This paper also investigates the benefits of dynamically reconfiguring the four ReAdapt operating modes at runtime for different levels of signal-to-noise ratio (SNR) for the processed signals. We also demonstrate that dynamically reconfiguring the ReAdapt operating modes during runtime results in an optimal energy-quality trade-off which is advantageous over the conventional single static mode. Pedro Tauã Lopes Pereira, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | AxPPA: Approximate Parallel Prefix AddersabstractAddition units are widely used in many computational kernels of several error-tolerant applications such as machine learning and signal, image, and video processing. Besides their use as stand-alone, additions are essential building blocks for other math operations such as subtraction, comparison, multiplication, squaring, and division. The parallel prefix adders (PPAs) is among the fastest adders. It represents a parallel prefix graph consisting of the carry operator nodes, called prefix operators (POs). The PPAs, in particular, are among the fastest adders because they optimize the parallelization of the carry generation ($G$) and propagation ($P$). In this work, we introduce approximate PPAs (AxPPAs) by exploiting approximations in the POs. To evaluate our proposal for approximate POs (AxPOs), we generate the following AxPPAs, consisting of a set of four PPAs: approximate Brent–Kung (AxPPA-BK), approximate Kogge–Stone (AxPPA-KS), Ladner-Fischer (AxPPA-LF), and Sklansky (AxPPA-SK). We compare four AxPPA architectures with energy-efficient approximate adders (AxAs) [i.e., Copy, error-tolerant adder I (ETAI), lower-part OR adder (LOA), and Truncation (trunc)]. We tested them generically in stand-alone cases and embedded them in two important signal processing application kernels: a sum of squared differences (SSDs) video accelerator and a finite impulse response (FIR) filter kernel. The AxPPA-LF provides a new Pareto front in both energy-quality and area-quality results compared to state-of-the-art energy-efficient AxAs. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | AppGNN: Approximation-Aware Functional Reverse Engineering Using Graph Neural NetworksabstractThe globalization of the Integrated Circuit (IC) market is attracting an ever-growing number of partners, while remarkably lengthening the supply chain. Thereby, security concerns, such as those imposed by functional Reverse Engineering (RE), have become quintessential. RE leads to disclosure of confidential information to competitors, potentially enabling the theft of intellectual property. Traditional functional RE methods analyze a given gate-level netlist through employing pattern matching towards reconstructing the underlying basic blocks, and hence, reverse engineer the circuit's function. Tim Bücher, Lilas Alrahis, Guilherme Paim, Sergio Bampi, Ozgur Sinanoglu, Hussam Amrouch |
ICCAD | 4 |
| 2022 | AxRSU: Approximate Radix-4 Squarer UnitabstractApproximate computing emerged as a design alternative to boost design efficiency by leveraging the intrinsic error resiliency of many applications. Several error-resilient and compute-intensive applications such as signal, image, and video processing, computer vision, and supervised machine learning perform mean squared error (MSE) estimation during the runtime demanding dedicated squarer logic units in their hardware accelerators. This work proposes an approximate Radix-4 squarer unit architecture (AxRSU). Our AxRSU proposal reduces the encoder complexity and the number of required partial products, which considerably boosts energy and circuit area savings. We demonstrate the AxRSU error-quality trade-off in an SSD (Sum Squared Difference) hardware accelerator as a case study targeting a video processing application. We offer a new Pareto front with eighth optimal AxRSU solutions ranging 52-97% of cross-correlation (i.e., accuracy) for savings of 15-47% in energy consumption and 12-32% in circuit area. Morgana Macedo Azevedo da Rosa, Guilherme Paim, Jorge Castro-Godínez, Eduardo A. C. da Costa, Rafael Soares, Sergio Bampi |
ISCAS | 6 |
| 2022 | Video Decoder Improvements with Near-Data Speculative Motion Compensation ProcessingabstractVideo decoder implementations are still evolving as they directly affect a large fraction of embedded systems nowadays. In this context, Versatile Video Coding (VVC) brings increased compression efficiency, which comes with extra over-head in terms of computational effort and energy consumption. At the same time, emerging Near-Data Processing (NDP) architectures promise drastic time and energy cuts for applications with data streaming behavior. In this paper, a speculative Motion Compensation (MC) is proposed to enable video decoders improvements through the exploitation of NDP. We adopted a large-vector SIMD-based NDP system (called VIMA) that provides high-performance operations over 2 K vectors. The proposed strategy leverages the correlation between the prediction modes and the motion data between spatially neighboring blocks within a frame to speculatively perform the MC for an entire region of 2Kx128 samples. MC interpolation kernels were implemented using VIMA and x86 AVX-256 SIMD libraries. Our NDP-based kernel implementation allows speedup of $1.9\times$ to $22\times$ compared to the x86 baseline solutions. Stepping forward, based on a coalescence estimation, our strategy can properly handle interpolation misses, achieving MC performance improvements from 7% to 64%. Garrenlus de Souza, José Rodrigo Azambuja, Bruno Zatt, Marco A. Z. Alves, Sergio Bampi, Felipe Sampaio |
ISCAS | 5 |
| 2022 | GPU-Acceleration of Affine Prediction in the Versatile Video CodingabstractThe Versatile Video Coding standard introduced a series of novel tools to improve the coding efficiency. However, these tools caused a massive increase in encoder computational complexity, and the affine prediction comprehends a significant share of this complexity. In this context, this work proposes an affine prediction modeling aiming at GPU implementation to accelerate the affine prediction by drawing the most parallelism out of such platforms. This modeling explores parallelism in two levels: conducting the prediction of multiple coding tree units simultaneously and breaking down the prediction into multiple highly-parallel stages. Since classical parallelization approaches for translational motion estimation are not very efficient for affine prediction, the proposed work explores the novel properties and parallelization possibilities introduced by affine prediction. Experimental results show that when applied to blocks 128x128, the proposed work can speed up the affine prediction by 57.21 times when compared to a fully sequential encoder, with a small coding efficiency penalty of 0.16% BD-BR. Iago Storch, Daniel Palomino 0001, Sergio Bampi |
ISCAS | 3 |
| 2022 | Mitigating Asynchronous QDI Drawbacks on MAC Operators with Approximate MultipliersabstractThis work explores the use of asynchronous approximate multiply-accumulate (MAC) operators and research ways to alleviate the inherent area overhead of such circuits, while leveraging on asynchronous circuits advantages. It analyzes three approximate MAC architectures with varying error rates and area trade-offs. Accurate and approximate, synchronous and asynchronous MAC operators are compared. Experiments show it is feasible to decrease the area overhead of the accurate asynchronous MAC from 8.1$\times$ down to 1.6$\times$ by recurring to approximate multipliers, for varying controlled error rates. The use of asynchronous quasi-delay insensitive (QDI) circuits allows applying extensive voltage scaling to all asynchronous MAC operators. Power and energy per operation can thus be significantly reduced, achieving savings of up to 2.66$\times$ in power and 3.17$\times$ in energy per operation when compared to the accurate asynchronous MAC. Rodrigo N. Wuerdig, Marcos L. L. Sartori, Brunno Abreu, Sergio Bampi, Ney Laert Vilar Calazans |
ISCAS | 4 |
| 2022 | Power-Throughput Trade-off Analysis for a Novel Multi-Boolean AV1 Arithmetic Encoder DesignabstractThe efficient and royalties-free video coding format AV1 was released in 2018 to cope with ultra-high-resolution video processing (e.g., 8K), avoiding the burden of royalty fees. The AV1 arithmetic encoding process, which is seen as a bottleneck due to its difficulty to work with parallelization, relies upon two primary operations: CDF Operation and Boolean Operation. This work presents a proposal and analysis of a novel fashion to parallelize the Boolean Operation of the AV1 arithmetic encoding block to achieve higher throughput rates, at the same time considering the power consumption to be at a compromise level. Furthermore, a novel design named MB-AV1, which utilizes the parallel Boolean approach and is able to accomplish 8K real-time video processing, is introduced. Tulio Pereira Bitencourt, Fábio Luís Livi Ramos, Sergio Bampi |
PCS | 3 |
| 2022 | Area and Power Efficient 8K Real-Time Design for AV1 Arithmetic DecodingabstractThe increasing demand for video media over limited bandwidth is currently an issue for both industry and academia. In the face of that challenge, AV1 is a recent video codec designed to compress videos efficiently. As the first step of AV1 playback, the codec utilizes a multi-alphabet arithmetic decoder as the entropy decoder kernel, with up to sixteen symbols. Hence, to accomplish a compromise trade-off between area, power, and real-time performance, we proposed MultiAD, an AV1 arithmetic decoder architecture. The presented design can decode a symbol of an alphabet with up to four elements in a single cycle, requiring more cycles for larger alphabets. When compared to the baseline prior-art architecture, MultiAD saves up to 58.3% in area and consumes 62.1% less power, while still accomplishing the 8K real-time video processing constraints. Jiovana Sousa Gomes, Tulio Pereira Bitencourt, Sergio Bampi, Fábio Luís Livi Ramos |
PCS | 3 |
| 2022 | A framework for designing power-efficient inference accelerators in tree-based learning applications
Brunno Abreu, Mateus Grellert, Sergio Bampi |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | A Framework for Crossing Temperature-Induced Timing Errors Underlying Hardware Accelerators to the Algorithm and Application LayersabstractTemperature rising is an unavoidable effect on VLSI and has always been a critical issue in any system-on-chip – especially when targeting compute-intensive applications. This effect increases the delay in hardware accelerators, resulting in timing errors due to unsustainable clock frequency, whose impact must be carefully evaluated on design time to measure the performance degradation of the hardware accelerator. Further, a hardware operating at a higher temperature accelerates device aging, which incurs in more timing errors. This issue is usually addressed with the inclusion of timing guardbands that compensate for the deleterious effects of temperature, ensuring the hardware accelerator works within a reliable zone, i.e., without any timing errors caused by temperature effects at runtime. However, guardbands directly result in considerable performance and efficiency losses because the circuit will be clocked at a frequency lower than its full potential. Accelerators on edge devices often dismiss such guardbands to explore the full potential of the designed circuits, posing an enormous design challenge as this approach requires a careful evaluation of the impact of timing errors on the quality of the target applications. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating hardware errors. Yet, these algorithms have a dynamic behavior (i.e., closed-loop) where a timing error can be propagated, affecting subsequent steps. Measuring the degradation-induced errors in these applications is very challenging given that an accurate gate-level simulation to investigate degradation-induced timing errors needs to be coupled dynamically with a system-level simulator to unveil how induced errors in the underlying hardware ultimately impact the algorithm execution in the hardware accelerator.This is the first work to achieve this goal. State-of-the-art works have studied accelerators under timing-errors when removing (or narrowing) guardbands. However, their approach was suitableonly for open-loop hardware accelerators which are entirely agnostic of complex interactions of the algorithms. Unlike prior work, this paper investigates temperature- and aging-induced timing-errors in the joint accelerator-algorithm interactions and their runtime impacts. Our framework investigates aging effects across the different layers starting from transistor physics all the way up to the algorithm layer. The hardware accelerator employed as a case study in this work is the sum of absolute differences (SAD), which is the most compute-intensive accelerator on commercial video encoder for mobile applications. Our results demonstrate the runtime behavior impacts of three advanced block-matching algorithms of the video encoder in a joint operation by a SAD accelerator under timing-errors induced by temperature and aging effects considering a 14nm FinFET technology. Guilherme Paim, Hussam Amrouch, Leandro M. G. Rocha, Brunno Abreu, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Computers | 6 |
| 2022 | C2PAx: Complexity-Aware Constant Parameter Approximation for Energy-Efficient Tree-Based Machine Learning AcceleratorsabstractTree-based machine learning models, like random forests and decision trees, are low-complexity solutions that provide an efficient prediction for a wide range of applications. These models are particularly interesting for energy-constrained platforms since they can be implemented with simple logical operations. Tree-based accelerators are also intrinsically resilient to errors, and this can be leveraged to boost energy efficiency with approximate computing techniques. The key operations in these models are comparisons to constants, making comparators excellent candidates for approximation. This paper presents a technique to approximate comparisons to constants called C2PAx, which is capable of reducing the area and energy of tree-based accelerators. The method consists in finding alternative constants that reduce circuit area while keeping an efficient prediction performance. It is also shown that the selection of the constant parameters directly influences both hardware complexity and model performance, demanding cross-layer optimization. For that, we extend an existing framework that generates VLSI tree-based accelerators, inserting our approximation proposal that allows selecting the constant parameters that maximize energy efficiency at the cost of minor accuracy drops. Simulation results demonstrate that C2PAx outperforms the Don’t Care logic approximation technique when accuracy and energy are jointly considered. C2PAx trades accuracy for significant reductions in the VLSI area, power, delay, and energy consumption compared to precise models. Brunno Abreu, Guilherme Paim, Mateus Grellert, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Energy-Quality Scalable Design Space Exploration of Approximate FFT Hardware ArchitecturesabstractThis paper presents a comprehensive design space exploration for boosting energy efficiency of a fast Fourier transform (FFT) VLSI accelerator, exploiting several approximate multipliers (AxM) combined with approximate adder (AxA) circuits. The FFT hardware herein presented consists of a fixed-point sequential architecture using a radix-2 butterfly with decimation in time. We explore a set of AxMs – namely Dynamic Range Unbiased (DRUM), Rounding-based Approximate (RoBA), leading one Bit-based Approximate (LoBA), and Truncated approach – jointly with the LOA, ETA-I, CopyA, CopyB, Trunc0, Trunc1 approximate adders. The approximate arithmetic operators are used in the butterfly kernel with exploration of the approximation levels (for the${L}$and${K}$least-significant bits, respectively, for the AxM and AxA), aiming at discovering the most energy-efficient configuration under a design-time QoR constraint. The mean square error and peak signal-to-noise ratio metrics define which approximate levels combining${L}$and${K}$variations will enable the FFT to process signals to generate spectrograms without significant losses. Our results show that the LoBA multiplier with$L$=8 together with the LOA, Trunc1 and Trunc0, at different approximation levels, provide most energy savings with controllable quality degradation, presenting a minimum decrease of 20.2% in power dissipation without degrading the spectrogram generation quality. Pedro Tauã Lopes Pereira, Patrícia Ücker, Guilherme da Costa Ferreira, Brunno Abreu, Guilherme Paim, Eduardo A. C. da Costa, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | Bridging the Gap Between Voltage Over-Scaling and Joint Hardware Accelerator-Algorithm Closed-LoopabstractVoltage over-scaling (VOS) optimizes energy while causing timing errors due to an unsustainable clock frequency. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating such errors. VOS has never been investigated in hardware accelerators running closed-loop algorithms. As the errors impact most decisions and actions in the subsequent steps, closed-loops dynamically change the execution flow. Timing errors should be evaluated by an accurate gate-level simulation, but a large gap still remains: how these timing errors propagate from the underlying hardware all the way up to the entire algorithm run, where they just may degrade the performance and quality of service of the application at stake? This paper tackles this issue showing a framework for VOS investigation, embracing any kind of application. Our framework simulates the VOS-induced timing errors at gate-level, dynamically linking the hardware result with the algorithm and vice versa during the evolution of the runtime of the application. The state-of-the-art VOS literature for video encoding application fails to assess the ultimate impacts of VOS-induced timing errors, as current works open the encoding loops. Unlike those, our work investigates the ultimate impact of a hardware accelerator dynamically carrying through to the video encoder all VOS-induced timing errors and preserving the full compliance to the standard. We employ a parallel sum of absolute differences (SAD) hardware accelerator as a case study. We assess the performance of the overall encoder under varying timing guardbands. Next, it is demonstrated that, under VOS, the ultimate impact in compression efficiency is related to the video’s motion intensity. Additionally, the advantages of timing guardband controlled reduction are clearly quantified in our results by virtue of the framework. Reducing at maximum 9.5% the clock frequency, energy savings (up to 16.5% in energy/operation) are achieved in SAD for video compression. Guilherme Paim, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | FastInter360: A Fast Inter Mode Decision for HEVC 360 Video CodingabstractThis paper presents FastInter360, a fast inter mode decision algorithm for accelerating the encoding of ERP 360 videos. The development of FastInter360 involves an in-depth and comprehensive set of evaluations performed to understand the differences in the encoder’s behavior when encoding 360 and conventional videos. These evaluations showed that due to the texture distortions resulting from projection, the encoder presents a specific behavior when encoding 360 videos, making it more likely to use a recurrent set of encoding modes when processing 360 videos. Besides, the coding efficiency is less sensible to approximations in some encoding steps depending on the frame region. FastInter360 is then proposed to reduce the encoding complexity by exploiting these differences. FastInter360 comprises three algorithms that accelerate the encoding by performing early decision by SKIP mode, reducing integer motion estimation search range, and adjusting fractional motion estimation precision. Furthermore, each of these algorithms behaves according to distortion intensity, performing greater complexity reduction in more distorted regions. When employed altogether, these algorithms compose FastInter360, which is able to achieve an average complexity reduction of 22.84% with a coding efficiency loss of 0.652% BD-BR, on average, making FastInter360 competitive with literature works. Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Logic Synthesis Meets Machine Learning: Trading Exactness for GeneralizationabstractLogic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning. Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee |
DATE | 27 |
| 2021 | Quality and Complexity Assessment of Learning-Based Image Compression SolutionsabstractThis work presents an analysis of state-of-the-art learning-based image compression techniques. We compare 8 models available in the Tensorflow Compression package in terms of visual quality metrics and processing time, using the KODAK data set. The results are compared with the Better Portable Graphics (BPG) and the JPEG2000 codecs. Results show that JPEG2000 has the lowest execution times compared with the fastest learning-based model, with a speedup of $1.46 \times$ in compression and $30 \times$ in decompression. However, the learning-based models achieved improvements over JPEG2000 in terms of quality, specially for lower bitrates. Our findings also show that BPG is more efficient in terms of PSNR, but the learning models are better for other quality metrics, and sometimes even faster. The results indicate that learning-based techniques are promising solutions towards a future mainstream compression method. João Dick, Brunno Abreu, Mateus Grellert, Sergio Bampi |
ICIP | 4 |
| 2021 | Fast Logic Optimization Using Decision TreesabstractThis work evaluates the use of Decision Trees (DTs) methods for a fast logic minimization of Boolean functions. The proposed DT approach is compared to traditional Espresso logic minimizer and the minimization algorithms available in the ABC tool. The methods are compared with respect to the execution time, number of nodes and number of logic levels. The DT methods proved to be a faster alternative, reducing time by an average of 52% and 5.5% when compared to Espresso and ABC respectively, while keeping competitive results in terms of AIG depth and number of nodes. Additionally, in order to obtain smaller circuits at the cost of approximate results we tested DTs with limited tree depth. The trade-offs between synthesis time, circuit area and accuracy are also discussed. Compared to ABC, limiting the maximum tree depth leads to time savings of up to 52%, up to 86% less number of nodes, and up to 48% lower AIG depth, while maintaining acceptable accuracy results. Brunno Abreu, Augusto Andre Souza Berndt, Isac de Souza Campos, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi |
ISCAS | 7 |
| 2021 | On the Resiliency of NCFET Circuits Against Voltage Over-ScalingabstractApproximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds. Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Approximate Pruned and Truncated Haar Discrete Wavelet Transform VLSI Hardware for Energy-Efficient ECG Signal ProcessingabstractThe approximate computing paradigm emerged as a key alternative for trading off accuracy and energy efficiency. Error-tolerant applications, such as multimedia and signal processing, can process the information with lower-than-standard accuracy at the circuit level while still fulfilling a good and acceptable service quality at the application level. The automatic detection of R-peaks in an electrocardiogram (ECG) signal is the essential step preceding ECG processing and analysis. The Haar discrete wavelet transform (HDWT) is a low-complexity pre-processing filter suitable to detect ECG R-peaks in embedded systems like wearable devices, which are incredibly energy-constrained. This work presents an approximate HDWT hardware architecture for ECG processing at very high energy efficiency. Our best-proposal employing pruning within the approximate HDWT hardware architecture requires just seven additions. The use of a truncation technique to improve energy efficiency is also investigated herein by observing the evolution of the signal-to-noise ratio and the ultimate impact in the ECG peak-detection application. This research finds that our HDWT approximate hardware architecture proposal accepts higher truncation levels than the original HDWT. In summary: Our results show about 9 times energy reduction when combining our HDWT matrix approximation proposal with the pruning and the highest acceptable level of truncation while still maintaining the R-peak detection performance accuracy of 99.68% on average. Henrique Seidel, Morgana Macedo Azevedo da Rosa, Guilherme Paim, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Energy-Throughput Configurable Design for Video Processing Binary Arithmetic EncoderabstractVideo encoding draws high research interest, due to the enormous demand for video traffic and real-time encoding for transmission. In video encoding standards such as HEVC (High-Efficiency Video Coding), the final step of the encoding stage is the CABAC (Context-Adaptive Binary Arithmetic Coding). The coding efficiency of the CABAC comes at the cost of increased computational complexity, especially for parallelization purposes, being the BAE (Binary Arithmetic Encoder) the critical part of CABAC. Thus, an important goal is to balance the real-time throughput requirements and the power/energy consumption in the design of BAE dedicated hardware. This work introduces a novel configurable high-throughput BAE design, named ET-BAE, utilizing a combination of a new modified Multiple-Bypass Bins Scheme (MBBS) and a power-saving approach into a single ASIC design with two-mode configuration. Synthesis and power-analysis results show that the configurable BAE design, the first of its kind with this feature, is more energy-efficient and less area consuming than utilizing non-configurable versions. The ET-BAE is able to accomplish the same real-time requirements of its competitors. Fábio Luís Livi Ramos, Bruno Zatt, Marcelo Schiavon Porto, Sergio Bampi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Architectural Exploration for Energy-Efficient Fixed-Point Kalman Filter VLSI DesignabstractEfficient Kalman filter (KF) designs for real-time mobile applications, such as nano-drones navigation, robots localization, spacecraft orbit control, GPS positioning, image recognition, and multisensor data fusion for wearable systems, are key technology goals. The KF is a compute-intensive kernel composed of consecutive complex matrix operations, like multiplications and matrix inversions. The most complex block in the KF is the Kalman gain (KG) function, which involves matrices inversion at each iteration, applying the determinant matrix calculation and division operations. In this article, we combine architectural solutions of different types, for which balancing conflicting low-power and high-performance requirements aiming at real-time KF processing is a key design issue. The key finding in our architectural exploration herein presented is that the KF architectures in semiparallel and sequential forms offer the best balance of circuit area size, power dissipation, and processing speed. Compared to the state-of-the-art solutions, our KF architecture is more efficient, with 2.8 times fewer arithmetic operators, requiring 3.3 times fewer clock cycles. The usefulness of the developed KF in digital signal processing (DSP) is shown herein by simulations of system identification, noise elimination, and state estimation applications. These figures highlight the results of the KF architecture: the speed of adaptation for the system identification applications with root mean square error (RMSE) of 0.01 after 12 samples, precision level in noise elimination applications with RMSE of 0.13, and reliability in state estimation processes with RMSE less than 10% of system peak response. Pedro Tauã Lopes Pereira, Guilherme Paim, Patrícia Ücker, Eduardo A. C. da Costa, Sérgio J. M. de Almeida, Sergio Bampi |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | VLSI Design of Tree-Based Inference for Low-Power Learning ApplicationsabstractThe use of Machine Learning techniques in battery-powered devices has increased in recent years. Therefore, evaluating power-accuracy trade-offs of inference models has become important. This work explores decision trees architectures, analyzing the effects of model complexity and approximation in power and accuracy. By quantizing the inputs to limited widths, we increased the accuracy of the models in up to 8.7% compared to precise cases. In terms of power, the variation in model complexity was more significant than in input width, as we obtained reductions of up to 97% by decreasing the tree depth, compared to 88% when decreasing the width. Brunno Abreu, Mateus Grellert, Sergio Bampi |
ISCAS | 3 |
| 2020 | Leveraging QDI Robustness to Simplify the Design of IoT CircuitsabstractInternet of Things devices require innovative power efficient design techniques that ensure correct operation in harsh environments, where using synchronous design can be challenging. The timing sign-off of synchronous circuits requires analysis and optimisation under multiple corners and operating modes. Considering that energy efficient circuits demand dynamic voltage ranges and harsh environments impose significant variations, design sign-off may become prohibitively expensive. An alternative is quasi-delay-insensitive asynchronous design, which presents robustness against timing variations, simplifying timing sign-off. This paper leverages recent developments in asynchronous circuits design automation to achieve higher degrees of energy efficiency using voltage scaling, while ensuring solid robustness to variability. Marcos L. L. Sartori, Rodrigo N. Wuerdig, Matheus T. Moreira, Sergio Bampi, Ney Laert Vilar Calazans |
ISCAS | 4 |
| 2020 | ERP-Based CTU Splitting Early Termination for Intra Prediction of 360 videosabstractThis work presents an Equirectangular projection (ERP) based Coding Tree Unit (CTU) splitting early termination algorithm for the High Efficiency Video Coding (HEVC) intra prediction of 360-degree videos. The proposed algorithm adaptively employs early termination in the HEVC CTU splitting based on distortion properties of the ERP projection, that generate homogeneous regions at the top and bottom portion of a video frame. Experimental results show an average of 24% time saving with 0.11% coding efficiency loss, significantly reducing the encoding complexity with minor impacts in the encoding efficiency. Besides, solution presents the best results considering the relation between time saving and coding efficiency when compared with all related works. Bernardo Beling, Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
VCIP | 5 |
| 2020 | A Cross-Layer Gate-Level-to-Application Co-Simulation for Design Space Exploration of Approximate Circuits in HEVC Video EncodersabstractA cross-layer design space exploration (DSE) method based on a proposed co-simulation technique is presented herein. The proposed method is demonstrated evaluating the impacts on both coding efficiency and power dissipation of applying distinct approximate logic operators in a sum of absolute differences (SAD) kernel that accelerates an H.265/HEVC (high-efficiency video coding) encoder. The proposed method simulates the gate-level circuit dynamically inside the application, with realistic results of the impact of the adder-tree approximate logic implementation on both quality and encoder bit-rate results. A comprehensive DSE is shown herein, with 13 types of 6 classes of approximate adders in the SAD accelerator hardware blocks. Over 3,000 logic variants of approximations at gate-level were developed. Actual video sequences as inputs to the x265 software encoder are co-simulated, to dynamically capture the video motion-estimation (ME) behavior in the presence of logic approximations. While the prior art that only estimates the impact of the approximate logic on power, area, and quality on static designs with statistical assumptions, which are agnostic to the actual algorithm data-dependent behavior in the application, our method explores accurately the trade-off between power dissipation and coding efficiency dynamically over the entire HEVC encoding. Our approach shows that the lower-part-or and error-tolerant adder I approximate adders, as well as truncation-to-zero deliver better compression-power trade-offs, with substantial differences from the static analysis. Guilherme Paim, Leandro M. G. Rocha, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Maximizing Side Channel Attack-Resistance and Energy-Efficiency of the STTL Combining Multi-Vt Transistors with Current and Capacitance BalancingabstractSecure triple track logic (STTL) is a circuit-level countermeasure to differential power analysis (DPA) attacks based on dual-rail precharge logic (DPL). STTL is robust to attacks due to the delay insensitive topology characteristic that avoids the glitches generated by the different path delays, before the logic gate inputs stabilize. However, the main STTL drawbacks are the validation of timing-robustness and the unbalanced and asymmetric transistors arrangement that result in variable internal capacitances and different internal paths to the current flow behaviors. The main contribution of this work is a new STTL-based topology called MT-BSTTL that combines multi-threshold with a set of circuit balancing improvements on capacitance, current paths, and fan-in, aiming to maximize the energy-efficiency while still preserving the side-channel attack-resistance. Three basic logic gates were implemented using the proposed strategy and other secure transistor topologies, all using the TSMC 40 nm technology. Results show that MT-BSTTL outperforms all state-of-the-art logic styles in terms of robustness against DPA attacks. Comparing to the baseline STTL, the proposed MT-BSTTL is, at least, 50% faster, has 53.5% higher energy-efficient, and it is 44% more robust, incurring in a 40% circuit area penalty. Vitor G. Lima, Guilherme Paim, Leandro M. G. Rocha, Leomar S. da Rosa Jr., Felipe S. Marques 0001, Eduardo A. C. da Costa, Vinícius V. Camargo, Rafael Soares, Sergio Bampi |
ISCAS | 9 |
| 2019 | Improving devices communication in Industry 4.0 wireless networks
Rafael Kunst, Leandro A. de Avila, Alécio Pedro Delazari Binotto, Edison Pignaton de Freitas, Sergio Bampi, Juergen Rochol |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Fast Coding Unit Partition Decision for HEVC Using Support Vector MachinesabstractDespite the several speedup methods proposed in the literature, the computational complexity of High Efficiency Video Coding (HEVC) video encoding is still a problem. This paper proposes a fast coding unit (CU) partition decision for use in HEVC encoders based on support vector machine (SVM)-trained offline. The SVM classifiers, features, and training procedures are described in detail, and a justification for the use of SVMs is provided. The trained classifiers are incorporated into a modified reference encoder in the form of a fast CU partition decision algorithm, which decides if the exhaustive search for the best partition is continued or terminated prematurely. Using the proposed method, an average complexity reduction of 48% is achieved with a 0.48% Bjontegaard-Delta bitrate (BD-BR) loss using the random access coding configuration, 44% reduction with a 0.62% BD-BR loss for the Low Delay B, and a 41% reduction with a 0.6% BD-BR loss for the Low Delay P configuration. We also tested our approach under constant bitrate conditions, achieving a 47% reduction in encoding time with a 1.11% loss in the BD-BR. In addition, a decision threshold adaptation is also proposed to allow adjusting the rate-distortion/complexity trade-off of our solution. With this approach, the computational complexity reduction can be varied from 34.9% (with a 0.13% loss in the BD-BR) up to 52.4% (with a 1.11% BD-BR loss) using the random access configuration. Compared with the state-of-the-art solutions, our decision scheme outperforms the related works in terms of combined rate distortion and complexity. Mateus Grellert, Bruno Zatt, Sergio Bampi, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Hybrid Scratchpad Video Memory Architecture for Energy-Efficient Parallel HEVCabstractA hybrid scratchpad video memory (Hy-SVM) for energy-efficient tiles-parallelized high-efficiency video coding (HEVC) is presented here. The key ideas behind the Hy-SVM include: application-specific design and management; combined multiple levels of private and shared memories that jointly exploit intra-tile and inter-tiles data reuse; scratchpad memories (SPMs) as on-chip data storage; SRAM; and STT-RAM hybrid design. We propose a design methodology for the Hy-SVM that leverages application-specific properties to properly define the SPMs parameters. The inter-tiles data reuse potential of parallel HEVC is exploited by our run-time overlap prediction scheme, which identifies the redundant memory access behavior by analyzing monitored past frames encoding. Based on the predicted overlap characteristics, the Hy-SVM integrates memory access management units to control the access dynamics to the private/shared SPM levels. Furthermore, adaptive access management units (APMUs) can strongly reduce on-chip energy consumption due to the predicted overlap formation. The experimental results demonstrate the Hy-SVM overall energy savings of 11%-64% (4-tile) and 8%-46% (8-tile) when compared with related works. From the external memory perspective, the Hy-SVM can improve data reuse, resulting in 14%-59% of off-chip energy consumption (compared with no inter-tiles data reuse scenarios). In addition, our APMU contributes by reducing on-chip energy consumption of the Hy-SVM by 58%, on average. Thus, compared with related works, the Hy-SVM presents the lowest on-chip energy consumption. Moreover, the overhead of implementing our management units insignificantly affects the performance- and energy-efficiency of the Hy-SVM. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Jörg Henkel, Sergio Bampi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Learning-Based Complexity Reduction and Scaling for HEVC EncodersabstractThis article proposes a fast Coding Unit (CU) partition decision for use in HEVC encoders based on Decision Tree classifiers. The trees are employed in a modified low-complexity encoder that implements a fast CU partition decision algorithm. Using the proposed method, an average complexity reduction of 47.8% is achieved with a Bjontegaard Delta bitrate (BD-BR) loss of 0.24% in the Random Access coding configuration, and a 42.8% complexity reduction with a 0.19% BD-BR loss in the Low Delay B configuration. A decision threshold analysis is also presented to assess the rate-distortion-complexity trade-off of the proposed method at different complexity points, varying the complexity reduction from 28% (with a 0.04% loss in BD-BR) up to 60% (with a 3.6% BD-BR loss) using the Random Access configuration. A comparison with related works shows that the proposed method outperforms competing solutions in terms of both rate-distortion efficiency and complexity reduction. Mateus Grellert, Sergio Bampi, Guilherme Corrêa 0001, Bruno Zatt, Luís Alberto da Silva Cruz |
ICASSP | 2 |
| 2018 | Exploring the Impact of Soft Errors on NoC-based Multiprocessor SystemsabstractSoftware reliability is an essential design metric in emerging large-scale multiprocessor embedded systems. Designers should identify soft error susceptibility of multiple applications executing in parallel early in the design time to ensure reliable system operation. This work proposes a non-intrusive fault injection engine that enables to conduct bespoke soft error analysis, allowing to identify and understand the soft error propagation through the processing elements (PEs). The proposed fault injection campaign evaluates the impact of soft errors considering real benchmarks in an RTL model of a distributed-memory NoC-based multiprocessor. Experiments demonstrate that 19% of soft errors are propagated to other PEs, where 31.6% of them led to erroneous computation and 58.4% to a system crash. Thus, the fault analysis must consider not only its local effect on the processor and memory but also how the fault propagates to other system components. Felipe T. Bortolon, Geancarlo Abich, Sergio Bampi, Ricardo Augusto da Luz Reis, Fernando Gehm Moraes, Luciano Ost |
ISCAS | 3 |
| 2018 | High-Throughput Binary Arithmetic Encoder using Multiple-Bypass Bins Processing for HEVC CABACabstractThe advance of massive video processing applications, devices and resolutions has led to new challenges in video encoding. The HEVC (High Efficiency Video Coding) standard emerges as one alternative in order to address the new video processing requirements. The HEVC allows only one type of entropy encoding algorithm, which is the CABAC (Context-Adaptive Binary Arithmetic Coding). The compression gains achieved by CABAC algorithm come at the cost of increasing complexity for implementation, due to intense data dependencies. The BAE (Binary Arithmetic Encoder) is the CABAC critical sub-block, in which the main part of the algorithm is executed. The present work proposes an 8-stage pipeline BAE architectural solution, named MB-BAE, with the addition of multiple-bypass bins processing, in order to increase throughput without compromising the critical path of the architecture. As a result, an average of 4.94 bins/cycle and around 2.6-Gbin/s of throughput are achieved in our BAE. This is the highest throughput found among related works in the literature, and being able to process 8K UHD videos with the lowest frequency when compared to the same related works. Fábio Luís Livi Ramos, Bruno Zatt, Marcelo Schiavon Porto, Sergio Bampi |
ISCAS | 4 |
| 2018 | Improving network resources allocation in smart cities video surveillance
Rafael Kunst, Leandro A. de Avila, Edison Pignaton de Freitas, Sergio Bampi, Juergen Rochol |
Comput. Networks | 4 |
| 2017 | A sub-1 V, nanopower, ZTC based zero-VT temperature-compensated current referenceabstractA nano-ampere current reference with temperature compensation operating is presented. The reference current is generated biasing a zero-VT transistor near its Zero-temperature coefficient (ZTC) point. Two versions were implemented in a 180 nm CMOS process. Both are designed using the same thermal compensation principle, but the second version uses an auxiliary circuit to compensate process variation. The circuits occupy 0.01 and 0.018 mm2 of silicon area while consuming around 30.5 and 122 nW at 27° C, respectively. Post-layout simulations present a reference current of 10.86 and 10.95 nA with a average temperature coefficient of 108 and 127 ppm/°C (100 Samples), under a temperature range from -20 to 120 °C, and a line sensitivity of 0.54 and 0.86 %/V at 0.9 V to 1.8 V of supply voltage, respectively. David Cordova, Arthur Campos de Oliveira, Pedro Toledo, Hamilton Klimach, Sergio Bampi, Eric E. Fabris |
ISCAS | 5 |
| 2016 | Segmentation and classification of melanocytic skin lesions using local and contextual featuresabstractThis work presents a novel approach for detecting and classifying melanocytic skin lesions on macroscopic images. We oversegment the skin lesions using superpixels, and classify independently each superpixel as a benign or malignant using local and contextual information. The overall superpixel classification results allow to calculate an index of malignancy or benignity for the skin lesion. Using the proposed approach it is possible to discriminate a malignant from a benign skin lesion by recognizing early signs of malignancy in parts of the segmented skin lesion. The experimental results are promising, and show a potential accuracy of 99.34% on a popular data set, outperforming the current state-of-art methods. Eliezer E. Bernart, Jacob Scharcanski, Sergio Bampi |
ICIP | 3 |
| 2016 | Ultra-low voltage wideband inductorless balun LNA with high gain and high IP2 for sub-GHz applicationsabstractThe design of an ultra-low voltage CMOS wideband LNA topology operating under a 0.6 V power supply with high gain and high IP2 is presented in this paper. The circuit performance is targeted towards use in direct conversion receivers, where a high IP2 is required. The LNA operates in sub-1 GHz applications reaching frequencies as low as 50 MHz, as it is required by IEEE 802.22 standard. The topology proposed here is based on the resistive-feedback inverter with a cascaded auxiliary amplifier to implement noise canceling. This cascaded amplifier is composed by a common-source (CS) amplifier and a very linear buffer at the output of the cascade. The high IP2 is obtained by a proper transistor biasing and by combining diodes with the load of the amplifiers with the highest non-linearity. This topology also works as a balun to comply with single-ended input from antennas and differential mixers as output loads. The post-layout simulations included bondwire inductances and pad capacitances parasitics. The results show a bandwidth of 50 MHz-1 GHz, voltage gain > 18 dB, NF11<; -13 dB and maximum IP2 of 38 dBm. The LNA power consumption is just 6.7 mW, excluding pad buffers. Arthur Liraneto Torres Costa, Hamilton Klimach, Sergio Bampi |
ISCAS | 3 |
| 2016 | Energy-aware cache assessment of HEVC decodingabstractThis paper presents a thorough analysis of energy consumption of a software HEVC decoder. The evaluation utilizes a framework developed herein specifically to estimate the energy consumption in all levels of cache hierarchies. Our framework is based on analytical models combined with memory profiling; tools. Energy analyses of several cache hierarchies executing HEVC decoding with different input bit streams were carried out. Our results point to the most suited cache parameters for each video resolution. The energy was estimated for a 32nm CMOS technology. Our study includes different tradeoffs between energy efficiency and capacity, associativity, and main memory bandwidth. Our detailed analysis shows that the higher are the cache features, the more efficient is the energy consumption. The main memory bandwidth evaluation shows that the energy consumption increases with the main memory bandwidth requirement. Full HD video resolutions require up to 90 times higher bandwidth and 57 times more energy than class D resolutions. Eduarda Monteiro, Mateus Grellert, Sergio Bampi, Bruno Zatt |
ISCAS | 3 |
| 2016 | A Resources Sharing Architecture for Heterogeneous Wireless Cellular NetworksabstractThe current static model used for allocating the spectrum of frequencies and the increasing demand for network resources imposed by modern applications and services may lead to a resources scarcity problem. Dealing with this problem demands optimized resources allocation. An alternative to provide this optimization is by allowing resources sharing among network operators. In this paper, an architecture is presented to encourage network operators to share their underutilized resources. A multilevel broker is proposed to control the resources sharing. This broker dynamically establishes a service level agreement that takes into account the quality of service requirements of resources renters and the cost of the resources. A performance evaluation shows that the implementation of the proposed architectures leads to better allocation of underutilized network resources. Rafael Kunst, Leandro A. de Avila, Edison Pignaton de Freitas, Sergio Bampi, Juergen Rochol |
LCN | 4 |
| 2016 | Complexity-scalable HEVC encodingabstractHEVC encoders impose several challenges in resource-constrained embedded applications, especially under real-time and battery constraints. This paper proposes a complexity-scalable encoder that is able to achieve considerable time savings while maintaining an efficient rate-distortion-complexity tradeoff. To design the system, a complexity analysis of HEVC-supported parameters as well as new ones introduced in this work is presented. To build the configurations for each target savings, a Complexity Target Satisfaction algorithm was designed. This algorithm was able to reduce the optimization space by approximately 600 times, producing efficient configurations with up to 90% time savings. The complete system was implemented and tested against a state-of-the-art complexity management implementation. The results proved the efficiency of our solution, as it achieves more time savings and better compression for savings of 60% and higher. Mateus Grellert, Sergio Bampi, Bruno Zatt |
PCS | 2 |
| 2016 | Improving QoS in multi-operator cellular networksabstractProviding high quality network access is challenging for network operators in the current static model used for allocating the spectrum of frequencies. Dealing with this challenge demands optimized resources allocation. One of the ways of providing this optimization is by allowing resources sharing among network operators which share the same geographical area. To allow and control the sharing of resources in such network scenarios, in this paper, a multilevel broker is presented to allow network operators to share their underutilized resources. This broker dynamically establishes a service level agreement that takes into account the quality of service requirements of the resources renters. A performance evaluation conducted in a scenario composed of multiple LTE-Advanced network operators shows that the implementation of the proposed architectures leads to more efficient allocation of underutilized network resources compared to two algorithms found in the literature. Rafael Kunst, Leandro A. de Avila, Edison Pignaton de Freitas, Sergio Bampi, Juergen Rochol |
WiMob | 4 |
| 2015 | Approximation-aware Multi-Level Cells STT-RAM cache architectureabstractCurrent manycore processors exhibit large on-chip last-level caches that may reach sizes of 32MB - 128MB and incur high power/energy consumption. The emerging Multi-Level Cells (MLC) STT-RAM memory technology improves the capacity and energy efficiency issues of large-sized memory banks. However, MLC STT-RAM incurs non-negligible protection overhead to ensure reliable operations when compared to the Single-Level Cells (SLC) STT-RAM. In this paper, we propose an approximation-aware MLC STT-RAM cache architecture, which is partially-protected to restrict the reliability overhead and in turn leverages variable resilience characteristics of different applications for adaptively curtailing the protection overhead under a given error tolerance level. It thereby improves the energy-efficiency of the cache while meeting the reliability requirements. Our cache architecture is equipped with a latency-aware hardware module for double-error correction. To achieve high energy efficiency, approximation-aware read and write policies are proposed that perform approximate storage management while tolerating some errors bounded within the user-provided tolerance level. The architecture also facilitates runtime control on the quality of applications' results. We perform a case study on the next-generation advanced video encoding that exhibit memory-intensive functional blocks with varying resilience properties and support for parallelism. Experimental results demonstrate that our approximation-aware MLC STT-RAM based cache architecture can improve the energy efficiency compared to state-of-the-art fully-protected caches (7%-19%, on average), while incurring minimal quality penalties in the output (-0.219% to -0.426%, on average). Furthermore, our architecture supports complete error protection coverage for all cache data when processing non-resilient application. The hardware overhead to implement our approximation-aware management negligibly affects the energy efficiency (0.15%-1.3% of overhead) and the access latency (only 0.02%-1.56% of overhead). Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
CASES | 4 |
| 2015 | A deblocking filter hardware architecture for the high efficiency video coding standard
Cláudio Machado Diniz, Muhammad Shafique 0001, Felipe Vogel Dalcin, Sergio Bampi, Jörg Henkel |
DATE | 4 |
| 2015 | 0.7 V supply self-biased nanoWatt MOS-only threshold voltage monitorabstractThis work presents a self-biased MOSFET threshold voltage VT0monitor. The threshold condition is defined based on a current-voltage relationship derived from a continuous physical model. The model is valid for any operating condition, from weak to strong inversion, and under triode or saturation regimes. The circuit consists in balancing two self-cascode cells operating at different inversion levels, where one of the transistors that compose these cells is biased at the threshold condition. The circuit is MOSFET-only (can be implemented in any standard digital process), and it operates with a power supply of less than 1 V, consuming tenths of nW. We propose a process independent design methodology, evaluating different trade-offs of accuracy, area and power consumption. Schematic simulation results, including Monte Carlo variability analysis, support the VT0monitoring behavior of the circuit with good accuracy on a 180 nm process. Oscar E. Mattia, Hamilton Klimach, Sergio Bampi, Márcio C. Schneider |
ISCAS | 3 |
| 2015 | Rate-distortion and energy performance of HEVC and H.264/AVC encoders: A comparative analysisabstractA quantitative, systematic, and detailed analysis of the energy impacts of the tools that comprise two of the most recent video coding standards: the High Efficiency Video Coding (HEVC) and the H.264/AVC is presented. Our comparative study measures the energy consumption effects of important video-coding parameters, like Search Range (SR), Quantization Parameter (QP), and video resolution on both encoders. The obtained results for HEVC showed, for the Random Access (RA) prediction structure, gains of 25% in BD-Rate over H.264/AVC at the expense of 17% higher energy consumption. A new metric we defined herein, called BD-Energy, was used in the SR analysis, and the results from this investigation showed HEVC achieved an energy consumption up to 37.6% higher for a BD-Rate gain of 32.2%. The QP analysis demonstrated that the energy consumption gap between both encoders varies greatly as QP increases, resulting in a 15.08% difference from QP 22 to QP 37, on average. The major finding from our work is that the HEVC encoder presents better results in the energy/compression trade-off, but this efficiency is reduced as encoding becomes more complex, as our results discovered that the HEVC energy consumption scales faster. Eduarda Monteiro, Mateus Grellert, Sergio Bampi, Bruno Zatt |
ISCAS | 3 |
| 2015 | A Reconfigurable Hardware Architecture for Fractional Pixel Interpolation in High Efficiency Video CodingabstractWe present a novel reconfigurable hardware architecture for interpolation filtering in high efficient video coding that adapts to run-time changes of the number of interpolation filter calls and thereby provides a high potential of energy efficiency. It employs a picture-based prediction scheme to estimate the number of interpolation filter calls at run-time by monitoring the group of pictures history based on video coding structure knowledge. Reconfigurable acceleration engines are developed that can adapt to different filter types. Dynamic composition of different instances of these engines enables different implementation versions with area versus throughput tradeoff. A run-time selection scheme determines the best implementation version for each picture based on the throughput requirements. Compared to state-of-the-art, our architecture reduces resource usage by 57% while supporting various throughputs and video resolutions. Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | dSVM: Energy-efficient distributed Scratchpad Video Memory Architecture for the next-generation High Efficiency Video CodingabstractAn energy-efficient distributed Scratchpad Video Memory Architecture (dSVM) for the next-generation parallel High Efficiency Video Coding is presented. Our dSVM combines private and overlapping (shared) Scratchpad Memories (SPMs) to support data reuse within and across different cores concurrently executing multiple parallel HEVC threads. We developed a statistical method to size and design the organization of the SPMs along with a supporting memory reading policy for energy efficiency. The key is to leverage the HEVC and video content knowledge. Furthermore, we integrate an adaptive power management policy for SPMs to manage the power states of different memory parts at run time depending upon the varying video content properties. Our experimental results illustrate that our dSVM architecture reduces the overall memory energy consumption by up to 51%-61% compared to parallelized state-of-the-art solutions [11]. The dSVM external memory energy savings increase with an increasing number of parallel HEVC threads and size of search window. Moreover, our SPM power management reacts to the current video properties and achieves up to 54% on-chip leakage energy savings. Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
DATE | 4 |
| 2014 | Run-time accelerator binding for tile-based mixed-grained reconfigurable architecturesabstractRun-time mixed-grained reconfigurable architectures emerged as an efficient solution to deal with the heterogeneous and at-design-time unpredictable nature of advanced applications. Due to interconnection limitations, the reconfigurable elements are grouped into tiles communicating through an on-chip network. State-of-the-art run-time accelerator binding schemes, i.e., mapping the accelerators to elements in the physical reconfigurable array, do not deal with such tile-based architectures. We propose a new scheme for run-time accelerator binding into our tile-based mixed-grained reconfigurable architecture. By means of an advanced video encoding application, we illustrate that our scheme reduces the inter-tile communication overhead by up to 44% (avg. 23%). Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
FPL | 3 |
| 2014 | Energy-efficient architecture for advanced video memoryabstractAn energy-efficient hybrid on-chip video memory architecture (enHyV) is presented that combines private and shared memories using a hybrid design (i.e., SRAM and emerging STT-RAM). The key is to leverage the application-specific properties to efficiently design and manage the enHyV. To increase STT-RAM lifetime, we propose a data management technique that alleviates the bit-toggling write occurrences. An adaptive power management is also proposed for static-energy savings. Experimental results illustrate that enHyV reduces on-chip static memory energy compared to SRAM-only version of enHyV and to state-of-art AMBER hybrid video memory [9] by 66%-75% and 55%-76%, respectively. Furthermore, negligible external memory energy consumption is required for reference frames communication (98% lower than state-of-the-art Level C+ technique [18]). Our data management significantly improves the enHyV STT-RAM lifetime, achieving 0.83 of normalized lifetime (near to the optimal case). Our hybrid memory design and management incur low overhead in terms of latency and dynamic energy. Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
ICCAD | 4 |
| 2014 | 2.3 ppm/°c 40 nW MOSFET-only voltage referenceabstractA MOSFET-only sub-bandgap voltage reference at less than 50 nW and with very low temperature coefficient is introduced. It consists of a threshold voltage extractor circuit and a proportional to absolute temperature voltage generator, using no resistors. The behavior of the circuit is analytically described, a design methodology is proposed and simulation results for a 0.13 µm CMOS process are presented. It allows a reference voltage below the bandgap, 625 mV in this design example, achieving a temperature coefficient of 2.3 ppm/°C for the -40 to 125 °C temperature range. The circuit consumes 40 nW at 27°C and under 1.2 V supply, being the implemented silicon area 0.0099 mm². Oscar E. Mattia, Hamilton Klimach, Sergio Bampi |
ISLPED | 3 |
| 2014 | Content-driven memory pressure balancing and video memory power management for parallel high efficiency video codingabstractWe present a novel content-driven memory pressure balancing and video memory power management scheme for parallel High Efficiency Video Coding (HEVC). The key is to leverage the application-specific knowledge to balance the (instant) access pressure on Scratchpad-based Video Memories (SVMs) for parallelized video processing. Our scheme accurately predicts the memory requirements of each processing core based on monitored memory usage and leverages this knowledge to perform a categorization of different video regions. Afterwards, it employs an adaptive policy for memory pressure balancing by rescheduling encoding of different video blocks based on their categories. This balancing also facilitates our scheme to perform efficient power-gating of unused parts of SVMs. Experimental results show that our scheme reduces the variations in the memory pressure by 37%-83% when compared to the traditional raster scan processing for 4- and 16-core parallelized HEVC encoder. Our content-driven power management saves 56% (on average) of SVM leakage energy. Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
ISLPED | 4 |
| 2013 | Energy-efficient memory hierarchy for motion and disparity estimation in multiview video codingabstractThis work presents an energy-efficient memory hierarchy for Motion and Disparity Estimation on Multiview Video Coding employing a Reference Frames-Centered Data Reuse (RCDR) scheme. In RCDR the reference search window becomes the center of the motion/disparity estimation processing flow and calls for processing all blocks requesting its data. By doing so, RCDR avoids multiple search window retransmissions leading to reduced number of external memory accesses, thus memory energy reduction. To deal with out-of-order processing and further reduce external memory traffic, a statistics-based partial results compressor is developed. The on-chip video memory energy is reduced by employing a statistical power gating scheme and candidate blocks reordering. Experimental results show that our reference-centered memory hierarchy outperforms the state-of-the-art [7][13] by providing reduction of up to 71% for external memory energy, 88% on-chip memory static energy, and 65% on-chip memory dynamic energy. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel |
DATE | 5 |
| 2013 | High-throughput interpolation hardware architecture with coarse-grained reconfigurable datapaths for HEVCabstractFractional-pel interpolation for motion estimation and motion compensation is one of the key computational hotspots in the new High Efficient Video Coding (HEVC) standard. This work presents a high-throughput interpolation hardware architecture to improve performance of HEVC encoding and decoding. It employs two acceleration engines for luma and chroma filtering, each with 12-pel-parallel coarse-grained reconfigurable interpolation datapaths. An adaptive scheduling scheme manages the operation of these interpolation datapaths in different ways depending upon the prediction unit (PU) size and the execution scenario (i.e. motion estimation or motion compensation). We have implemented our hardware architecture in 150 nm technology. Compared to state-of-the-art techniques [12], our architecture required 49% less hardware area, while processing QFHD (3840×2160) resolution @ 30 fps. Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICIP | 3 |
| 2013 | Content-adaptive reference frame compression based on intra-frame prediction for multiview video codingabstractThis paper presents a content-adaptive reference frame compression scheme to alleviate the large overhead of external memory communication during the Motion and Disparity Estimation process in Multiview Video Coding (MVC). Our scheme is based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples. The intra-prediction residue is compressed by a path composed of non-linear quantization and Huffman-based entropy encoder. Four different quantization strengths and Huffman tables were statistically defined. They are dynamically selected according to a content adaptation strategy, which classifies the original blocks based on their spatial homogeneity. Experimental results show that the proposed content-adaptive compression scheme is able to reduce the external memory accesses by up to 63% along with negligible losses in the MVC encoder rate-distortion performance. Compared to the best available related work [12] our content-adaptive reference frame compression achieves 39% reduced external memory accesses, while still providing a BD-PSNR increase of 0.03dB. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Jörg Henkel, Sergio Bampi |
ICIP | 6 |
| 2013 | Adaptive content-based Tile partitioning algorithm for the HEVC standardabstractThis paper proposes a content-based Tile partitioning algorithm designed to exploit video properties in the tiling process, aiming at the reduction of the coding losses generated by the use of Tiles. In the proposed algorithm two steps are performed to define the vertical and horizontal Tile boundaries willing to group the high correlated samples into the same Tile partition. The definition of the Tiles' boundaries is performed by analyzing the variance map extracted from the raw picture information. Experimental results have shown that the proposed algorithm is able to reduce the inherent coding efficiency losses of using Tiles when compared to the conventional uniform spaced Tile partitions. Cauane Blumenberg, Daniel Palomino 0001, Sergio Bampi, Bruno Zatt |
PCS | 3 |
| 2013 | Iterative random search: a new local minima resistant algorithm for motion estimation in high-definition videos
Marcelo Schiavon Porto, Cassio Cristani, Pargles Dall'Oglio, Mateus Grellert, Júlio C. B. de Mattos, Sergio Bampi, Luciano Volcan Agostini |
Multim. Tools Appl. | 6 |
| 2013 | Model Predictive Hierarchical Rate Control With Markov Decision Process for Multiview Video CodingabstractThis paper presents a novel hierarchical rate control (HRC) for the Multiview Video Coding standard targeting improved bandwidth usage and high video quality. The HRC is designed to jointly address the rate control at both frame level and basic unit (BU) level. The proposed scheme is able to exploit the bitrate distribution correlation with neighboring frames to efficiently predict the future bitrate behavior by employing a model predictive control that defines a proper control action through quantization parameter (QP) adaptation. To provide a fine-grained tuning, the QP is further adapted within each frame by a Markov decision process implemented at BU level able to take into consideration a map of the regions of interest. A coupled frame/BU level feedback is performed in order to guarantee the system consistency. Experimental results show the superiority of our HRC compared to state-of-the-art solutions in terms of bitrate allocation accuracy and rate distortion while delivering smooth video quality at frame and BU levels. Bruno Boessio Vizzotto, Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Adaptive power management of on-chip video memory for multiview video codingabstractAn adaptive power management of on-chip video memory for Multiview Video Coding is presented. It leverages texture, motion and disparity properties of objects and their correlations in the 3D-neighborhood. It groups different Macroblocks of a frame and predicts the highly-probable motion/disparity search direction in order to power-gate idle memory regions. Exploited are the statistical properties of Macroblock groups to predict idle sectors. Our approach achieves on average 32% and 61% energy reduction (averaged over various video sequences) compared to state-of-the-art DSW [7] and Level C [12], respectively. The Motion/Disparity Estimation architecture with video memory and power management scheme is implemented using an ASIC flow (IBM-65nm Low-Power technology) and it processes 4-view [email protected] Muhammad Shafique 0001, Bruno Zatt, Fabio Leandro Walter, Sergio Bampi, Jörg Henkel |
DAC | 4 |
| 2012 | Real-time block matching motion estimation onto GPGPUabstractThis work presents an efficient method to map Motion Estimation (ME) algorithms onto General Purpose Graphic Processing Unit (GPGPU) architectures using CUDA programming model. Our method jointly exploits the massive parallelism available in current GPGPU devices and the parallelization potential of ME algorithms: Full Search (FS) and Diamond Search (DS). Our main goal is to evaluate the feasibility of achieving real-time high-definition video encoding performance running on GPUs. For comparison reasons, multi-core parallel and distributed versions of these algorithms were developed using OpenMP and MPI (Message Passing Interface) libraries, respectively. The CUDA-based solutions achieve the highest speed-up in comparison with OpenMP and MPI versions for both algorithms and, when compared to the state-of-the-art, our FS and DS solutions reach up to 18x and 11x speed-up, respectively. Eduarda Monteiro, Marilena Maule, Felipe Sampaio, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi |
ICIP | 6 |
| 2012 | A memory aware and multiplierless VLSI architecture for the complete Intra Prediction of the HEVC emerging standardabstractThis work proposes a hardware architecture for the Intra Frame Prediction of the emerging High Efficiency Video Coding (HEVC) standard. The architecture was designed considering all innovative features of the Intra Prediction included in the HEVC, i.e. all modes and all Prediction Units (PU) sizes. Performance and memory accesses are a problem in the HEVC intra prediction and hardware architecture designs are good alternative to solve these issues, especially when energy-efficient solutions are targeted. Buffers and internal memories were used in the designed architecture to decrease the number of external memory accesses. Two independent data paths processing eight samples in parallel and a deep and multiplierless pipeline were designed to increase the throughput. The architecture was synthesized using an IBM 65nm CMOS technology. The results have shown that the architecture is able to process 30 HD720p frames per second and 13 HD1080p frames per second when running at 500 MHz, reducing in 95% the accesses to the external memory. Daniel Palomino 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin |
ICIP | 4 |
| 2012 | Motion Vectors Merging: Low Complexity Prediction Unit Decision Heuristic for the Inter-prediction of HEVC EncodersabstractThis paper presents the Motion Vectors Merging (MVM) heuristic, which is a method to reduce the HEVC inter-prediction complexity targeting the PU partition size decision. In the HM test model of the emerging HEVC standard, computational complexity is mostly concentrated in the inter-frame prediction step (up to 96% of the total encoder execution time, considering common test conditions). The goal of this work is to avoid several Motion Estimation (ME) calls during the PU inter-prediction decision in order to reduce the execution time in the overall encoding process. The MVM algorithm is based on merging NxN PU partitions in order to compose larger ones. After the best PU partition is decided, ME is called to produce the best possible rate-distortion results for the selected partitions. The proposed method was implemented in the HM test model version 3.4 and provides an execution time reduction of up to 34% with insignificant rate-distortion losses (0.08 dB drop and 1.9% bitrate increase in the worst case). Besides, there is no related work in the literature that proposes PU-level decision optimizations. When compared with works that target CU-level fast decision methods, the MVM shows itself competitive, achieving results as good as those works. Felipe Sampaio, Sergio Bampi, Mateus Grellert, Luciano Volcan Agostini, Júlio C. B. de Mattos |
ICME | 2 |
| 2012 | Spread and Iterative Search: A High Quality Motion Estimation Algorithm for High Definition Videos and Its VLSI DesignabstractThis paper presents the Spread and Iterative Search (S&IS) motion estimation algorithm, which uses a random spread evaluation together with a central iterative evaluation to avoid local minima falls and to increase the image quality for high definition videos. Considering Full HD videos, S&IS reached an average PSNR gain of 1.41dB when compared to Diamond Search (DS), with an increase of about four times in the number of evaluated blocks. When compared to Full Search (FS), the S&IS achieved an average PSNR loss of 1.56 dB, evaluating 73 times less blocks than FS. An efficient architecture for the S&IS algorithm is also presented in this paper. The architecture was designed targeting in real time processing (30 frames per seconds) for QFHD videos (3840×2160 pixels). The architecture was described in VHDL and synthesized for and Altera Stratix 4 FPGA and for ST90nm standard cells technology. Booth syntheses show that the architecture is able to process QFHD frames in real time. The standard cells version is able to reach also a good trade-off among area, memory and power consumption, processing QFHD videos with 62.2 mW. Gustavo Sanchez, Luciano Volcan Agostini, Felipe Sampaio, Marcelo Schiavon Porto, Sergio Bampi |
ICME | 5 |
| 2012 | A Model Predictive Controller for Frame-Level Rate Control in Multiview Video CodingabstractIn this work, we present a novel frame-level Rate Control algorithm for Multiview Video Coding encoder that adopts the Model Predictive Control technique in order to provide low bitrate fluctuation and high video quality. Our Model Predictive Rate Control (MPRC) predicts the bitrate for a frame by employing (i) inter-view inter-GOP (Group of Pictures) phase-based bitrate prediction, and (ii) temporal (intra-GOP) target bitrate linear weighting. Moreover, the MPRC also defines an optimal control action through frame-level QP value selection. Experimental results demonstrate that our MPRC bitrate prediction incurs a Mean Bit Estimation Error (MBEE) of 1.13% compared to 2.46% provided by single view-based Rate Control and 1.61% provided by the state-of-the-art MVC Rate Control. Our solution also provides on average 0.876dB BD-PSNR increase and 28.92% BD-Bitrate reduction while providing smoother quality and bitrate variations when compared to state-of-the-art. Bruno Boessio Vizzotto, Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICME | 4 |
| 2012 | A low-cost and high efficiency entropy encoder architecture for H.264/AVC
Cristiano Thiele, Bruno Boessio Vizzotto, André Luís Machado Martinez, Vagner Santos Da Rosa, Sergio Bampi |
VLSI-SoC | 5 |
| 2011 | Run-time adaptive energy-aware motion and disparity estimation in multiview video codingabstractThis paper presents a novel run-time adaptive energy-aware Motion and Disparity Estimation (ME, DE) architecture for Multiview Video Coding (MVC). It incorporates efficient memory access and data prefetching techniques for jointly reducing the on/off-chip memory energy consumption. A dynamically expanding search window is constructed at run time to reduce the off-chip memory accesses. Considering the multi-stage processing nature of advanced fast ME/DE schemes, a reduced-sized multi-bank on-chip memory is employed which can be power-gated depending upon the video properties. As a result, when tested for various video sequence, our approach provides a dynamic energy reduction of 82--96% for the off-chip memory and a leakage energy reduction of 57--75% for the on-chip memory compared to the Level-C and Level-C+ [7] prefetching techniques (which are the prominent data reuse and prefetching techniques in ME for video coding). The proposed ME/DE architecture is synthesized using a 65nm IBM low power technology. Compared to state-of-the-art MVC ME/DE hardware [14], our architecture provides 66% and 72% reduction in the area and power consumption, respectively. Moreover, our scheme achieves 30fps ME/DE 4-view HD1080p encoding with a power consumption of 74mW. Bruno Zatt, Muhammad Shafique 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel |
DAC | 5 |
| 2011 | Multi-level pipelined parallel hardware architecture for high throughput motion and disparity estimation in Multiview Video CodingabstractThis paper presents a novel motion and disparity estimation (ME, DE) scheme in Multiview Video Coding (MVC) that addresses the high throughput challenge jointly at the algorithm and hardware levels. Our scheme is composed of a fast ME/DE algorithm and a multi-level pipelined parallel hardware architecture. The proposed fast ME/DE algorithm exploits the correlation available in the 3D-neighborhood (spatial, temporal, and view). It eliminates the search step for different frames by prioritizing and evaluating the neighborhood predictors. It thereby reduces the coding computations by up to 83% with 0.1 dB quality loss. The proposed hardware architecture further improves the throughput by using parallel ME/DE modules with a shared array of SAD (Sum of Absolute Differences) accelerators and by exploiting the four levels of parallelism inherent to the MVC prediction structure (view, frame, reference frame, and macroblock levels). A multi-level pipeline schedule is introduced to reduce the pipeline stalls. The proposed architecture is implemented for a Xilinx Virtex-6 FPGA and as an ASIC with an IBM 65nm low power technology. It is compared to state-of-the-art at both algorithm and hardware levels. Our scheme achieves a real-time (30fps) ME/DE in 4-view High Definition (HD1080p) encoding with a low power consumption of 81 mW. Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
DATE | 3 |
| 2011 | A low-power memory architecture with application-aware power management for motion & disparity estimation in Multiview Video CodingabstractA low-power architecture for an on-chip multi-banked video memory for motion and disparity estimation in Multiview Video Coding is proposed. The memory organization (size, banks, sectors, etc.) is driven by an extensive analysis of memory-usage behavior for various 3D-video sequences. Considering a multiple-sleep state model, an application-aware power management scheme is employed to reduce the leakage energy of the on-chip memory. The knowledge of motion and disparity estimation algorithm in conjunction with video properties are considered to predict the memory requirements of each Macroblock. A cost function is evaluated to determine an appropriate sleep mode for the idle memory sectors, while considering the wakeup overhead (latency and energy). The complete motion and disparity estimation architecture is implemented in a 65nm low power IBM technology. The experiments (for various test video sequences) demonstrate that our architecture provides up to 80% leakage energy reduction compared to state-of-the-art. Our scheme processes motion and disparity estimation of four HD1080p views encoding at 30fps with a power consumption of 57mW. Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICCAD | 3 |
| 2011 | A multi-level dynamic complexity reduction scheme for multiview video codingabstractIn this paper, we propose a novel scheme for dynamically reducing the computational complexity of MVC. Our scheme exploits the coding mode correlation available in the 3D-neighborhood (i.e., spatial, temporal, and view) along with the rate-distortion proper- ties of the neighboring Macroblocks. Our scheme incorporates a multi-level mode decision process based on a mode-ranking mechanism that categorizes more-probable and less-probable coding modes. In order to react to the changing bitrates, our scheme deploys Quantization Parameter based threshold equations which are formulated using an offline statistical analysis. Compared to the exhaustive Rate-Distortion-Optimized Mode Decision (RDO-MD), our scheme achieves a complexity reduction of up to 80% (68% on average) with an average PSNR loss of 0.075 dB. Compared to state-of-the-art fast RDO-MD, our scheme achieves a complexity reduction of up to 34% with an average PSNR gain of 0.007 dB. Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICIP | 3 |
| 2011 | SHBS: A heuristic for fast inter mode decision of H.264/AVC standard targeting VLSI designabstractIn the Rate-Distortion Optimization technique for H.264/AVC, the process of choosing the best mode is performed through exhaustive executions of the whole encoding process, which increases significantly the encoder complexity, sometimes even forbidding its use in real time video coding applications. In order to reduce the number of calculations necessary to determine the best inter-frame mode, this work proposes the SHBS (Stationarity, Heterogeneity and Border Strength) heuristic. The use of SHBS causes a reduction of 168 times in the encoding iterations, with a better PSNR, at the cost of a relatively small bit-rate increase. The SHBS heuristic was designed in hardware targeting FPGAs and this architecture achieved an operation frequency of 118 MHz, being able to process up to 438 HD 1080p frames per second. Guilherme Corrêa 0001, Daniel Palomino 0001, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
ICME | 5 |
| 2011 | A high throughput H.264/AVC intra-frame encoding loop architecture for HD1080pabstractIn this work we present a high throughput hardware architecture for the H.264/AVC intra-frame encoder exploiting the parallelism of intra prediction, forward and inverse transforms and quantization. Since there is a strong data dependency between the intra prediction and the image reconstruction loop, the latency of this path is a key design issue in order to provide high performance coding. Considering that 77% of the total intra-encoding computation is spent in these modules, our architecture handles a 4-pixel wide intra prediction module and a 16-pixel wide reconstruction loop. Compared to the state-of-the-art our approach reduces by 47% the number of cycles to process a macroblock. Running at 150 MHz our architecture guarantees encoding of 61 HD1080p frames per second. The developed architecture requires 73.4 MHz to real-time encode HD1080p, which is a 46% reduction of the frequency requirement compared to the state-of-the-art. Cláudio Machado Diniz, Bruno Zatt, Cristiano Thiele, Altamiro Amadeu Susin, Sergio Bampi, Felipe Sampaio, Daniel Palomino 0001, Luciano Volcan Agostini |
ISCAS | 5 |
| 2011 | Applying CUDA Architecture to Accelerate Full Search Block Matching Algorithm for High Performance Motion Estimation in Video EncodingabstractThis work presents a parallel GPU-based solution for the Motion Estimation (ME) process in a video encoding system. We propose a way to partition the steps of Full Search block matching algorithm in the CUDA architecture. A comparison among the performance achieved by this solution with a theoretical model and two other implementations (sequential and parallel using OpenMP library) is made as well. We obtained a O(n^2/log^2n) speed-up which fits the proposed theoretical model considering different search areas. It represents up to 600x gain compared to the serial implementation, and 66x compared to the parallel OpenMP implementation. Eduarda Monteiro, Bruno Boessio Vizzotto, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi |
SBAC-PAD | 5 |
| 2010 | Gop structure adaptive to the video content for efficient H.264/AVC encodingabstractThis paper presents a new method for high efficiency video coding using an adaptive GOP structure based on video content for the H.264/AVC standard. The available H.264/AVC encoders typically use static GOP sizes that define how the frames I (Intra), P (Predictive) and B (Bi-predictive) are positioned during de coding process. However, by analyzing the video content it is possible to identify the optimum position for each type of frame inside the GOP. The proposed method analyses the video content and finds the best position for inserting I frames in the video sequence. Thus the GOP structure can assume different sizes, depending on the video content. The results for test sequences and real videos show that the proposed method can significantly reduce the required bit rate, comparing to the static GOP sizes, with reduced PSNR losses. The proposed adaptive GOP presents a gain, in terms of bit rate reduction for real movies, of 8.6%, 15%, 24.7% and 40.8% in comparison with static GOP sizes 32, 16, 8 and 4, respectively. Bruno Zatt, Marcelo Schiavon Porto, Jacob Scharcanski, Sergio Bampi |
ICIP | 4 |
| 2010 | Power-aware complexity-scalable multiview video coding for mobile devicesabstractWe propose a novel power-aware scheme for complexity-scalable multiview video coding on mobile devices. Our scheme exploits the asymmetric view quality which is based on the binocular suppression theory. Our scheme employs different quality-complexity classes (QCCs) and adapts at run time depending upon the current battery state. It thereby enables a run-time tradeoff between complexity and video quality. The experimental results show that our scheme is superior to state-of-the-art and it provides an up to 87% complexity reduction while keeping the PSNR close to the exhaustive mode decision. We have demonstrated the power-aware adaptivity between different QCCs using a laptop with battery charging and discharging scenarios. Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
PCS | 3 |
| 2010 | An adaptive early skip mode decision scheme for multiview video codingabstractIn this work a novel scheme is proposed for adaptive early SKIP mode decision in the multiview video coding based on mode correlation in the 3D-neighborhood, variance, and ratedistortion properties. Our scheme employs an adaptive thresholding mechanism in order to react to the changing values of Quantization Parameter (QP). Experimental results demonstrate that our scheme provides a consistent time saving over a wide range of QP values. Compared to the exhaustive mode decision, our scheme provides a significant reduction in the encoding complexity (up to 77%) at the cost of a small PSNR loss (0.172 dB in average). Compared to state-of-the-art, our scheme provides an average 2x higher complexity reduction with a relatively higher PSNR value (avg. 0.2 dB). Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
PCS | 3 |
| 2010 | Fast forward and inverse transforms for the H.264/AVC standard using hierarchical adder compressorsabstractThis paper presents fast architectures for the forward and inverse transforms of the H.264/AVC video compression standard. These transforms can be computed exactly as in integer arithmetic, thus avoiding mismatch problems between the encoder and decoder. They are inserted into the T and T-1block of the H.264/AVC and they can be computed by using only additions and shifts. Since the transforms algorithms are composed by a large number of addition/subtraction, fast architectures for the 4×4 Discrete Cosine transforms and 4×4 and 2×2 Hadamard transforms are proposed using efficient hierarchical adder compressors. The designs were described in VHDL and mapped to TSMC 0.18μm CMOS standard cells. Experimental results show that the architectures using 8-2 adder compressor can reach high frequency operation, high throughput and they are more efficient than the solutions of the literature. João S. Altermann, Eduardo A. C. da Costa, Sergio Bampi |
VLSI-SoC | 3 |
| 2010 | Timing and interface communication analysis of H.264/AVC encoder using SystemC modelabstractThis work presents a detailed timing and communication analysis for an H.264/AVC video encoder architecture using a SystemC model. The model was described using different abstraction levels in order to evaluate specific characteristics of each component module. The target encoder is defined to be able for H.264/AVC real-time encoding for 1080p video sequences at 30 fps and was modeled as a two-stage macro-pipeline system composed by eight component modules: Macroblock buffer, Intra- and Inter-Frame Predictors, Mode Decision, Forward and Inverse Transforms and Quantization, Reference Memory Write and Entropy Encoder (CAVLC). The bandwidth of each internal connection and of external memory interface was evaluated. The timing behavior and the data dependencies were characterized and summarized in a timing diagram in order to define design constraints and provide an accurate system specification when compared to a H.264/AVC encoder in the literature. Bruno Zatt, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
VLSI-SoC | 4 |
| 2009 | A real time H.264/AVC intra frame prediction hardware architecture for HDTV 1080P videoabstractThis work presents an intra frame prediction hardware architecture for H.264/AVC baseline/main profile encoder which performs real time processing of HDTV 1080p videos. It is achieved by exploring the parallelism of intra prediction and by reducing the latency for Intra 4times4 processing, which is the intra encoding bottleneck. Synthesis results on Xilinx Virtex-II Pro FPGA and TSMC 0.18 mum standard-cells indicate that this architecture is able to real time encode HDTV 1080p video operating at 110 MHz. Our architecture can encode HD1080p, 720p and SD video in real time at a frequency 25% lower when compared to similar works. Cláudio Machado Diniz, Bruno Zatt, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi |
ICME | 5 |
| 2009 | A High Performance H.264 Deblocking FilterabstractAlthough the H.264 Deblocking Filter process is a relatively small piece of code in a software implementation, profile results shows it cost about a third of the total CPU time in the decoder. This work presents a high performance architecture for implementing a H.264 Deblocking Filter IP that can be used either in the decoder or in the encoder as a hardware accelerator for a processor or embedded in a full-hardware codec. A developed IP using the proposed architecture support multiple high definition processing flows in real-time. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Vagner Santos Da Rosa, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 3 |
| 2009 | Challenges and Emerging Technologies for System Integration beyond the End of the Roadmap of Nano-CMOS
Sergio Bampi, Ricardo Augusto da Luz Reis |
VLSI-SoC | 1 |
| 2009 | Techniques for Architecture Design for Binary Arithmetic Decoder Engines Based on Bitstream Flow Analysis
Dieison Antonello Deprá, Sergio Bampi |
VLSI-SoC | 2 |
| 2008 | A high throughput and low cost diamond search architecture for HDTV motion estimationabstractThis paper presents a high throughput and low cost architecture for motion estimation using a sub-sampled diamond search algorithm (SDS). The quality of SDS was compared with full search through software implementations and the results are presented. The designed hardware considered a search area of 100times100 samples, with blocks of 16times16 pixels. The architecture was described in VHDL and mapped to a Xilinx Virtex-4 FPGA. Synthesis results indicate that SDS is able to run at 185.7 MHz, using only 3541 LUTs. This architecture can reach real time for HDTV (1920times1080 pixels) in the worst case, and it can process 120 HDTV frames per second in the average case. Marcelo Schiavon Porto, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin |
ICME | 3 |
| 2008 | HP422-MoCHA: A H.264/AVC High Profile motion compensation architecture for HDTVabstractThis work presents the HP422-MoCHA, the first published hardware architecture that implements a full compliant H.264/AVC motion compensator for high profile 4:2:2. The hardware is composed by three main modules: Motion Vector Predictor, Memory Access and Sample Interpolator. The designed architecture was described in VHDL and mapped to a Xilinx Virtex-II PRO FPGA. This architecture reaches the throughput to decode HDTV Level 4 (1080p) @ 30fps. Bruno Zatt, Altamiro Amadeu Susin, Sergio Bampi, Luciano Volcan Agostini |
ISCAS | 3 |
| 2007 | MoCHA: a Bi-Predictive Motion Compensation Hardware for H.264/AVC Decoder Targeting HDTVabstractThis paper presents the MoCHA (motion compensation hardware architecture) design. MoCHA is an architectural design for bi-predictive motion compensation of the H.264/AVC decoder. The designed architecture features a memory hierarchy to reduce the memory bandwidth and the number of memory access cycles. The architecture uses a single datapath to process bi-predictive reference areas and it processes luma and chroma samples in parallel. The design was mapped to a Xilinx Virtex II Pro FPGA and it is able to run at 100MHz. The throughput is enough to support more than 30 bi-predictive HDTV frames per second. Arnaldo Azevedo, Bruno Zatt, Luciano Volcan Agostini, Sergio Bampi |
ISCAS | 4 |
| 2007 | A 4-Bits Trimmed CMOS Bandgap Reference with an Improved Matching Modeling DesignabstractComponent tolerances and mismatches due to process variations severely degrade the performance of bandgap reference (BGR) circuits. In this paper, the authors describe the design of a BGR considering the Pelgrom's mismatch model. The main purpose of our methodology is to convey the design to reach a good trade-off between area and mismatch. Implemented in standard 0.35μm CMOS technology, the circuit also includes a straightforward 4-bits trimming circuit to achieve more process variations independence. Its Monte Carlo temperature coefficient average is 40-ppm/°C and the reference output voltage average is 1.230V. The area of the BGR is 400x350μm2due to our design matching requirements. Juan Pablo Martinez Brito, Sergio Bampi, Hamilton Klimach |
ISCAS | 2 |
| 2007 | High Throughput Hardware Architecture for Motion Estimation with 4: 1 Pel Subsampling Targeting Digital Television Applications
Marcelo Schiavon Porto, Luciano Volcan Agostini, Leandro Rosa, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 5 |
| 2007 | A Pipelined 8x8 2-D Forward DCT Hardware Architecture for H.264/AVC High Profile Encoder
Thaísa Leal da Silva, Cláudio Machado Diniz, João Alberto Vortmann, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 6 |
| 2007 | Motion Compensation Hardware Accelerator Architecture for H.264/AVC
Bruno Zatt, Valter Ferreira, Luciano Volcan Agostini, Flávio Rech Wagner, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 6 |
| 2007 | An HDTV H.264 deblocking filter in FPGA with RGB video outputabstractThis paper presents an architecture for implementing the H.264 Deblocking Filter with RGB output in FPGA, exceeding HDTV requirements when synthesized to a target FPGA. The goal of the design was to achieve the HDTV requirements, designing a deep pipelined architecture that makes a balanced use of the resources available in the target FPGA architecture. When synthesized to VirtexII-pro FPGA, the developed architecture used only 1800 logic cells and achieved 71 frames per second at 1080p HDTV resolution (1920x1080). Vagner Santos Da Rosa, Altamiro Amadeu Susin, Sergio Bampi |
VLSI-SoC | 3 |
| 2007 | A new array architecture for signed multiplication using Gray encoded radix-2m operands
Eduardo A. C. da Costa, José Monteiro 0001, Sergio Bampi |
Integr. | 3 |
| 2006 | FPGA Design of A H.264/AVC Main Profile Decoder for HDTVabstractThis paper presents the architecture, design, validation, and prototyping of inverse transforms and quantization, intra prediction, motion compensation and loop filter, for a main profile H.264/AVC decoder. These architectures were designed to reach high throughputs and to be easily integrated with the other H.264/AVC modules. The architectures, all fully H.264/AVC compliant, were completely described in VHDL and further validated through simulations down to prototyping. The architectures were prototyped using a Digilent XUP V2P board, containing a Virtex-II Pro XC2VP30 Xilinx FPGA. The post place-and-route synthesis results indicate that the designed architectures are able to process 114 million of samples per second and, in the worst case, they are able to process 64 HDTV frames (1080×1920) per second, allowing their use in H.264/AVC decoders targeting real time HDTV applications. Luciano Volcan Agostini, Arnaldo Azevedo, Vagner Santos Da Rosa, Eduardo A. Berriel, Tatiana Gadelha Serra dos Santos, Sergio Bampi, Altamiro Amadeu Susin |
FPL | 6 |
| 2006 | FPGA Based Architectures for H. 264/AVC Video Compression StandardabstractThe H.264/AVC (as known as MPEG-4 part 10) [1, 2] is a video coding standard that has been developed to achieve significant improvements, in the compression performance, over the existing standards. The main blocks of a H.264/AVC encoder are the motion estimation, the motion compensation, the intra prediction, the loop filter, the entropy coder, the forward and inverse quantization and the forward and inverse transforms. The H.264/AVC decoder is formed by entropy decoder, motion compensation, intra prediction, loop filter, inverse quantization and inverse transforms [1]. This work focuses on the design of high performance architectures for the H.264/AVC standard. Luciano Volcan Agostini, Sergio Bampi |
FPL | 2 |
| 2006 | High throughput architecture for H.264/AVC forward transforms blockabstractThis paper presents a high throughput hardware for the complete H.264/AVC forward transforms block. There are three different transform inside this block and the presented architecture synchronizes these transforms, generating a constant processing rate in its outputs. This is an important characteristic of this architecture that was designed to be easily integrated to the other H.264/AVC blocks. The architecture does not use memory bits and the transforms in two dimensions are calculated directly, without the use of the separability property. The architecture was described in VHDL and was validated and prototyped using a Xilinx Virtex II Pro FPGA. The synthesis was directed to a VP30 FPGA and to a TSMC 0.35μm standard-cell technology. The throughputs of the T block architecture for these two different technologies reaches a processing rate higher than 120 million of samples per second, allowing its use in H.264/AVC codecs directed to HDTV. Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, Sergio Bampi, Leandro Rosa, José Luís Güntzel, Ivan Saraiva Silva |
ACM Great Lakes Symposium on VLSI | 3 |
| 2006 | High throughput multitransform and multiparallelism IP for H.264/AVC video compression standardabstractThis paper presents the design of a high throughput multitransform and multiparallelism IP for H.264/AVC standard. This solution supports the five H.264/AVC transforms and it supports five different levels of parallelism. The proposed architecture were described in VHDL and synthesized to Altera Stratix and Xilinx Virtex-II Pro FPGAs and to TSMC 0.35/spl mu/m standard cells. The multitransform and multiparallelism architecture mapped to FPGAs could process from 124 millions to 3.2 billions of samples per second, depending on the parallelism level selected. The standard cells version could process from 218.7 millions to 3.5 billions of samples per second. These results indicate that the proposed solution presents a high flexibility and that this solution is able to be used in various H.264/AVC codecs with different performance requirements. The performance results of all experiments realized indicated that this architecture is able to be used in high definition applications, like HDTV. Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, José Luís Güntzel, Ivan Saraiva Silva, Sergio Bampi |
ISCAS | 5 |
| 2006 | Reconfigurable analog interface for mixed signal SOCabstractThis work discusses the development, modeling and implementation results of a programmable architecture to be employed for analog signals interfacing in mixed-signal SOC. The proposed architecture is able to achieve wide frequency range, covering a large range of applications with constant performance, allied to digital configuration compatibility. The proposed approach utilizes the concept of frequency translation (mixing) and SigmaDelta modulation leading to a fairly constant analog block for an input signal uniform treatment from DC to high frequencies. The interface performance theoretical model is addressed for supporting the design space exploration and also the physical design. An interface prototype using a fourth order continuous time band-pass SigmaDelta modulator is built and characterized validating the proposed performance model. The usage of this interface as a multi-band parametric ADC is presented Eric E. Fabris, Luigi Carro, Sergio Bampi |
ISCAS | 3 |
| 2006 | A tool for automatic design of analog circuits based on gm/ID methodologyabstractThe goal of this paper is to present a transistor optimization methodology for analog integrated CMOS circuits, based on the physics-based gm/IDcharacteristics provided by the ACM compact MOS model. This methodology is implemented in a design tool, exploiting all the design space with the use of simulated annealing optimization process. A single technology dependent curve and accurate expressions for transconductance and current in all operations regions are integrated in the methodology, providing solutions close to the optimum. The advantage of constraining the optimization within a power budget is of great importance for low-power applications. As an example, we show the optimization results obtained for the design of a folded-cascode operational amplifier and a comparison with a typical hand-made design procedure Alessandro Girardi, Fernando da Rocha Paixão Cortes, Sergio Bampi |
ISCAS | 3 |
| 2006 | Motion Compensation Decoder Architecture for H.264/AVC Main Profile Targeting HDTVabstractThis work presents the design, the validation and the prototyping of a motion compensation architecture for a H.264/AVC video decoder. The designed architecture supports the main profile level 4.0 and it targets high resolution applications, like HDTV. This design considers the sample processing of the motion compensation block, which includes quarter-pel interpolation, weighted prediction, average to bi-predictive processing and clipping. The architecture processes luma and chroma samples in parallel, with independent luma and chroma datapaths. The design uses a single interpolator to process bi-predictive macroblocks. The design was synthesized to FPGA and standard cell technologies. The synthesis results had indicated that this architecture reaches 100 MHz in both technologies, allowing real time to decode HDTV videos with 1920times1080 pixels. The prototype was targeted to a Xilinx Virtex-II PRO FPGA Arnaldo Azevedo, Bruno Zatt, Luciano Volcan Agostini, Sergio Bampi |
VLSI-SoC | 4 |
| 2006 | A VHDL Generation Tool for Optimized Parallel FIR FiltersabstractThis paper presents generation tool and performance results on a method to minimize the amount of hardware needed to implement a parallel digital finite impulse response (FIR) filters for hardwired (fixed coefficients) implementation targeted for high performance. The generation tool employ a combination of two approaches: first, the reduction of the coefficients to n-power-of-two (NPT) terms, using cannonical signed digit (CSD) as an option, followed by common subexpression elimination (CSE) among multipliers. Synthesis results for a range of different filter specifications, using Quartus II FPGA synthesis tool and Cadence PKS standard cell synthesis tool are presented Vagner Santos Da Rosa, Eduardo A. C. da Costa, Sergio Bampi |
VLSI-SoC | 3 |
| 2006 | Architecture of an HDTV Intraframe Predictor for a H.264 DecoderabstractMultimedia applications need larger and larger bandwidth. The only way to face these demands is to provide more efficient compression algorithms, with the expense of computational complexity. The most efficient compression standard available today is the H.264/AVC. On the architectural point of view, an H.264 decoder can be seen as a system of six main modules: entropy decoder, inverse quantization, inverse transform, motion compensation, intraframe prediction and deblocking filter. These modules can be designed independently and enclosed in IPs, which could be used later to build an H.264 SoC. This paper presents an architecture design of intraframe prediction module. The system was completely built in VHDL language and prototyped over a Xilinx Virtex II Pro FPGA of 27,392 logic elements. The proposed architecture attained the required performance to decode HDTV video stream in realtime, i.e. 1920 times 1080 pixels at 30 frames per second, and used just 20% of the chip Wagston T. Staehler, Eduardo A. Berriel, Altamiro Amadeu Susin, Sergio Bampi |
VLSI-SoC | 4 |
| 2005 | Area and Throughput Trade-Offs in the Design of Pipelined Discrete Wavelet Transform ArchitecturesabstractThe JPEG2000 standard defines the discrete wavelet transform (DWT) as a linear space-to frequency transform of the image domain in an irreversible compression. This irreversible discrete wavelet transform is implemented by an FIR filter using 9/7 Daubechies coefficients or a lifting scheme of factorized coefficients from 9/7 Daubechies coefficients. The paper investigates the tradeoffs between area, power and data throughput (or operating frequency) of several implementations of the discrete wavelet transform using the lifting scheme in various pipeline designs. The paper shows the results of five different architectures synthesized and simulated in FPGAs. It concludes that the descriptions with pipelined operators provide the best area-power-operating frequency trade-off over non-pipelined operator descriptions. Those descriptions require around 40% more hardware to increase the maximum operating frequency up to 100% and reduce power consumption to less than 50%. Starting from behavioral HDL descriptions provides the best area-power-operating frequency trade-off, improving hardware cost and maximum operating frequency by around 30% in comparison to structural descriptions for the same power requirement. Sandro V. Silva, Sergio Bampi |
DATE | 2 |
| 2005 | A FPGA Based Design of a Multiplierless and Fully Pipelined JPEG CompressorabstractThis paper presents the design and implementation of a multiplierless JPEG compressor for gray scale images. The modules of this architecture were fully pipelined and targeted to FPGA device implementation. The designed architectures are detailed in this paper and they were described in VHDL, simulated and physically mapped to Altera Flex10KE FPGAs. The JPEG compressor pipeline has a minimum latency of 238 clock cycles, given the full modular pipeline depth. The minimum compressor period is 26.6ns and the compressor is able to process 37.6 millions of pixels per second. For example, the compressor can process a 640x480 pixels still image in 8.2 ms, reaching a maximum processing rate of 122.4 frames per second. Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, Sergio Bampi, Ivan Saraiva Silva |
DSD | 3 |
| 2005 | Reusing Traces in a Dynamic Conditional Execution ArchitectureabstractThe cost of control and data dependences in superscalar processors is still an open issue, for what no definitive solution was yet found. Moreover, the cost of branch mispredictions is getting worse due to the increasing number of pipeline stages. The dynamic conditional execution (DCE) is a new approach to address this problem. The basic idea is to fetch and execute all paths produced by a branch that obey certain restrictions regarding complexity and size. As a consequence, a smaller number of predictions is performed, and therefore, a smaller number of branches is mispredicted. Although the execution of multiple paths of certain branches allows for a reduction in branch misprediction penalties, it implies on an increase in the number of executed instructions. Thus, an alternative to reduce the overhead created by DCE pipeline is to reuse previously executed values, freeing up resources for more useful instructions. The goal of this work is to analyze the impact of value reuse in DCE architecture. As it is presented, this effectively reduces the overhead produced by the architecture, increasing the overall performance. This paper shows that, in some cases, the speedup gain exceeds 60% over the original DCE architecture. Tatiana Gadelha Serra dos Santos, Sergio Bampi, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 2 |
| 2005 | A Comparison of Layout Implementations of Pipelined and Non-Pipelined Signed Radix-4 Array Multiplier and Modified Booth Multiplier Architectures
Leonardo Londero de Oliveira, Cristiano Santos, Daniel Lima Ferrão, Eduardo A. C. da Costa, José Monteiro 0001, João Baptista dos Santos Martins, Sergio Bampi, Ricardo Augusto da Luz Reis |
VLSI-SoC | 7 |
| 2004 | A Run-Time Reconfigurable Datapath Architecture for Image Processing ApplicationsabstractThis paper describes a run-time reconfigurable architecture targeted to flexible low-level image processing functions. The purpose is to present the evolution of the DRIP (dynamically reconfigurable image processor) architecture from a statically configurable datapath design to a dynamically reconfigurable approach. The methodology used to redefine the datapath basic building blocks and the hardware units developed to provide an efficient and flexible image processing system are also discussed. An important issue is the granularity of the basic processing elements of the datapath, in view of the combination of programmable function by hardware control-the classical datapath paradigm-and the dynamic reconfiguration. DRIP can perform a large set of digital image processing algorithms with real-time performance to fulfill the requirements of contemporary complex applications. Marcos R. Boschetti, Ivan Saraiva Silva, Sergio Bampi |
DATE | 3 |
| 2004 | Design of Very Deep Pipelined Multipliers for FPGAsabstractThis work investigates the use of very deep pipelines for implementing circuits in FPGAs, where each pipeline stage is limited to a single FPGA logic element (LE). The architecture and VHDL design of a parameterized integer array multiplier is presented and also an IEEE 754 compliant 32-bit floating-point multiplier. We show how to write VHDL cells that implement such approach, and how the array multiplier architecture was adapted. Synthesis and simulation were performed for Altera Apex20KE devices, although the VHDL code should be portable to other devices. For this family, a 16 bit integer multiplier achieves a frequency of 266 MHz, while the floating point unit reaches 235 MHz, performing 235 MFLOPS in an FPGA. Additional cells are inserted to synchronize data, what imposes significant area penalties. This and other considerations to apply the technique in real designs are also addressed. Alex Panato, Sandro V. Silva, Flávio Rech Wagner, Marcelo O. Johann, Ricardo Augusto da Luz Reis, Sergio Bampi |
DATE | 6 |
| 2004 | Throughput and Reconfiguration Time Trade-Offs: From Static to Dynamic Reconfiguration in Dedicated Image Filters
Marcos R. Boschetti, Sergio Bampi, Ivan Saraiva Silva |
FPL | 2 |
| 2004 | Analog Signal Processing Reconfiguration for Systems-on-Chip Using a Fixed Analog Cell Approach
Eric E. Fabris, Luigi Carro, Sergio Bampi |
FPL | 3 |
| 2003 | LIT - An Automatic Layout Generation Tool for Trapezoidal Association of Transistors for Basic Analog Building Blocks
Alessandro Girardi, Sergio Bampi |
DATE | 2 |
| 2003 | Complex Branch Profiling for Dynamic Conditional ExecutionabstractBranch predictors are widely used as an alternative to deal with conditional branches. Despite the high accuracy rates, misprediction penalties are still large in any superscalar pipeline. DCE, or dynamic conditional execution, is an alternative to reduce the number of predicted branches by executing both paths of certain branches, reducing the number of predictions and, therefore, the occurrence of mispredictions. The goal of this work is to analyze the complexity of branch structures and determine the number of branches that can be predicated in DCE and the distribution of mispredictions according to the proposed classification. The complex branch classification proposed extends the classification presented by Klauser [A. Klauser, et al., (1998)]. As result, we show that an average of 35% of all branches can be predicated in DCE and around 32% of all mispredictions fall into these branches. Rafael R. dos Santos, Tatiana Gadelha Serra dos Santos, Maurício L. Pilla, Philippe Olivier Alexandre Navaux, Sergio Bampi, Mario Nemirovsky |
SBAC-PAD | 5 |
| 2003 | Applying the GM/ID method in the analysis and design of Miller Amplifier, Comparator and GM-C PASS-B
Fernando da Rocha Paixão Cortes, Eric E. Fabris, Sergio Bampi |
VLSI-SOC | 3 |
| 2003 | Gray Encoded Arithmetic Operators Applied to FFT and FIR Dedicated Datapaths
Eduardo A. C. da Costa, José Monteiro 0001, Sergio Bampi |
VLSI-SOC | 3 |
| 2002 | A New Architecture for Signed Radix-2m Pure Array MultipliersabstractWe present a new architecture for signed multiplication which maintains the pure form of an array multiplier, exhibiting a much lower overhead than the Booth architecture. This architecture is extended for radix-2/sup m/ encoding, which leads to a reduction of the number of partial lines, enabling a significant improvement in performance and power consumption. The flexibility of our architecture allows for the easy construction of multipliers for different values of m, as opposed to the Booth architecture for which implementations for m > 2 are complex. The results we present show that the proposed architecture with radix-4 compares favorably in performance and power with the Modified Booth multiplier. We have experimented our architecture with different values of m and concluded that m = 4 minimizes both delay and power. Eduardo A. C. da Costa, Sergio Bampi, José Monteiro 0001 |
ICCD | 2 |
| 2000 | TAT transistors on SOT array for mixed analog/digital applicationsabstractThis paper presents the advantages of the trapezoidal association of transistor (TAT) technique for analog applications. Using only an association of fixed minimum size channel length transistors available on a Sea-of-Transistors (SOT), the TAT transistor achieves almost the same performance as its equivalent single transistor. Several structures of TATs and single transistors of electrically equivalent sizes, and operational transconductance amplifiers (OTA) in both methodologies were implemented to allow better performance comparisons. The experimental results obtained are herein shown for 1.0 /spl mu/m digital technology. Jung Hyun Choi, Sergio Bampi |
ISCAS | 2 |
| 2000 | A frame stream controller IPabstractThe IP (Intellectual Property) for frame stream control is a soft IP module designed to handle data frames of variable lengths, with different transmission rates. It works in a store and forward way, using an external SDRAM memory of arbitrary capacity. There is a receiver and a transmitter interface where frames are respectively stored and forwarded. The goal of this IP module is to handle the different kinds of protocols used in modern networks. Through the standard interface a bridge can be built between a synchronous and an asynchronous channel, or a faster and a slower rate port. A simple protocol is implemented on both reception and transmission interface, making the IP reusable for different designs. It can be used, for instance, to build an IP (Internet Protocol) LAN bridge. A protocol-specific function (Ethernet, Token Ring, etc.) has to be implemented in order to deliver the IP frame to the frame stream controller IP, and another specific protocol function to receive the IP frames from the frame controller and to transmit them through the physical media. Fernanda Lima Kastensmidt, Marcelo Barcelos, Juergen Rochol, Sergio Bampi, Ricardo Augusto da Luz Reis |
ISCAS | 4 |
| 1999 | Dynamically Reconfigurable Architecture for Image Processor ApplicationsabstractThis work presents an overview of the principles that underlie the speed-up achievable by dynamic hardware reconfiguration, proposes a more precise taxonomy for the execution models for reconfigurable platforms, and demonstrates the advantage of dynamic reconfiguration in the new implementation of a neighborhood image processor, called DRIP. It achieves a realtime performance, which is 3 times faster than its pipelined nonreconfigurable version Alexandro M. S. Adário, Eduardo L. Roehe, Sergio Bampi |
DAC | 3 |
| 1999 | OTA Amplifiers Design on Digital Sea-of-Transistors ArrayabstractThis paper presents measurement results of OTA (Operational Transconductance Amplifiers) designed in 1.0 /spl mu/m CMOS digital technology implemented in two different methodologies: in a fixed-size transistors array and in a full-custom design. Some characteristic parameters of OTA'S are compared with HSPICE simulations. Jung Hyun Choi, Sergio Bampi |
DATE | 2 |
| 1994 | A Timing Model for VLSI CMOS Circuits Verification and OptimizationabstractThis paper presents a semi-empirical model (TIME) for timing verification of CMOS VLSI digital circuits. Its novelty lies in the introduction of a new timing parameter (latency time), which is added to the usual /spl Rfr/C effective time constant. The effects that bear in the delay of all gates are included in the TIME model: type, geometry, body effect of transistors, slope of the gate input signals, capacitance loads and threshold logic voltage (VSW). Results are shown that compare circuit delays, analytically estimated with this model, within 10% of SPICE simulations and with more accuracy than Horowitz's analytical timing model. Other results illustrate the application of TIME to W-size optimization.> Luís Felipe Uebel, Sergio Bampi |
ISCAS | 2 |
| 1990 | An experiment on PC-based automated extraction of electrical parameters for VLSI MOSFETs: Methods, algorithms, and implementation
Sergio Bampi |
Microprocessing and Microprogramming | 1 |