VLDB 2026 Research / reviewers in the wild / expert
Bruno Zatt
dblp:33/2753
· DBLP profile ↗
95ranked-venue papers
9as first author
17since 2021 · last 2025
0000-0002-8045-957XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Approximate SRAM-Based Memory System for Intra-Frame Prediction in VVC EncodersabstractThis paper introduces a memory system for the Versatile Video Coding (VVC) standard, aimed at optimizing energy consumption in intra-frame prediction through approximate storage techniques. We propose a system that employs data approximation due to dynamic voltage scaling Static Random Access Memory (SRAM) in two distinct buffers: the Neighbor Samples Buffer (NeighSB) and the Original Samples Buffer (OrigSB). Our findings indicate the feasibility of approximate storage exploitation of our memory system, with minor coding efficiency drops at controlled error rates. As NeighSB and OrigSB presented distinct resilience behaviors, we evaluate different approximation strengths to obtain the most suitable coding efficiency, enabling still effective energy savings. Finally, we investigate the energy savings achieved through different SRAM operation levels, yielding reductions of 42%-62%. Yasmin Camargo, Matheus Isquierdo, Daniel Palomino 0001, Bruno Zatt, Felipe Sampaio |
ISCAS | 4 |
| 2025 | Machine Learning-Driven Multiple Transform Selection for Low-Complexity VVC EncodingabstractThe H.266/VVC video coding standard achieves significant compression rates, but its deployment faces important challenges due to high computational demands. This is particularly evident in the Multiple Transform Selection (MTS) tool, which applies various combinations of transforms to enhance the energy concentration of the prediction residue. This work first presents an analysis of the MTS tool and shows that the distribution of transform mode decisions varies significantly. Next, a machine learning approach based on decision trees is employed to avoid the need for exhaustively testing all transform mode combinations in MTS. The predictive models are trained on features extracted directly from the encoding process, thus avoiding additional processing overhead. The proposed strategy reduces the overall encoding time by 30.32% in the All-Intra configuration, with only a minor increase of 1.17% in BD-rate, achieving a better balance between efficiency and complexity. Caroline Camargo, Bianca Silveira, Bruno Zatt, Guilherme Corrêa 0001 |
ISCAS | 3 |
| 2025 | Exploiting Approximate SRAM for Energy-Efficient Integer Motion Estimation on VVC EncodersabstractVideo coding is a critical technology for enabling many modern applications. However, the high complexity of state-of-the-art encoders leads to energy consumption challenges, particularly in memory systems. This paper proposes an integer motion estimation (IME) system using approximate SRAM memories to enhance the energy efficiency of VVC encoders. The approximate SRAM memories are employed for both the current block and the search area buffers by reducing the supply voltage. The proposed IME system is evaluated using a customized tool that simulates the approximation effects on both buffers based on real energy measurements from a 28nm SRAM. This tool is integrated with VVenC, a fast VVC implementation, to assess the impact on coding efficiency and energy consumption. Experimental results show that the IME system using approximate SRAM can reduce energy consumption in reading operations by up to 55%, with a low impact on coding efficiency. Matheus Isquierdo, Felipe Sampaio, Bruno Zatt, Nikil Dutt, Daniel Palomino 0001 |
ISCAS | 3 |
| 2024 | A systematic literature review on video transcoding acceleration: challenges, solutions, and trends
Bruno Zatt, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
Multim. Tools Appl. | 2 |
| 2023 | High-Throughput and Multiplierless Hardware Design for the AV1 Local Warped MC InterpolationabstractMost of the current video codecs support only translational motion models. However, real motion is often complex and cannot be precisely estimated using only translational models. To handle complex motions like panning, zooming, scaling, shearing and rotation, AOMedia AV1 encoder counts with two tools, called Global and Local Warped Motion Compensation (LWMC). This paper presents two dedicated hardware designs for the AV1 LWMC interpolation filters. The presented hardware can process up to UHD 8K videos at 60fps. The architecture was synthesized for 40nm TSMC standard cells, requiring 454.37K gates with a power dissipation of 189.35mW. To the best of the authors’ knowledge, this is the first work in the literature targeting a dedicated hardware design for LWMC AV1 tool. Robson Domanski, William Kolodziejski, Wagner Penny, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 5 |
| 2023 | H.264-to-AV1 Video Transcoding Acceleration Based on Lightweight Machine LearningabstractVideo streaming platforms have been using the H.264/AVC standard for a long time, even though it was released almost 20 years ago and much more efficient codecs are currently available. The AOMedia Video 1 (AV1) format is an alternative with significant coding efficiency gains in comparison to H.264/AVC, besides being a royalty-free format. However, migrating legacy content from older to newer formats is a costly task, which requires long processing times. This work presents a solution for accelerating the H.264-to-AV1 transcoder based on machine learning. Sixteen decision tree models trained with data gathered during the H.264/AVC decoding and the AV1 encoding processes are proposed and implemented in the libaom reference software, leading to a complexity reduction of 18.96% at the cost of coding efficiency losses of 2.85% on average. To the best of the authors' knowledge, this is the first H.264-to-AV1 transcoding acceleration solution published in the literature. Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001 |
ISCAS | 3 |
| 2023 | Fast Intra Mode Decision Using Machine Learning for the Versatile Video Coding StandardabstractThis paper presents a fast intra mode decision solution for the VVC standard using machine learning. The idea is to reorder the evaluation of modes performed by the Rate-Distortion Optimization (RDO) process according to the modes occurrence rate. Based on the new evaluation order, three Decision Tree models were trained to skip the modes less likely to be chosen. The results show that the proposed solution achieves time savings of up to 15.57% with coding efficiency degradation of only 0.41% on average. When compared with related works, the proposed solution shows competitive results. Adson Duarte, Bruno Zatt, Guilherme Corrêa 0001, Daniel Palomino 0001 |
ISCAS | 2 |
| 2023 | High-Throughput Design for a Multi-Size DCT-II Targeting the AV1 EncoderabstractThis paper presents a dedicated multi-size hardware design for the Discrete Cosine Transform type II (DCT-II) of AV1 encoder. The DCT-II is one of four transform kernels supported by AV1; however, DCT-II is used in all configurations defined by AV1. Moreover, the 1D DCT-II can be applied for five different sizes ranging from 4-point up to 64-point. The 1D multi-size DCT-II was designed to process multiple transform sizes in parallel, always processing 64 samples in parallel for any size. The presented solution can process UHD 8K videos at 60 frames per second when running at 46.6 MHz, with a power dissipation of 44.48 mW and an area of 261.28 Kgates. To the best of authors' knowledge, this is the first work in the literature presenting a hardware design for the AV1 DCT-II transform. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 3 |
| 2023 | A High-Throughput Hardware Design for the AV1 Decoder IntrapredictionabstractThe Alliance for Open Media (AOMedia) (AV1) was released in 2018 as a royalty-free and open-source video codec. AV1 was developed by the AOMedia that is composed of many leading tech companies. AV1 has the goal to process ultrahigh definition (UHD) 8K (7680$\times4320$pixels) and 4K videos (3840$\times2160$pixels) and to achieve high coding efficiency, which leads to increased complexity when compared to other codecs in the market, such as VP9, HEVC, and H.264. This article presents the AV1 intraprediction decoder (AVID), a dedicated high-throughput hardware design for the AV1 decoder intraprediction supporting the AV1 68 prediction modes and 19 block sizes. The proposed architecture can decode UHD 4K videos at 120 frames/s in the worst case, requiring an operation frequency of 279.93 MHz and demanding a total area of 234.45 kgates with a power dissipation of 27.74 mW. The comparison with related works showed that AVID reached the smallest area and very competitive power results. To the best of the authors’ knowledge, this is the first article detailing the hardware design of a complete decoder for intraprediction targeting the AV1 codec. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Low-Complexity Multi-Type Tree Partitioning for Versatile Video Coding Based on Machine LearningabstractThe Versatile Video Coding (VVC) standard introduces new types of frame partitioning structures, such as QuadTrees (QT) and Multi-Type Trees (MTT). To achieve the best compression efficiency, for each block of pixels the encoder performs a recursive search over the partitioning possibilities, which also impacts significantly on the encoding complexity and processing time. This work proposes a machine learning-based solution for quick block partitioning decisions. A set of fourteen Random Forests were trained using data gathered during the encoding process and the models were employed to decide whether vertical and horizontal partitions are required for each candidate block. The proposed solution leads to an average encoding time reduction of 34.76% at the cost of a compression efficiency loss of 1.03%. Matheus Lindino, Bruno Zatt, Mateus Grellert, Guilherme Corrêa 0001 |
ICIP | 2 |
| 2022 | Fast Affine Motion Estimation for VVC using Machine-Learning-Based Early Search TerminationabstractThe Affine Motion Estimation (AME) was introduced in the Versatile Video Coding (VVC) standard to allow for the detection of non-translational transformations during inter-frame prediction. Although providing important coding efficiency gains, this new tool represents 43% of the motion estimation (ME) complexity. However, an analysis over the AME step shows that the Affine motion vectors are often generated without resulting in the best ME prediction. This paper proposes a AME early search termination based on supervised machine learning. Six Random Forest models were trained with features obtained during the encoding process to accurately predict whether the AME step should be executed, partially executed or skipped, avoiding unnecessary calculations. As result, the proposed solution achieves an average time saving of 46.94% in the AME step with a coding efficiency loss of only 0.18%. Adson Duarte, Luciano Volcan Agostini, Bruno Zatt, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Daniel Palomino 0001 |
ISCAS | 4 |
| 2022 | Improving Content-Aware Video Streaming in Congested Networks with In-Network ComputingabstractNetwork congestion and packet loss pose an ever-increasing challenge to video streaming. Despite the research efforts toward making video encoding schemes resilient to lossy network conditions, forwarding devices have not considered monitoring packet content to prioritize packets and minimize the impact of packet loss on video transmission. In this work, we advocate in favor of in-network computing employing a packet drop algorithm and an in-network hardware module to devise a solution for improving content-aware video streaming in congested network. Results show that our approach can reduce intra-predicted packet loss by over 80% at negligible resource usage and performance costs. Leonardo Gobatto, Mateus Saquetti, Cláudio Machado Diniz, Bruno Zatt, Weverton Luis da Costa Cordeiro, José Rodrigo Azambuja |
ISCAS | 4 |
| 2022 | A High-Throughput Design for the H.266/VVC Low-Frequency Non-Separable TransformabstractThis paper presents a high throughput hardware design for the Low-Frequency Non-Separable Transform (LFNST) of the Versatile Video Coding (H.266/VVC) standard. The LFNST is a secondary transform used to transform the coefficients already transformed by the DCT-II as primary transform over the residues from the directional intra prediction. The LFNST architecture was designed to process Ultra-High Definition (UHD) videos with $4098 \times 2160$ pixels (4K) at 60 frames per second. Our solution presents an area utilization of 99.13 kgates and a power dissipation of 38.50 mW, when running at 186.62 MHz and considering the worst-case operation (processing the LFNST $4\times 4$ through TU size of $4\times 4$). Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2022 | Video Decoder Improvements with Near-Data Speculative Motion Compensation ProcessingabstractVideo decoder implementations are still evolving as they directly affect a large fraction of embedded systems nowadays. In this context, Versatile Video Coding (VVC) brings increased compression efficiency, which comes with extra over-head in terms of computational effort and energy consumption. At the same time, emerging Near-Data Processing (NDP) architectures promise drastic time and energy cuts for applications with data streaming behavior. In this paper, a speculative Motion Compensation (MC) is proposed to enable video decoders improvements through the exploitation of NDP. We adopted a large-vector SIMD-based NDP system (called VIMA) that provides high-performance operations over 2 K vectors. The proposed strategy leverages the correlation between the prediction modes and the motion data between spatially neighboring blocks within a frame to speculatively perform the MC for an entire region of 2Kx128 samples. MC interpolation kernels were implemented using VIMA and x86 AVX-256 SIMD libraries. Our NDP-based kernel implementation allows speedup of $1.9\times$ to $22\times$ compared to the x86 baseline solutions. Stepping forward, based on a coalescence estimation, our strategy can properly handle interpolation misses, achieving MC performance improvements from 7% to 64%. Garrenlus de Souza, José Rodrigo Azambuja, Bruno Zatt, Marco A. Z. Alves, Sergio Bampi, Felipe Sampaio |
ISCAS | 3 |
| 2022 | FastInter360: A Fast Inter Mode Decision for HEVC 360 Video CodingabstractThis paper presents FastInter360, a fast inter mode decision algorithm for accelerating the encoding of ERP 360 videos. The development of FastInter360 involves an in-depth and comprehensive set of evaluations performed to understand the differences in the encoder’s behavior when encoding 360 and conventional videos. These evaluations showed that due to the texture distortions resulting from projection, the encoder presents a specific behavior when encoding 360 videos, making it more likely to use a recurrent set of encoding modes when processing 360 videos. Besides, the coding efficiency is less sensible to approximations in some encoding steps depending on the frame region. FastInter360 is then proposed to reduce the encoding complexity by exploiting these differences. FastInter360 comprises three algorithms that accelerate the encoding by performing early decision by SKIP mode, reducing integer motion estimation search range, and adjusting fractional motion estimation precision. Furthermore, each of these algorithms behaves according to distortion intensity, performing greater complexity reduction in more distorted regions. When employed altogether, these algorithms compose FastInter360, which is able to achieve an average complexity reduction of 22.84% with a coding efficiency loss of 0.652% BD-BR, on average, making FastInter360 competitive with literature works. Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Low-Power and High-Throughput Approximated Architecture for AV1 FME InterpolationabstractModern video encoders like the AOM Video 1 (AV1) implement several complex tools to allow the required high level of compression efficiency. The Fractional Motion Estimation (FME) is one of these tools and in AV1 the FME defines 90 different filters. To handle such complexity, hardware acceleration using approximate computing has become an alternative to be explored. This paper presents an approximate solution for the AV1 FME interpolation filters based on the approximation of the original filter coefficients intending to generate more hardware friendly coefficients. The approximated version was designed in hardware and can achieve real-time interpolation for UHD 8K videos at 30 frames per second, when synthesized using 40nm TSMC standard-cells technology. The designed architecture dissipates 26.79mW which represents more than 80% power reduction when compared to the original precise solution. The approximation implied in a small average coding efficiency degradation of 0.54% in BD-BR. When comparing with related works, this architecture reaches an expressive power reduction (2.1 to 4.8 times) even supporting more complex tools. Robson Domanski, William Kolodziejski, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 5 |
| 2021 | Energy-Throughput Configurable Design for Video Processing Binary Arithmetic EncoderabstractVideo encoding draws high research interest, due to the enormous demand for video traffic and real-time encoding for transmission. In video encoding standards such as HEVC (High-Efficiency Video Coding), the final step of the encoding stage is the CABAC (Context-Adaptive Binary Arithmetic Coding). The coding efficiency of the CABAC comes at the cost of increased computational complexity, especially for parallelization purposes, being the BAE (Binary Arithmetic Encoder) the critical part of CABAC. Thus, an important goal is to balance the real-time throughput requirements and the power/energy consumption in the design of BAE dedicated hardware. This work introduces a novel configurable high-throughput BAE design, named ET-BAE, utilizing a combination of a new modified Multiple-Bypass Bins Scheme (MBBS) and a power-saving approach into a single ASIC design with two-mode configuration. Synthesis and power-analysis results show that the configurable BAE design, the first of its kind with this feature, is more energy-efficient and less area consuming than utilizing non-configurable versions. The ET-BAE is able to accomplish the same real-time requirements of its competitors. Fábio Luís Livi Ramos, Bruno Zatt, Marcelo Schiavon Porto, Sergio Bampi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | RDE-MOGA: Automatic Selection of Rate-Distortion-Energy Control Points for Video Encoders Using Muti-Objetive Genetic AlgorithmabstractControlling energy consumption of video encoders is a complex multi-objective optimization problem of great importance. In this work we propose the RDE-MOGA, an multi-objective genetic algorithm capable of finding energetically efficient configurations for the HEVC encoder and replacing the current sensitivity analysis methodologies in the development of energy controllers. The utilization of our algorithm improved its efficiency in 60% whereas increasing the range of achievable reductions of the controller in at least 50%. Furthermore, the algorithm proved capable of sustaining 30% energy reduction at a cost of 3.45 BD-BR loss. Italo Machado, Marilton S. de Aguiar, Marcelo Schiavon Porto, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ICASSP | 6 |
| 2020 | Spatially Adaptive Intra Mode Pre-Selection for ERP 360 Video CodingabstractIn this work, we propose a spatially adaptive HEVC intra mode pre-selection for equirectangular (ERP) 360 video coding. The proposed technique exploits the spatial characteristics of 360 video in the ERP projection to reduce the complexity of intra prediction mode selection. The number of intra modes evaluated in Rate-Distortion Optimization is reduced based on a score technique that is adaptive to the frame region being encoded. Results show that the proposed technique achieves a complexity reduction of 16.5% with low coding efficiency penalties. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
ICASSP | 2 |
| 2020 | Memory Assessment Of Versatile Video CodingabstractThis paper presents a memory assessment of the next-generation Versatile Video Coding (VVC). The memory analyses are performed adopting as a baseline the state-of-the-art High-Efficiency Video Coding (HEVC). The goal is to offer insights and observations of how critical the memory requirements of VVC are aggravated, compared to HEVC. The adopted methodology consists of two sets of experiments: (1) an overall memory profiling and (2) an inter-prediction specific memory analysis. The results obtained in the memory profiling show that VVC access up to 13.4x more memory than HEVC. Moreover, the inter-prediction module remains (as in HEVC) the most resource-intensive operation in the encoder: 60%-90% of the memory requirements. The inter-prediction specific analysis demonstrates that VVC requires up to 5.3x more memory accesses than HEVC. Furthermore, our analysis indicates that up to 23% of such growth is due to VVC novel-CU sizes (larger than 64x64). Arthur Cerveira, Luciano Volcan Agostini, Bruno Zatt, Felipe Sampaio |
ICIP | 3 |
| 2020 | Low-Power and Memory-Aware Approximate Hardware Architecture for Fractional Motion Estimation Interpolation on HEVCabstractNowadays, current video coding standards like the High Efficiency Video Coding (HEVC) implement several complex coding tools, like the Fractional Motion Estimation (FME). An alternative to improve performance and save power is allying the hardware acceleration with approximate computing solutions, focusing on such complex tools. In this work, we present a low-power and memory-aware hardware architecture for the HEVC FME interpolator, proposing the development of two novel hardware designs for the interpolation filters, called Approximate Unified FME Filters (AUFF). These solutions exploit the usage of approximate computing at both algorithmic and data levels, leading to a reduction in dissipated power and memory bandwidth. The proposed design is capable of real-time interpolation of UHD (Ultra High Definition) 4K and 8K videos when synthesized using a 40 nm standard-cell library, with a power dissipation ranging from 22.04 to 62.06 mW. Wagner Penny, Guilherme Corrêa 0001, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Gabriel L. Nazar, Bruno Zatt |
ISCAS | 7 |
| 2020 | A Reliability-Oriented Machine Learning Strategy for Heterogeneous Multicore Application MappingabstractWe propose a methodology to transparently estimate near-optimal application mappings aiming at increasing the Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network (ANN) capable of estimating the vulnerability factor of RISC-V cores at runtime, which allows for efficient and dynamic application-to-core mappings targeting better MWTF and MWTF/energy tradeoffs. Results show that our ANN-based mapping yields very close-to-optimal solutions, with a difference in MWTF of only 3% when compared to the optimal mapping. When compared to a homogeneous architecture composed of only big cores, heterogeneous architectures may provide improvement in MWTF of up to 20.5% while impacting 12.2% on performance. Rafael Billig Tonetto, Hiago Rocha, Bruno Zatt, Antonio Carlos Schneider Beck, Gabriel L. Nazar |
ISCAS | 3 |
| 2020 | ERP-Based CTU Splitting Early Termination for Intra Prediction of 360 videosabstractThis work presents an Equirectangular projection (ERP) based Coding Tree Unit (CTU) splitting early termination algorithm for the High Efficiency Video Coding (HEVC) intra prediction of 360-degree videos. The proposed algorithm adaptively employs early termination in the HEVC CTU splitting based on distortion properties of the ERP projection, that generate homogeneous regions at the top and bottom portion of a video frame. Experimental results show an average of 24% time saving with 0.11% coding efficiency loss, significantly reducing the encoding complexity with minor impacts in the encoding efficiency. Besides, solution presents the best results considering the relation between time saving and coding efficiency when compared with all related works. Bernardo Beling, Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
VCIP | 4 |
| 2020 | 4D-DCT Hardware Architecture for JPEG Pleno Light Field CodingabstractThis paper presents a 4D-DCT hardware architecture for Light Field Coding according to the JPEG Pleno standard. It is composed of two instances of 2D-DCT engines and a novel 4D Transposition Memory organization. Experimentally-defined fixed-point representation and LSB pruning techniques are employed do reduce hardware area and power dissipation. The proposed architecture operates over 4D-hypercubes of up to 8x8x8x8 samples and reaches performance to process 30 Lytro-like light fields per second dissipating 145.32mW at 825.75MHz. This is the first known 4D-DCT hardware architecture for light field coding and demonstrates the feasibility of such solutions on real-world systems. Matheus Jahnke, Jones Goebel, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 7 |
| 2020 | Power/QoS-Adaptive HEVC FME Hardware using Machine Learning-Based Approximation ControlabstractThis paper presents a machine learning-based adaptive approximate hardware design targeting the fractional motion estimation (FME) of HEVC encoder. Hardware designs targeting multiple levels of approximation are proposed, by changing FME filters coefficients and/or discarding taps. The level of approximation is defined by a decision tree, generated taking into account the behavior of several parameters of the encoding in order to predict homogeneous blocks, more suitable for more aggressive approximation without significant losses on quality of service (QoS). Instead of applying a specific level of approximation over the full video, different approximate FME accelerators are dynamically selected. Such a strategy is able to provide up to 50.54% of power reduction while keeping the QoS losses at 1.18% BD-BR. Wagner Penny, Daniel Palomino 0001, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 4 |
| 2020 | Complexity and compression efficiency assessment of 3D-HEVC encoder
Mário Saldanha, Ruhan A. Conceição, Vladimir Afonso, Giovanni Avila, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001, Luciano Volcan Agostini |
Multim. Tools Appl. | 7 |
| 2019 | FastIntra360: A Fast Intra-Prediction Technique for 360-Degrees Video Codingabstract360-degrees videos represent a whole sphere and enable the user to feel as if he is inside the scene. These videos demand more data than conventional videos to be represented, therefore they also must be compressed to be handled properly. However, current video coding standards only process rectangular videos, thus 360 videos must be represented in a flat fashion to be encoded. There are several projections to perform this and the currently most used one is the equirectangular projection (ERP), which transforms each parallel from the sphere into a row of the rectangle, resulting in a faithful representation of the equatorial area, and a stretched representation of the polar regions. This stretching in the polar regions tends to impact the behavior of intra-frame prediction, which is used to exploit the spatial redundancies in each frame. Therefore, this paper proposes FastIntra360 to accelerate the encoding of 360 videos. FastIntra360 is implemented in HEVC video coding standard [1], which is a recently established standard and poses high computational demand. During the development of FastIntra360, a set of videos were encoded and the behavior of the intra-prediction throughout the frame was extracted. Then, a statistical analysis was conducted over such data and it concluded that when encoding the polar regions of the frame, the prediction modes which exploit horizontal directions are selected more frequently than the remaining modes, whereas in the center of the frame all prediction modes present similar occurrence rates. FastIntra360 exploits this behavior to reduce the number of prediction modes evaluated in different regions of the frame to accelerate the encoding. FastIntra360 is developed in two variants: one considering three bands and other considering five bands, where each band is a horizontal stripe of the frame. Each band divides the frame samples into three or five stripes and performs the statistical analysis over these stripes individually. Both implementations were evaluated and compared against the HEVC Test Model version 16.16 (HM-16.16) according to time reduction and coding efficiency (considering BD-BR), where BD-BR represents the bitrate increase of the proposed technique. Experimental results showed that both implementations present good performance, reaching up to 16.5% complexity reduction with negligible BD-BR, that is, they present considerable complexity reduction whereas posing no harm to the video quality. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
DCC | 2 |
| 2019 | Coding Tree Early Termination for Fast HEVC Transrating Based on Random ForestsabstractVideo transrating has become an essential task in streaming service providers that need to transmit and deliver different versions of the same content for a multitude of users that operate under different network conditions. As the transrating operation is comprised of a decoding and an encoding step in sequence, a huge computational cost is required in such large-scale services, especially when considering the use of complex state-of-the-art codecs, such as the High Efficiency Video Coding (HEVC). This work proposes an early-termination method for complexity reduction of the HEVC transrating based on Random Forests, which use features obtained from the HEVC decoding process to accelerate the coding tree decisions during the re-encoding process. Experimental results show that the proposed method achieves an average transrating time reduction of 47.09% at the cost of a negligible bitrate increase of 0.292%. Thiago Luiz Alves Bubolz, Mateus Grellert, Bruno Zatt, Guilherme Corrêa 0001 |
ICASSP | 3 |
| 2019 | Fast Hevc-to-Av1 Transcoding Based On Coding Unit Depth InheritanceabstractWith the advent of the recently launched AOMedia Video 1 (AV1) bitstream specification, there is currently a need for converting legacy content encoded with the state-of-the-art High Efficiency Video Coding (HEVC) standard to the new format. However, transcoding is a complex task composed of a decoding and an encoding process in sequence, which requires long processing times and high energy consumption. This paper proposes the first HEVC-to-AV1 transcoding solution, which is based on the high correlation between block size decisions in HEVC and AV1. The solution allows the AV1 encoder to inherit Coding Unit (CU) depth information from the HEVC bitstream to constrain the AV1 re-encoding process. Experimental results show an average transcoding time reduction of 35.41% at the cost of a compression efficiency loss of 4.54%. Bruno Zatt, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 2 |
| 2019 | High Throughput Hardware Design for AV1 Paeth and Smooth Intra ModesabstractDeveloped by AOMedia industry consortium and released in June 2018, AV1 is an open-source and royalty-free video coding format. The main goal of AV1 is to deliver substantial compression gains over state-of-the-art codecs such as VP9 and HEVC, while keeping a practical decoding complexity, hardware feasibility and its open and free status. This paper presents a high throughput hardware architecture for four important AV1 intra prediction coding modes: Paeth, Smooth, Smooth Vertical and Smooth Horizontal. The proposed architecture was designed to support all 19 block sizes specified by AV1 and to process every single combination of these blocks according to the 10-way partition tree, with a throughput of UHD 4K (3840×2160 pixels) videos at up to 30 frames per second. When synthesized to the TSMC 40nm cell library targeting a frequency of 648MHz, the proposed design used 109.57K gates and showed a power dissipation and an energy efficiency of 16.1mW and 1.23pJ/sample respectively. No other works were found in the literature describing hardware designs for AV1 intra prediction. Marcel Moscarelli Corrêa, Bianca Waskow, Bruno Zatt, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 3 |
| 2019 | TITAN: Tile Timing-Aware Balancing Algorithm for Speeding Up the 3D-HEVC Intra CodingabstractThis paper presents the Tile Timing-Aware balancing algorithm (TITAN) for speeding up the 3D-High Efficiency Video Coding (3D-HEVC) video encoding. The design of the TITAN algorithm is based on the premise that the encoding time of tiles partitioning of neighbor frames are similar. Therefore, based on the encoding time of the last-encoded frame, TITAN controls the tiles boundaries aiming to maximize the balance among tiles. Our software evaluation with TITAN implemented in 3D-HEVC Test Model 16.0 achieved an average of 6.2% higher speedup compared to the uniform-sized tiles. This is the first work in the literature proposing to speed up the 3D-HEVC video encoding by balancing the tiles workload. Mário Saldanha, Gustavo Sanchez, Bruno Zatt, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 3 |
| 2019 | Energy-Aware Motion and Disparity Estimation System for 3D-HEVC With Run-Time Adaptive Memory HierarchyabstractThe popularization of multimedia services has pushed forward the development of 2D/3D video-capable embedded mobile devices. Such devices require efficient energy/memory-management strategies to deal with severe memory/processing requirements and limited energy supply. Therefore, we propose a motion and disparity estimation (ME and DE) system—the most memory/processing demanding encoding steps—for the 3D High Efficiency Video Coding (3D-HEVC) standard. It was designed for low energy consumption, featuring a run-time adaptive memory hierarchy. The processing unit employs flexible coding order and optimizations to reduce the computational effort by exploring the inter-channel and inter-view redundancies. The memory hierarchy features window-based prefetching, data reuse, subsampling, and dynamic voltage scaling controlled by our depth-based dynamic search window resizing algorithm. Memory results demonstrate an average on-chip energy reduction of 79% in comparison to the widely used Level-C solution for a 45-nm technology. The proposed energy-aware ME and DE system dissipates 7.55 W while processing three HD 1080p views (video + depth) at 30 frames per second and presents a mean energy consumption of 0.107 J per access unit. To the best of our knowledge, this is the first work that proposes a real-time ME/DE system for the 3D-HEVC standard with an adaptive memory hierarchy. Vladimir Afonso, Ruhan A. Conceição, Mário Saldanha, Luciano Almeida Braatz, Murilo R. Perleberg, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt, Altamiro Amadeu Susin |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2019 | Fast Coding Unit Partition Decision for HEVC Using Support Vector MachinesabstractDespite the several speedup methods proposed in the literature, the computational complexity of High Efficiency Video Coding (HEVC) video encoding is still a problem. This paper proposes a fast coding unit (CU) partition decision for use in HEVC encoders based on support vector machine (SVM)-trained offline. The SVM classifiers, features, and training procedures are described in detail, and a justification for the use of SVMs is provided. The trained classifiers are incorporated into a modified reference encoder in the form of a fast CU partition decision algorithm, which decides if the exhaustive search for the best partition is continued or terminated prematurely. Using the proposed method, an average complexity reduction of 48% is achieved with a 0.48% Bjontegaard-Delta bitrate (BD-BR) loss using the random access coding configuration, 44% reduction with a 0.62% BD-BR loss for the Low Delay B, and a 41% reduction with a 0.6% BD-BR loss for the Low Delay P configuration. We also tested our approach under constant bitrate conditions, achieving a 47% reduction in encoding time with a 1.11% loss in the BD-BR. In addition, a decision threshold adaptation is also proposed to allow adjusting the rate-distortion/complexity trade-off of our solution. With this approach, the computational complexity reduction can be varied from 34.9% (with a 0.13% loss in the BD-BR) up to 52.4% (with a 1.11% BD-BR loss) using the random access configuration. Compared with the state-of-the-art solutions, our decision scheme outperforms the related works in terms of combined rate distortion and complexity. Mateus Grellert, Bruno Zatt, Sergio Bampi, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Hybrid Scratchpad Video Memory Architecture for Energy-Efficient Parallel HEVCabstractA hybrid scratchpad video memory (Hy-SVM) for energy-efficient tiles-parallelized high-efficiency video coding (HEVC) is presented here. The key ideas behind the Hy-SVM include: application-specific design and management; combined multiple levels of private and shared memories that jointly exploit intra-tile and inter-tiles data reuse; scratchpad memories (SPMs) as on-chip data storage; SRAM; and STT-RAM hybrid design. We propose a design methodology for the Hy-SVM that leverages application-specific properties to properly define the SPMs parameters. The inter-tiles data reuse potential of parallel HEVC is exploited by our run-time overlap prediction scheme, which identifies the redundant memory access behavior by analyzing monitored past frames encoding. Based on the predicted overlap characteristics, the Hy-SVM integrates memory access management units to control the access dynamics to the private/shared SPM levels. Furthermore, adaptive access management units (APMUs) can strongly reduce on-chip energy consumption due to the predicted overlap formation. The experimental results demonstrate the Hy-SVM overall energy savings of 11%-64% (4-tile) and 8%-46% (8-tile) when compared with related works. From the external memory perspective, the Hy-SVM can improve data reuse, resulting in 14%-59% of off-chip energy consumption (compared with no inter-tiles data reuse scenarios). In addition, our APMU contributes by reducing on-chip energy consumption of the Hy-SVM by 58%, on average. Thus, compared with related works, the Hy-SVM presents the lowest on-chip energy consumption. Moreover, the overhead of implementing our management units insignificantly affects the performance- and energy-efficiency of the Hy-SVM. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Jörg Henkel, Sergio Bampi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Octagonal-Axis Raster Pattern for Improved Test Zone Search Motion EstimationabstractTest Zone Search (TZS) is considered the current state-of-the-art fast Motion Estimation algorithm because it presents the best tradeoff between compression efficiency and complexity in comparison to the Full Search strategy. However, it is still one of the most computationally-demanding tools of current video coding standards, such as the High Efficiency Video Coding (HEVC). This paper presents an analysis on the search area opportunities and best match distributions in TZS, which led to the proposal of a novel search pattern in its most complex step, the Raster Search (RS). The new pattern, named Octagonal-Axis Raster Pattern (OARP), allowed an average complexity reduction of 61 % in TZS, with a negligible BD-rate increase of 0.0371 % in comparison to the original algorithm. Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001 |
ICASSP | 3 |
| 2018 | Learning-Based Complexity Reduction and Scaling for HEVC EncodersabstractThis article proposes a fast Coding Unit (CU) partition decision for use in HEVC encoders based on Decision Tree classifiers. The trees are employed in a modified low-complexity encoder that implements a fast CU partition decision algorithm. Using the proposed method, an average complexity reduction of 47.8% is achieved with a Bjontegaard Delta bitrate (BD-BR) loss of 0.24% in the Random Access coding configuration, and a 42.8% complexity reduction with a 0.19% BD-BR loss in the Low Delay B configuration. A decision threshold analysis is also presented to assess the rate-distortion-complexity trade-off of the proposed method at different complexity points, varying the complexity reduction from 28% (with a 0.04% loss in BD-BR) up to 60% (with a 3.6% BD-BR loss) using the Random Access configuration. A comparison with related works shows that the proposed method outperforms competing solutions in terms of both rate-distortion efficiency and complexity reduction. Mateus Grellert, Sergio Bampi, Guilherme Corrêa 0001, Bruno Zatt, Luís Alberto da Silva Cruz |
ICASSP | 4 |
| 2018 | LF-CAE: Context-Adaptive Encoding for Lenslet Light Fields Using HEVCabstractLight fields can outperform the representability of current imaging technologies by considering the angle of light-rays striking in the camera in addition to the color information. The additional information brings a set of challenges related to the amount of data required to represent light fields, arising the need for efficient compression schemes. This work proposes a novel and efficient scheme to encode lenslet light fields, called Light Fields Context-Adaptive Encoding(LF-CAE). LF -CAE is a video-based compression solution that defines a flexible and dynamic scheme to create intermediate video sequences from light fields in order to efficiently exploit the HEVC inter-frame encoder structure. This scheme reduces the inter-frame prediction residue leading to compression efficiency gains in light-field coding. LF -CAE reaches an average compression rate of 99.56% in relation to the uncompressed light field, with an average PSNR of 37.61dB. LF-CAE surpasses all related works that decompose the light field into an intermediate video sequence, reaching the highest PSNR and with BD-Rate gains ranging from 6.07% up to 20.50%. Ruhan A. Conceição, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 3 |
| 2018 | Hardware-Friendly Unidirectional Disparity-Search Algorithm for 3D-HEVCabstractThis paper presents a novel hardware-friendly Unidirectional Disparity-Search (UDS) algorithm for the 3D-HEVC. This algorithm explores the typical camera arrangements used in 3D-HEVC. UDS was evaluated in two operation points reaching a computational effort reduction from 32.7% to 61.8%, with a BD-Rate increase from 0.3123% to 0.4803%, when compared to TZS. Estimated hardware results showed memory-size and leakage-energy reductions of 66.7%, a dynamic-energy reduction from 34.5% to 62.1%, and an energy-consumption reduction from 32.7% to 61.8% in SAD calculations, when compared to TZS. To the best of the authors' knowledge, this is the first proposed DE algorithm that explores the 3D-HEVC typical camera arrangement. Vladimir Afonso, Altamiro Amadeu Susin, Murilo R. Perleberg, Ruhan A. Conceição, Guilherme Corrêa 0001, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 7 |
| 2018 | High-Throughput and Low-Power Integrated Direct/Inverse HEVC Quantization Hardware DesignabstractThis paper presents a high-throughput and low-power integrated HEVC direct/inverse quantization hardware design. The main focus of this design is to allow the evaluation of multiple coding modes during the residual encoding process of the HEVC for real-time Ultra-High Definition (UHD) video processing. The ASIC synthesis results, for a Nangate 45nm standard cell library, presented a maximum operational frequency of 1679.51MHz and a processing rate of 53.74 Gsps (giga samples per second). This throughput allows processing of real-time up to 72 coding modes for UHD 4K@60fps or up to nine coding modes for the UHD 8K@120fps while dissipating 369.37mW. Luciano Almeida Braatz, Bruno Zatt, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 2 |
| 2018 | OTED: Encoding Optimization Technique Targeting Energy-Efficient HEVC DecodingabstractThis work exploits the encoding for decoding concept by proposing an encoding optimization technique targeting energy-efficient HEVC decoding (called OTED). OTED changes the Rate-Distortion Optimization (RDO) calculation at the encoder side by adding the decoding energy estimation as a new variable to be considered. Along with this estimation, we use the Running Average Power Limit (RAPL) energy measurement tool to present real energy results and prove the efficiency of OTED. Experimental results show that the proposed algorithm can reach an energy reduction of up to 17.7% at the HEVC decoder with a small cost in coding efficiency. Besides being compliant with the standard HEVC decoder, the algorithm presents high levels of energy reduction for different encoding configurations. Furthermore, when compared with state-of-the-art works OTED presents the best relation between energy reduction consumption at the encoder side and coding efficiency. Douglas Corrêa, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ISCAS | 4 |
| 2018 | Configurable Cache Memory Architecture for Low-Energy Motion EstimationabstractThe popularization of mobile devices and the increased demand for video applications from these devices necessitates the design of efficient video encoders such as HEVC. Since the Motion Estimation (ME) is the most processing and memory intensive unit in a video encoder, our focus is in the communication between external memory and the ME unit. The TZS algorithm is widely used in video encoders and has an unpredictable behavior, which leads to an unknown pattern of memory accesses, making SPMs ineffective solutions, for example. Therefore, this work proposes a configurable cache memory architecture for fast ME algorithms. This cache has settings that suit different video encoding scenarios. Six optimal cache configurations were defined based on our evaluation considering 23 video sequences, 4 QPs, and 32 different cache settings. External memory bandwidth savings of up to 96.84% were reached, representing a reduction from 25.48GB/s to 548.53MB/s in the best case. When compared to Level-C SPM and to a static 16KB 8-way associative cache, the proposed configurable cache achieves energy savings of up to 86.91% and 78.09%, respectively. Anderson Martins, Wagner Penny, Matheus Weber, Luciano Volcan Agostini, Marcelo Schiavon Porto, Daniel Palomino 0001, Júlio C. B. de Mattos, Bruno Zatt |
ISCAS | 8 |
| 2018 | High-Throughput Binary Arithmetic Encoder using Multiple-Bypass Bins Processing for HEVC CABACabstractThe advance of massive video processing applications, devices and resolutions has led to new challenges in video encoding. The HEVC (High Efficiency Video Coding) standard emerges as one alternative in order to address the new video processing requirements. The HEVC allows only one type of entropy encoding algorithm, which is the CABAC (Context-Adaptive Binary Arithmetic Coding). The compression gains achieved by CABAC algorithm come at the cost of increasing complexity for implementation, due to intense data dependencies. The BAE (Binary Arithmetic Encoder) is the CABAC critical sub-block, in which the main part of the algorithm is executed. The present work proposes an 8-stage pipeline BAE architectural solution, named MB-BAE, with the addition of multiple-bypass bins processing, in order to increase throughput without compromising the critical path of the architecture. As a result, an average of 4.94 bins/cycle and around 2.6-Gbin/s of throughput are achieved in our BAE. This is the highest throughput found among related works in the literature, and being able to process 8K UHD videos with the lowest frequency when compared to the same related works. Fábio Luís Livi Ramos, Bruno Zatt, Marcelo Schiavon Porto, Sergio Bampi |
ISCAS | 2 |
| 2017 | Low-power and high-throughput hardware design for the 3D-HEVC depth intra skipabstractThis paper presents a low-power and high-throughput hardware design for the 3D-HEVC (Three Dimensional High Efficiency Video Coding) Depth Intra Skip coding tool. A strategy to reduce the computational effort was employed based on an analysis using the 3D-HEVC reference software. The proposed strategy consists of replacing the SVDC (Synthesized View Distortion Change) for the SAD (Sum of Absolute Differences) as the similarity criterion. This way, the number of arithmetic operations related with the similarity criterion is reduced over 71%, and a rendering process is avoided at the cost of only 0.21% increase in the BD-Rate. The hardware was described in VHDL and synthesized for ASIC technology. The synthesis results for the 45nm Nangate standard cells demonstrate that the architecture can process 60 UHD 2160p frames per second (five views) with a power dissipation of 19.57mW. Vladimir Afonso, Altamiro Amadeu Susin, Luan Audibert, Mário Saldanha, Ruhan A. Conceição, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 7 |
| 2017 | A multiplierless parallel HEVC quantization hardware for real-time UHD 8K video codingabstractOne step required several times for current video encoders is the residual coding loop, composed of the direct transformation, direct quantization, inverse quantization, and inverse transformation. These operations demand high throughput and low latency since their outputs must be processed by other steps of the coder. This paper proposes a high-throughput parallel and multiplierless hardware architecture for the HEVC direct quantization targeting real-time processing of Ultra-High Definition 8K videos. The proposed architecture support frequency dependent quantization steps. The binary multiplications were replaced by multiple constant multiplications in order to improve the throughput and to reduce the area and power dissipation. The developed design is able to process 32 samples in parallel, which represents one line of the biggest HEVC transform block. The ASIC synthesis results, obtained with Nangate 45nm standard cells library, show that the proposed architecture is able to quantize about 8 billion coefficients per second, when running at 186.6 MHz, with a gate count of 168,330. This throughput is enough to process UHD 8K videos at 120 fps. Luciano Almeida Braatz, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 3 |
| 2017 | High-throughput HEVC intrapicture prediction hardware design targeting UHD 8K videosabstractThis paper presents a high-throughput hardware architecture for the HEVC intrapicture prediction targeting the processing of UHD 8K (7680×4320 pixels) videos at 120 frames per second. The proposed design supports all intra prediction modes and all block sizes. It implements an internal mode decision algorithm that is more hardware friendly than RMD at the cost of a negligible 0.17% BD-rate impact (intra-only). When synthesized to the NanGate 45nm 0.95v cell library targeting a frequency of 529MHz, the proposed design used 4952K gates and showed a power dissipation and an energy efficiency of 363mW and 32.02pJ/sample respectively. Marcel Moscarelli Corrêa, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 2 |
| 2017 | Multiple early-termination scheme for TZ search algorithm based on data mining and decision treesabstractThe latest video compression standards, such as the H.264/AVC and the High Efficiency Video Coding (HEVC), provide fast Motion Estimation (ME) algorithms in their reference software aiming at complexity reduction. Test Zone Search (TZS) is the state-of-the-art fast ME algorithm, currently deployed in the reference HEVC encoder due to its great coding efficiency. However, ME is still one of the main sources of complexity in HEVC. This paper proposes an early-termination scheme for TZS, called e-TZS, based on an extensive data mining process on ME attributes. The data mining process allowed identifying the most relevant information during the encoding process to build a set of decision tree models that terminate TZS in different steps of its execution. The e-TZS scheme was implemented in the HEVC reference software and achieved average decision precision of 94.2%. Experimental results showed an average complexity reduction of 62.53% in TZS, with a negligible BD-rate increase of only 0.49%, in comparison to the original algorithm. Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
MMSP | 4 |
| 2017 | Energy-efficient motion estimation with approximate arithmeticabstractEnergy efficiency has become a primary concern in the design of multimedia digital systems, particularly when targeting mobile devices. Approximate computing is a highly promising approach to address this challenge. This paper presents an architectural exploration in a variable block size motion estimation (VBSME) architecture using imprecise Lower-Part-OR Adders (LOA). These adders were applied to Sum of Absolute Differences units (SAD) in order to reduce the energy consumption while introducing a minimum impact on the coding efficiency. Three VBSME architectures with LOA operators were developed by considering different imprecision levels. The conducted evaluations, performed using the High-Efficiency Video Coding standard (HEVC) reference software, showed that this technique introduces a negligible impact on the coding efficiency (between 0.6% and 2.5% increase of the BD-Rate). Nevertheless, when the designed architectures were synthesized for a 45nm standard cells technology, significant power savings were observed (between 7% and 11.5%, depending on the used LOA version), demonstrating the viability and significant gains of the proposed approach. Roger Endrigo Carvalho Porto, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto, Nuno Roma, Leonel Sousa |
MMSP | 3 |
| 2017 | Application-Guided Power-Efficient Fault Tolerance for H.264 Context Adaptive Variable Length CodingabstractThis paper presents a fault-tolerance technique for H.264's Context-Adaptive Variable Length Coding (CAVLC) on unreliable computing hardware. The application-specific knowledge is leveraged at both algorithm and architecture levels to protect the CAVLC process (especially context adaptation and coding tables) in a reliable yet power-efficient manner. Specifically, the statistical analysis of coding syntax and video content properties are exploited for: (1) selective redundancy of coefficient/header data of video bitstreams; (2) partitioning the coding tables into various sub-tables to reduce the power overhead of fault tolerance; and (3) run-time power management of memory parts storing the sub-tables and their parity computations. Experimental results demonstrate that leveraging application-specific knowledge reduces area and performance overhead by 2x compared to a double-parity table protection technique. For functional verification and area comparison, the complete H.264 CAVLC architecture is prototyped on a Xilinx Virtex-5 FPGA (though not limited to it). Muhammad Shafique 0001, Semeen Rehman, Florian Kriebel, Muhammad Usman Karim Khan, Bruno Zatt, Arun Subramaniyan 0001, Bruno Boessio Vizzotto, Jörg Henkel |
IEEE Trans. Computers | 5 |
| 2016 | Complexity reduction for 3D-HEVC depth map coding based on early Skip and early DIS schemeabstractThis paper presents a novel early Skip/DIS mode decision for 3D-HEVC depth encoding which aims at reducing the complexity effort of this process. The proposed solution is based on an adaptive threshold model, which takes into consideration the occurrence rate of both Skip and DIS modes. Occurrence analysis showed that the lower is the Skip and DIS Rate-Distortion cost, the higher is the probability of these modes being chosen. Furthermore, software evaluations showed that the proposed early Skip/DIS scheme is capable of reducing the depth coder complexity in 24.4% for a target hit rate of 99%, and in 33.7% for a target hit rate of 95%, leading to a negligible coding efficiency penalty in both scenarios. Ruhan A. Conceição, Giovanni Avila, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 5 |
| 2016 | High-throughput and memory-aware hardware of a sub-pixel interpolator for multiple video coding standardsabstractReal-time operation and low-power dissipation in video coding systems have become important research challenges, especially in mobile devices with limited battery and computational resources. There are many video coding standards coexisting in the market nowadays, so it is important for current devices to support different video coding standards. This paper presents a multi-standard luminance sub-samples interpolator hardware design for the Motion Compensation (MC) and Fractional Motion Estimation (FME), with support to MPEG-2/4, H.264/AVC, HEVC, and AVS/2 video coding standards. Our design is able to save hardware resources through an optimized filter organization, totally compliant with the focused standards and capable to interpolate samples for UHD 4320p@60fps at real time. The 45nm standard-cell library implementation dissipates 10mW, when processing according MPEG-2 standard, up to 46.4mW when processing AVS2. Guilherme Paim, Jones Goebel, Wagner Penny, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 4 |
| 2016 | An efficient sub-sample interpolator hardware for VP9-10 standardsabstractThis paper presents a hardware design for the sub-sample interpolator used in FME (Fractional Motion Estimation) and MC (Motion Compensation) stages according to the VP9 and VP10 video-coding standards. The proposed architecture is able to save hardware resources through an optimized-filter organization whereas reaching high-throughput and low-power dissipation. The hardware design was described in Verilog and synthesized for ASIC technology. The synthesis results were generated for 45nm Nangate standard cells and demonstrate that the developed architecture is able to process 2160p@60fps videos with a power dissipation of 2.34mW focusing on a VP9-10 decoder. Guilherme Paim, Wagner Penny, Jones Goebel, Vladimir Afonso, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 7 |
| 2016 | Pareto-based energy control for the HEVC encoderabstractThe current state-of-art video coding standard, the High Efficiency Video Coding (HEVC), brings many innovations as a way to improve the coding performance. However, the improvement on performance also brought higher computational effort and energy consumption. Since most of devices that handle digital videos are battery powered, the energy consumption became an important issue that demands efficient solutions. This way, controlling energy consumption is strongly desirable to adapt the encoding process to the energy availability. This goal is a hard task due the heterogeneous dynamic behavior of HEVC encoder. This work presents the development of a Pareto-based dynamic energy controller for the HEVC encoder, reaching up to 70% energy saving with small losses on coding efficiency for most of the cases. Wagner Penny, Italo Machado, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt |
ICIP | 5 |
| 2016 | Speedup-aware history-based tiling algorithm for the HEVC standardabstractThis paper proposes a history-based tiling algorithm aiming at the increase of speedup when using Tiles. The algorithm is composed of two independent steps that use workload history information to define the vertical and horizontal boundaries of the Tiles. The workload distribution of previous frames are used as reference to perform the tiling of the current frame exploiting the temporal similarity between neighboring frames. Experimental results show that the proposed algorithm outperforms the speedup when compared to uniform tiling by 6.85% on average for tested sequences, besides, there is no significant complexity increase and similar coding efficiency results. When compared to related works the proposed solution also presents better speedup results. Iago Storch, Daniel Palomino 0001, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 3 |
| 2016 | An HEVC multi-size DCT hardware with constant throughput and supporting heterogeneous CUsabstractThis paper presents an efficient hardware design for the Discrete Cosine Transform (DCT) of High Efficiency Video Coding standard (HEVC). This hardware supports all HEVC transform sizes: 4×4, 8×8, 16×16, and 32×32 including any combination of the Transform Unit (TU) sizes. The proposed DCT architecture has a constant throughput of 32 coefficients per cycle, independently of the transform sizes combination. The architecture was synthesized for a Nangate 45nm standard-cell library and the power analysis was made considering real input vectors. The synthesis results show a very good tradeoff between area, power dissipation and processing rates. The architecture is able to process 1.6G coeff/s when running at 50MHz dissipating 24.2 mW. These results allow a processing rate of 30 HD 1080p frames per second when evaluating 17 HEVC prediction modes. Jones Goebel, Guilherme Paim, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2016 | Energy-aware cache assessment of HEVC decodingabstractThis paper presents a thorough analysis of energy consumption of a software HEVC decoder. The evaluation utilizes a framework developed herein specifically to estimate the energy consumption in all levels of cache hierarchies. Our framework is based on analytical models combined with memory profiling; tools. Energy analyses of several cache hierarchies executing HEVC decoding with different input bit streams were carried out. Our results point to the most suited cache parameters for each video resolution. The energy was estimated for a 32nm CMOS technology. Our study includes different tradeoffs between energy efficiency and capacity, associativity, and main memory bandwidth. Our detailed analysis shows that the higher are the cache features, the more efficient is the energy consumption. The main memory bandwidth evaluation shows that the energy consumption increases with the main memory bandwidth requirement. Full HD video resolutions require up to 90 times higher bandwidth and 57 times more energy than class D resolutions. Eduarda Monteiro, Mateus Grellert, Sergio Bampi, Bruno Zatt |
ISCAS | 4 |
| 2016 | Complexity-scalable HEVC encodingabstractHEVC encoders impose several challenges in resource-constrained embedded applications, especially under real-time and battery constraints. This paper proposes a complexity-scalable encoder that is able to achieve considerable time savings while maintaining an efficient rate-distortion-complexity tradeoff. To design the system, a complexity analysis of HEVC-supported parameters as well as new ones introduced in this work is presented. To build the configurations for each target savings, a Complexity Target Satisfaction algorithm was designed. This algorithm was able to reduce the optimization space by approximately 600 times, producing efficient configurations with up to 90% time savings. The complete system was implemented and tested against a state-of-the-art complexity management implementation. The results proved the efficiency of our solution, as it achieves more time savings and better compression for savings of 60% and higher. Mateus Grellert, Sergio Bampi, Bruno Zatt |
PCS | 3 |
| 2015 | Approximation-aware Multi-Level Cells STT-RAM cache architectureabstractCurrent manycore processors exhibit large on-chip last-level caches that may reach sizes of 32MB - 128MB and incur high power/energy consumption. The emerging Multi-Level Cells (MLC) STT-RAM memory technology improves the capacity and energy efficiency issues of large-sized memory banks. However, MLC STT-RAM incurs non-negligible protection overhead to ensure reliable operations when compared to the Single-Level Cells (SLC) STT-RAM. In this paper, we propose an approximation-aware MLC STT-RAM cache architecture, which is partially-protected to restrict the reliability overhead and in turn leverages variable resilience characteristics of different applications for adaptively curtailing the protection overhead under a given error tolerance level. It thereby improves the energy-efficiency of the cache while meeting the reliability requirements. Our cache architecture is equipped with a latency-aware hardware module for double-error correction. To achieve high energy efficiency, approximation-aware read and write policies are proposed that perform approximate storage management while tolerating some errors bounded within the user-provided tolerance level. The architecture also facilitates runtime control on the quality of applications' results. We perform a case study on the next-generation advanced video encoding that exhibit memory-intensive functional blocks with varying resilience properties and support for parallelism. Experimental results demonstrate that our approximation-aware MLC STT-RAM based cache architecture can improve the energy efficiency compared to state-of-the-art fully-protected caches (7%-19%, on average), while incurring minimal quality penalties in the output (-0.219% to -0.426%, on average). Furthermore, our architecture supports complete error protection coverage for all cache data when processing non-resilient application. The hardware overhead to implement our approximation-aware management negligibly affects the energy efficiency (0.15%-1.3% of overhead) and the access latency (only 0.02%-1.56% of overhead). Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
CASES | 3 |
| 2015 | A multi-standard interpolation hardware solution for H.264 and HEVCabstractAttending real-time constraints in video coding systems represents a big challenge for nowadays systems, especially for high definition videos at mobile systems. The Fractional Motion Estimation (FME) and Motion Compensation (MC) are responsible for a large share of processing effort in both state-of-the-art video coding standards, the High Efficiency Video Coding (HEVC), and its predecessor, the H.264. This work proposes a multi-standard hardware solution for the fractional sample interpolation used in FME/MC processing of the HEVC and H.264 standards. The hardware design is composed of four IP (Intellectual Property) cores able to process 1080p@60fps videos independently. The whole architecture can process 2160p@60fps with 80.69mW, considering bi-prediction. Henrique Maich, Guilherme Paim, Vladimir Afonso, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ICIP | 5 |
| 2015 | Rate-distortion and energy performance of HEVC and H.264/AVC encoders: A comparative analysisabstractA quantitative, systematic, and detailed analysis of the energy impacts of the tools that comprise two of the most recent video coding standards: the High Efficiency Video Coding (HEVC) and the H.264/AVC is presented. Our comparative study measures the energy consumption effects of important video-coding parameters, like Search Range (SR), Quantization Parameter (QP), and video resolution on both encoders. The obtained results for HEVC showed, for the Random Access (RA) prediction structure, gains of 25% in BD-Rate over H.264/AVC at the expense of 17% higher energy consumption. A new metric we defined herein, called BD-Energy, was used in the SR analysis, and the results from this investigation showed HEVC achieved an energy consumption up to 37.6% higher for a BD-Rate gain of 32.2%. The QP analysis demonstrated that the energy consumption gap between both encoders varies greatly as QP increases, resulting in a 15.08% difference from QP 22 to QP 37, on average. The major finding from our work is that the HEVC encoder presents better results in the energy/compression trade-off, but this efficiency is reduced as encoding becomes more complex, as our results discovered that the HEVC energy consumption scales faster. Eduarda Monteiro, Mateus Grellert, Sergio Bampi, Bruno Zatt |
ISCAS | 4 |
| 2015 | Complexity reduction for the 3D-HEVC depth maps codingabstractThis paper presents a qualitative discussion of the depth maps properties that can be considered to achieve complexity reduction for 3D-High Efficiency Video Coding (3D-HEVC) depth maps coding. Both intra and inter-frame predictions are considered in this discussion that conduced to the proposition of two simple complexity reduction techniques: the Simplified Edge Detector (SED) and the Diamond Search (DS) simplified inter-prediction. The SED anticipates the blocks that are likely to be better predicted by the HEVC intra-prediction, avoiding evaluations of Depth Modeling Modes (DMM). The DS and SED were compared to anchor results and experimental analysis showed that the proposed algorithms are able to achieve a time saving of 11.3% encoding time reduction, with acceptable impact on the BD-Rate of the synthesized views of 0.6%. Mário Saldanha, Gustavo Sanchez, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 3 |
| 2015 | A real-time architecture for reference frame compression for high definition video codersabstractCurrent battery-powered devices that manipulate digital videos must consider the energy consumption of this process as an important issue, especially when high or ultra-high definition videos are handled. In this scenario, this paper proposes a solution to reduce the energy consumption in video coding systems by reducing the external memory communication during the motion estimation. The scheme presented in this paper is called Differential Reference Frame Coder and it implements an algorithm that combines two techniques to reduce the memory bandwidth: a differential coding based on a simplified intra-prediction process, to reduce the spatial redundancy of the reconstructed samples, and a semi-fixed length coding applied in the residues generated by the differential coding step. This solution reaches an average lossless compression ratio higher than 57% for the evaluated HD 1080p video sequences whereas supporting random access to reference frame blocks. The proposed hardware architectures (Coder and Decoder) were described in VHDL and synthesized targeting ASIC for 65nm and 180nm TSMC standard-cell libraries. The results show that with 65nm, the architectures are able to process UHD 2160p (3840×2160 samples) at 30 fps or HD 1080p (1920×1080 samples) at 120 fps with a power dissipation of 0.885mW. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2014 | dSVM: Energy-efficient distributed Scratchpad Video Memory Architecture for the next-generation High Efficiency Video CodingabstractAn energy-efficient distributed Scratchpad Video Memory Architecture (dSVM) for the next-generation parallel High Efficiency Video Coding is presented. Our dSVM combines private and overlapping (shared) Scratchpad Memories (SPMs) to support data reuse within and across different cores concurrently executing multiple parallel HEVC threads. We developed a statistical method to size and design the organization of the SPMs along with a supporting memory reading policy for energy efficiency. The key is to leverage the HEVC and video content knowledge. Furthermore, we integrate an adaptive power management policy for SPMs to manage the power states of different memory parts at run time depending upon the varying video content properties. Our experimental results illustrate that our dSVM architecture reduces the overall memory energy consumption by up to 51%-61% compared to parallelized state-of-the-art solutions [11]. The dSVM external memory energy savings increase with an increasing number of parallel HEVC threads and size of search window. Moreover, our SPM power management reacts to the current video properties and achieves up to 54% on-chip leakage energy savings. Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
DATE | 3 |
| 2014 | A low-complexity and lossless reference frame encoder algorithm for video codingabstractThis paper presents a lossless coding solution to reduce the large overhead of external memory communication during the motion estimation process in current video coders. Our solution is called Differential Reference Frame Coder (DRFC), and uses two techniques together to compress the reference frame: a differential coding based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples, and a simple VLC applied to differential coding residues. The proposed solution reaches an average compression rate higher than 45% for the evaluated HD 1080p video sequences. This is a lossless and low-complexity solution, and could easily be implemented in hardware. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ICASSP | 4 |
| 2014 | Energy-efficient architecture for advanced video memoryabstractAn energy-efficient hybrid on-chip video memory architecture (enHyV) is presented that combines private and shared memories using a hybrid design (i.e., SRAM and emerging STT-RAM). The key is to leverage the application-specific properties to efficiently design and manage the enHyV. To increase STT-RAM lifetime, we propose a data management technique that alleviates the bit-toggling write occurrences. An adaptive power management is also proposed for static-energy savings. Experimental results illustrate that enHyV reduces on-chip static memory energy compared to SRAM-only version of enHyV and to state-of-art AMBER hybrid video memory [9] by 66%-75% and 55%-76%, respectively. Furthermore, negligible external memory energy consumption is required for reference frames communication (98% lower than state-of-the-art Level C+ technique [18]). Our data management significantly improves the enHyV STT-RAM lifetime, achieving 0.83 of normalized lifetime (near to the optimal case). Our hybrid memory design and management incur low overhead in terms of latency and dynamic energy. Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
ICCAD | 3 |
| 2014 | Complexity reduction for 3D-HEVC depth maps intra-frame prediction using simplified edge detector algorithmabstractThis paper presents a new mode decision for the depth maps intra-frame prediction in 3D-HEVC. The proposed technique decides if the traditional High Efficiency Video Coding-based (HEVC) intra-frame prediction should be performed or skipped. This technique is inspired by the fact that traditional intra-frame prediction may generate artifacts in the synthesized views when an edge is encoded. The Simplified Edge Detector (SED) algorithm has been proposed to classify if a block contains an edge or a nearly constant region demanding a minimum processing overhead. Through software evaluations, SED algorithm was capable to obtain an average complexity reduction of 23.8% for depth maps coding with no quality losses. Gustavo Sanchez, Mário Saldanha, Gabriel Balota, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 4 |
| 2014 | A new differential and lossless Reference Frame Variable-Length Coder: An approach for high definition video codersabstractThis paper presents a novel solution for external memory bandwidth reduction in video coding systems. The approach is based on reference frame compression, using a differential coding and a hardware-aware adaptation of the traditional Huffman algorithm, besides, it is a lossless solution fully compliant to state-of-art video coding standards, as H.264/AVC and HEVC. This solution is called DRFVLC (Differential Reference Frame Variable-Length Coder) and it uses differential coding to concentrate the samples values distribution. With the samples concentrated, an efficient static Huffman coding is applied to represent them in fewer bits. The DRFVLC reaches an average compression rate higher than 60% for the evaluated HD 1080p video sequences. This compression rate also indicates the external memory bandwidth reduction achieved with our technique. This solution can be easily implemented in hardware demanding one differentiator and a simple variable-length coder. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ICIP | 4 |
| 2014 | Power efficient and high troughtput multi-size IDCT targeting UHD HEVC decodersabstractThis paper is focused on the inverse transforms defined in the HEVC (High Efficiency Video Coding) standard. The HEVC standard allows the use of four transform sizes, including novel transforms applied over bigger block sizes (16×16 and 32×32). The hardware architecture presented in this paper was planned to reach real-time processing (at 30 frames per second) for ultra-higher solution videos, exploiting high level of parallelism. As a secondary goal, the architecture was also planned to reach low cost in terms of hardware consumption and power dissipation. Thus, the architecture was designed in a purely combinational way, using a multiplierless approach and employing an optimization algorithm through operations reuse and sub-expressions sharing. The synthesis targeted an Altera Stratix V FPGA and ASIC 90nm standard-cells technology. The synthesis results show that the designed architecture has the best performance results among all related works, being able to achieve real-time decoding for UHD videos (7680×4320 pixels) with a power consumption from 33.8mW to 339.2 mW. Ruhan A. Conceição, J. Claudio de Souza, Ricardo Jeske, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 5 |
| 2014 | Memory bandwidth reduction for H.264 and HEVC encoders using lossless reference frame codingabstractThis paper presents a hardware-efficient algorithm for external memory bandwidth reduction focusing on the state-of-the-art video encoders, like H.264/AVC and HEVC. The proposed approach is a lossless solution based on an adaptation of the traditional Huffman algorithm. This solution is entitled RFCAVLC8T (Reference Frame Context Adaptive Variable-Length Coder with 8 Tables) and is based on the use of off-line defined static Huffman tables. The RFCAVLC8T is a hardware-efficient version of the Huffman algorithm that employs eight static tables to avoid the cost of the on-the-fly Huffman statistical analysis. The best table to encode a block is selected at run time using a context evaluation, resulting in a context-adaptive configuration. The use of RFCAVLC8T reaches an average compression rate higher than 35% for the evaluated video sequences, with computational cost of a single VLC. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2014 | Content-driven memory pressure balancing and video memory power management for parallel high efficiency video codingabstractWe present a novel content-driven memory pressure balancing and video memory power management scheme for parallel High Efficiency Video Coding (HEVC). The key is to leverage the application-specific knowledge to balance the (instant) access pressure on Scratchpad-based Video Memories (SVMs) for parallelized video processing. Our scheme accurately predicts the memory requirements of each processing core based on monitored memory usage and leverages this knowledge to perform a categorization of different video regions. Afterwards, it employs an adaptive policy for memory pressure balancing by rescheduling encoding of different video blocks based on their categories. This balancing also facilitates our scheme to perform efficient power-gating of unused parts of SVMs. Experimental results show that our scheme reduces the variations in the memory pressure by 37%-83% when compared to the traditional raster scan processing for 4- and 16-core parallelized HEVC encoder. Our content-driven power management saves 56% (on average) of SVM leakage energy. Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
ISLPED | 3 |
| 2014 | Sample adaptive offset filter hardware design for HEVC encoderabstractThis work presents a hardware design for the Sample Adaptive Offset filter, which is an innovation brought by the new video coding standard HEVC. The architectures focus on the encoder side and include both classification methods used in SAO, the Band Offset and Edge Offset, and also the statistical calculations for the offset generation. The proposed architectures feature two sample buffers, classification units for both SAO types and the statistical collection unit. The architectures were described in VHDL and synthesized to an Altera Stratix V FPGA. The synthesis results show that the proposed architectures achieve 364MHz and are capable to process 44 QFHD (3840×2160) frames per second using 8,040 ALUTs of the target device hardware resources. Fabiane Rediess, Ruhan A. Conceição, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
VCIP | 3 |
| 2014 | A complexity reduction algorithm for depth maps intra prediction on the 3D-HEVCabstractThis paper proposes a complexity reduction algorithm for the depth maps intra prediction of the emerging 3D High Efficiency Video Coding standard (3D-HEVC). The 3D-HEVC introduces a new set of specific tools for the depth map coding that includes four Depth Modeling Modes (DMM) and these new features have inserted extra effort on the intra prediction. This extra effort is undesired and contributes to increasing the power consumption, which is a huge problem especially for embedded-systems. For this reason, this paper proposes a complexity reduction algorithm for the DMM 1, called Gradient-Based Mode One Filter (GMOF). This algorithm applies a filter to the borders of the encoded block and determines the best positions to evaluate the DMM 1, reducing the computational effort of DMM 1 process. Experimental analysis showed that GMOF is capable to achieve, in average, a complexity reduction of 9.8% on depth maps prediction, when evaluating under Common Test Conditions (CTC), with minor impacts on the quality of the synthesized views. Gustavo Sanchez, Mário Saldanha, Gabriel Balota, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
VCIP | 4 |
| 2013 | Energy-efficient memory hierarchy for motion and disparity estimation in multiview video codingabstractThis work presents an energy-efficient memory hierarchy for Motion and Disparity Estimation on Multiview Video Coding employing a Reference Frames-Centered Data Reuse (RCDR) scheme. In RCDR the reference search window becomes the center of the motion/disparity estimation processing flow and calls for processing all blocks requesting its data. By doing so, RCDR avoids multiple search window retransmissions leading to reduced number of external memory accesses, thus memory energy reduction. To deal with out-of-order processing and further reduce external memory traffic, a statistics-based partial results compressor is developed. The on-chip video memory energy is reduced by employing a statistical power gating scheme and candidate blocks reordering. Experimental results show that our reference-centered memory hierarchy outperforms the state-of-the-art [7][13] by providing reduction of up to 71% for external memory energy, 88% on-chip memory static energy, and 65% on-chip memory dynamic energy. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel |
DATE | 2 |
| 2013 | Content-adaptive reference frame compression based on intra-frame prediction for multiview video codingabstractThis paper presents a content-adaptive reference frame compression scheme to alleviate the large overhead of external memory communication during the Motion and Disparity Estimation process in Multiview Video Coding (MVC). Our scheme is based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples. The intra-prediction residue is compressed by a path composed of non-linear quantization and Huffman-based entropy encoder. Four different quantization strengths and Huffman tables were statistically defined. They are dynamically selected according to a content adaptation strategy, which classifies the original blocks based on their spatial homogeneity. Experimental results show that the proposed content-adaptive compression scheme is able to reduce the external memory accesses by up to 63% along with negligible losses in the MVC encoder rate-distortion performance. Compared to the best available related work [12] our content-adaptive reference frame compression achieves 39% reduced external memory accesses, while still providing a BD-PSNR increase of 0.03dB. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Jörg Henkel, Sergio Bampi |
ICIP | 2 |
| 2013 | Adaptive content-based Tile partitioning algorithm for the HEVC standardabstractThis paper proposes a content-based Tile partitioning algorithm designed to exploit video properties in the tiling process, aiming at the reduction of the coding losses generated by the use of Tiles. In the proposed algorithm two steps are performed to define the vertical and horizontal Tile boundaries willing to group the high correlated samples into the same Tile partition. The definition of the Tiles' boundaries is performed by analyzing the variance map extracted from the raw picture information. Experimental results have shown that the proposed algorithm is able to reduce the inherent coding efficiency losses of using Tiles when compared to the conventional uniform spaced Tile partitions. Cauane Blumenberg, Daniel Palomino 0001, Sergio Bampi, Bruno Zatt |
PCS | 4 |
| 2013 | Model Predictive Hierarchical Rate Control With Markov Decision Process for Multiview Video CodingabstractThis paper presents a novel hierarchical rate control (HRC) for the Multiview Video Coding standard targeting improved bandwidth usage and high video quality. The HRC is designed to jointly address the rate control at both frame level and basic unit (BU) level. The proposed scheme is able to exploit the bitrate distribution correlation with neighboring frames to efficiently predict the future bitrate behavior by employing a model predictive control that defines a proper control action through quantization parameter (QP) adaptation. To provide a fine-grained tuning, the QP is further adapted within each frame by a Markov decision process implemented at BU level able to take into consideration a map of the regions of interest. A coupled frame/BU level feedback is performed in order to guarantee the system consistency. Experimental results show the superiority of our HRC compared to state-of-the-art solutions in terms of bitrate allocation accuracy and rate distortion while delivering smooth video quality at frame and BU levels. Bruno Boessio Vizzotto, Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Adaptive power management of on-chip video memory for multiview video codingabstractAn adaptive power management of on-chip video memory for Multiview Video Coding is presented. It leverages texture, motion and disparity properties of objects and their correlations in the 3D-neighborhood. It groups different Macroblocks of a frame and predicts the highly-probable motion/disparity search direction in order to power-gate idle memory regions. Exploited are the statistical properties of Macroblock groups to predict idle sectors. Our approach achieves on average 32% and 61% energy reduction (averaged over various video sequences) compared to state-of-the-art DSW [7] and Level C [12], respectively. The Motion/Disparity Estimation architecture with video memory and power management scheme is implemented using an ASIC flow (IBM-65nm Low-Power technology) and it processes 4-view [email protected] Muhammad Shafique 0001, Bruno Zatt, Fabio Leandro Walter, Sergio Bampi, Jörg Henkel |
DAC | 2 |
| 2012 | Power-efficient error-resiliency for H.264/AVC Context-Adaptive Variable Length CodingabstractTechnology scaling has led to unreliable computing hardware due to high susceptibility against soft errors. In this paper, we propose an error-resilient architecture for Context-Adaptive Variable Length Coding (CAVLC) in H.264/AVC. Due to its context-adaptive nature and intricate control flow CAVLC is very sensitive to soft errors. An error during the CAVLC process (especially during the context adaptation or in VLC tables) may result in severe mismatch between encoder and decoder. The primary goal in our error-resilient CAVLC architecture is to protect codeword/codelength tables and context adaptation in a reliable yet power efficient manner. For reducing the power over-head, the tables are partitioned in various sub-tables each protected with variable-sized parity. Moreover, for further power reduction, our approach incorporates state-retentive power-gating of different sub-tables at run time depending upon the statistical distribution of syntax elements. Compared to the unprotected case, our scheme provides a video quality improvement of 18dB (averaged over various fault injection cases and video sequences) at the cost of a 35% area overhead and 45% performance overhead due to the error-detection logic. However, partitioned sub-tables increase the potential for power-gating, thus bring a leakage energy saving of 58%. Compared to state-of-the-art table protection, our scheme provides 2x reduced area and performance overhead. For function-al verification and area comparison, the architecture is prototyped on a Xilinx Virtex-5 FPGA, though not limited to it. For the soft errors experiments, evaluation of error-resiliency and power efficiency, we have developed a fault injection and simulation setup. Muhammad Shafique 0001, Bruno Zatt, Semeen Rehman, Florian Kriebel, Jörg Henkel |
DATE | 2 |
| 2012 | Real-time block matching motion estimation onto GPGPUabstractThis work presents an efficient method to map Motion Estimation (ME) algorithms onto General Purpose Graphic Processing Unit (GPGPU) architectures using CUDA programming model. Our method jointly exploits the massive parallelism available in current GPGPU devices and the parallelization potential of ME algorithms: Full Search (FS) and Diamond Search (DS). Our main goal is to evaluate the feasibility of achieving real-time high-definition video encoding performance running on GPUs. For comparison reasons, multi-core parallel and distributed versions of these algorithms were developed using OpenMP and MPI (Message Passing Interface) libraries, respectively. The CUDA-based solutions achieve the highest speed-up in comparison with OpenMP and MPI versions for both algorithms and, when compared to the state-of-the-art, our FS and DS solutions reach up to 18x and 11x speed-up, respectively. Eduarda Monteiro, Marilena Maule, Felipe Sampaio, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi |
ICIP | 5 |
| 2012 | A Model Predictive Controller for Frame-Level Rate Control in Multiview Video CodingabstractIn this work, we present a novel frame-level Rate Control algorithm for Multiview Video Coding encoder that adopts the Model Predictive Control technique in order to provide low bitrate fluctuation and high video quality. Our Model Predictive Rate Control (MPRC) predicts the bitrate for a frame by employing (i) inter-view inter-GOP (Group of Pictures) phase-based bitrate prediction, and (ii) temporal (intra-GOP) target bitrate linear weighting. Moreover, the MPRC also defines an optimal control action through frame-level QP value selection. Experimental results demonstrate that our MPRC bitrate prediction incurs a Mean Bit Estimation Error (MBEE) of 1.13% compared to 2.46% provided by single view-based Rate Control and 1.61% provided by the state-of-the-art MVC Rate Control. Our solution also provides on average 0.876dB BD-PSNR increase and 28.92% BD-Bitrate reduction while providing smoother quality and bitrate variations when compared to state-of-the-art. Bruno Boessio Vizzotto, Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICME | 2 |
| 2012 | A complexity reduction scheme with adaptive search direction and mode elimination for multiview video codingabstractA novel complexity reduction scheme for Multiview Video Coding (MVC) is presented that adaptively eliminates the less-probable Motion or Disparity Estimation (ME, DE) search directions and less-probable block coding modes. Based on their texture difference w.r.t. the current Macroblock, matching neighbors in the 3D-neighborhood (spatial, temporal, view domains) are identified. Our scheme employs a multi-level decision process to predict the more-probable ME/DE search direction based on the texture and (motion/disparity) activity classification of the matching neighbors. For a predicted search direction, more-probable block coding modes are predicted depending upon the texture classification and RD-Cost of the current Macroblock. Quantization Parameter based thresholds are formulated using an offline statistical analysis of texture, motion/disparity, and RD-Cost properties. Our scheme achieves a complexity reduction of up to 81% and 40% compared to the exhaustive Rate-Distortion-Optimized Mode Decision and state-of-the-art, respectively, at the cost of an average BD-PSNR loss of 0.03 dB. Muhammad Shafique 0001, Bruno Zatt, Jörg Henkel |
PCS | 2 |
| 2011 | Run-time adaptive energy-aware motion and disparity estimation in multiview video codingabstractThis paper presents a novel run-time adaptive energy-aware Motion and Disparity Estimation (ME, DE) architecture for Multiview Video Coding (MVC). It incorporates efficient memory access and data prefetching techniques for jointly reducing the on/off-chip memory energy consumption. A dynamically expanding search window is constructed at run time to reduce the off-chip memory accesses. Considering the multi-stage processing nature of advanced fast ME/DE schemes, a reduced-sized multi-bank on-chip memory is employed which can be power-gated depending upon the video properties. As a result, when tested for various video sequence, our approach provides a dynamic energy reduction of 82--96% for the off-chip memory and a leakage energy reduction of 57--75% for the on-chip memory compared to the Level-C and Level-C+ [7] prefetching techniques (which are the prominent data reuse and prefetching techniques in ME for video coding). The proposed ME/DE architecture is synthesized using a 65nm IBM low power technology. Compared to state-of-the-art MVC ME/DE hardware [14], our architecture provides 66% and 72% reduction in the area and power consumption, respectively. Moreover, our scheme achieves 30fps ME/DE 4-view HD1080p encoding with a power consumption of 74mW. Bruno Zatt, Muhammad Shafique 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel |
DAC | 1 |
| 2011 | Multi-level pipelined parallel hardware architecture for high throughput motion and disparity estimation in Multiview Video CodingabstractThis paper presents a novel motion and disparity estimation (ME, DE) scheme in Multiview Video Coding (MVC) that addresses the high throughput challenge jointly at the algorithm and hardware levels. Our scheme is composed of a fast ME/DE algorithm and a multi-level pipelined parallel hardware architecture. The proposed fast ME/DE algorithm exploits the correlation available in the 3D-neighborhood (spatial, temporal, and view). It eliminates the search step for different frames by prioritizing and evaluating the neighborhood predictors. It thereby reduces the coding computations by up to 83% with 0.1 dB quality loss. The proposed hardware architecture further improves the throughput by using parallel ME/DE modules with a shared array of SAD (Sum of Absolute Differences) accelerators and by exploiting the four levels of parallelism inherent to the MVC prediction structure (view, frame, reference frame, and macroblock levels). A multi-level pipeline schedule is introduced to reduce the pipeline stalls. The proposed architecture is implemented for a Xilinx Virtex-6 FPGA and as an ASIC with an IBM 65nm low power technology. It is compared to state-of-the-art at both algorithm and hardware levels. Our scheme achieves a real-time (30fps) ME/DE in 4-view High Definition (HD1080p) encoding with a low power consumption of 81 mW. Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
DATE | 1 |
| 2011 | A low-power memory architecture with application-aware power management for motion & disparity estimation in Multiview Video CodingabstractA low-power architecture for an on-chip multi-banked video memory for motion and disparity estimation in Multiview Video Coding is proposed. The memory organization (size, banks, sectors, etc.) is driven by an extensive analysis of memory-usage behavior for various 3D-video sequences. Considering a multiple-sleep state model, an application-aware power management scheme is employed to reduce the leakage energy of the on-chip memory. The knowledge of motion and disparity estimation algorithm in conjunction with video properties are considered to predict the memory requirements of each Macroblock. A cost function is evaluated to determine an appropriate sleep mode for the idle memory sectors, while considering the wakeup overhead (latency and energy). The complete motion and disparity estimation architecture is implemented in a 65nm low power IBM technology. The experiments (for various test video sequences) demonstrate that our architecture provides up to 80% leakage energy reduction compared to state-of-the-art. Our scheme processes motion and disparity estimation of four HD1080p views encoding at 30fps with a power consumption of 57mW. Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICCAD | 1 |
| 2011 | A multi-level dynamic complexity reduction scheme for multiview video codingabstractIn this paper, we propose a novel scheme for dynamically reducing the computational complexity of MVC. Our scheme exploits the coding mode correlation available in the 3D-neighborhood (i.e., spatial, temporal, and view) along with the rate-distortion proper- ties of the neighboring Macroblocks. Our scheme incorporates a multi-level mode decision process based on a mode-ranking mechanism that categorizes more-probable and less-probable coding modes. In order to react to the changing bitrates, our scheme deploys Quantization Parameter based threshold equations which are formulated using an offline statistical analysis. Compared to the exhaustive Rate-Distortion-Optimized Mode Decision (RDO-MD), our scheme achieves a complexity reduction of up to 80% (68% on average) with an average PSNR loss of 0.075 dB. Compared to state-of-the-art fast RDO-MD, our scheme achieves a complexity reduction of up to 34% with an average PSNR gain of 0.007 dB. Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
ICIP | 1 |
| 2011 | A high throughput H.264/AVC intra-frame encoding loop architecture for HD1080pabstractIn this work we present a high throughput hardware architecture for the H.264/AVC intra-frame encoder exploiting the parallelism of intra prediction, forward and inverse transforms and quantization. Since there is a strong data dependency between the intra prediction and the image reconstruction loop, the latency of this path is a key design issue in order to provide high performance coding. Considering that 77% of the total intra-encoding computation is spent in these modules, our architecture handles a 4-pixel wide intra prediction module and a 16-pixel wide reconstruction loop. Compared to the state-of-the-art our approach reduces by 47% the number of cycles to process a macroblock. Running at 150 MHz our architecture guarantees encoding of 61 HD1080p frames per second. The developed architecture requires 73.4 MHz to real-time encode HD1080p, which is a 46% reduction of the frequency requirement compared to the state-of-the-art. Cláudio Machado Diniz, Bruno Zatt, Cristiano Thiele, Altamiro Amadeu Susin, Sergio Bampi, Felipe Sampaio, Daniel Palomino 0001, Luciano Volcan Agostini |
ISCAS | 2 |
| 2011 | Applying CUDA Architecture to Accelerate Full Search Block Matching Algorithm for High Performance Motion Estimation in Video EncodingabstractThis work presents a parallel GPU-based solution for the Motion Estimation (ME) process in a video encoding system. We propose a way to partition the steps of Full Search block matching algorithm in the CUDA architecture. A comparison among the performance achieved by this solution with a theoretical model and two other implementations (sequential and parallel using OpenMP library) is made as well. We obtained a O(n^2/log^2n) speed-up which fits the proposed theoretical model considering different search areas. It represents up to 600x gain compared to the serial implementation, and 66x compared to the parallel OpenMP implementation. Eduarda Monteiro, Bruno Boessio Vizzotto, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi |
SBAC-PAD | 4 |
| 2010 | Gop structure adaptive to the video content for efficient H.264/AVC encodingabstractThis paper presents a new method for high efficiency video coding using an adaptive GOP structure based on video content for the H.264/AVC standard. The available H.264/AVC encoders typically use static GOP sizes that define how the frames I (Intra), P (Predictive) and B (Bi-predictive) are positioned during de coding process. However, by analyzing the video content it is possible to identify the optimum position for each type of frame inside the GOP. The proposed method analyses the video content and finds the best position for inserting I frames in the video sequence. Thus the GOP structure can assume different sizes, depending on the video content. The results for test sequences and real videos show that the proposed method can significantly reduce the required bit rate, comparing to the static GOP sizes, with reduced PSNR losses. The proposed adaptive GOP presents a gain, in terms of bit rate reduction for real movies, of 8.6%, 15%, 24.7% and 40.8% in comparison with static GOP sizes 32, 16, 8 and 4, respectively. Bruno Zatt, Marcelo Schiavon Porto, Jacob Scharcanski, Sergio Bampi |
ICIP | 1 |
| 2010 | Power-aware complexity-scalable multiview video coding for mobile devicesabstractWe propose a novel power-aware scheme for complexity-scalable multiview video coding on mobile devices. Our scheme exploits the asymmetric view quality which is based on the binocular suppression theory. Our scheme employs different quality-complexity classes (QCCs) and adapts at run time depending upon the current battery state. It thereby enables a run-time tradeoff between complexity and video quality. The experimental results show that our scheme is superior to state-of-the-art and it provides an up to 87% complexity reduction while keeping the PSNR close to the exhaustive mode decision. We have demonstrated the power-aware adaptivity between different QCCs using a laptop with battery charging and discharging scenarios. Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel |
PCS | 2 |
| 2010 | An adaptive early skip mode decision scheme for multiview video codingabstractIn this work a novel scheme is proposed for adaptive early SKIP mode decision in the multiview video coding based on mode correlation in the 3D-neighborhood, variance, and ratedistortion properties. Our scheme employs an adaptive thresholding mechanism in order to react to the changing values of Quantization Parameter (QP). Experimental results demonstrate that our scheme provides a consistent time saving over a wide range of QP values. Compared to the exhaustive mode decision, our scheme provides a significant reduction in the encoding complexity (up to 77%) at the cost of a small PSNR loss (0.172 dB in average). Compared to state-of-the-art, our scheme provides an average 2x higher complexity reduction with a relatively higher PSNR value (avg. 0.2 dB). Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel |
PCS | 1 |
| 2010 | Timing and interface communication analysis of H.264/AVC encoder using SystemC modelabstractThis work presents a detailed timing and communication analysis for an H.264/AVC video encoder architecture using a SystemC model. The model was described using different abstraction levels in order to evaluate specific characteristics of each component module. The target encoder is defined to be able for H.264/AVC real-time encoding for 1080p video sequences at 30 fps and was modeled as a two-stage macro-pipeline system composed by eight component modules: Macroblock buffer, Intra- and Inter-Frame Predictors, Mode Decision, Forward and Inverse Transforms and Quantization, Reference Memory Write and Entropy Encoder (CAVLC). The bandwidth of each internal connection and of external memory interface was evaluated. The timing behavior and the data dependencies were characterized and summarized in a timing diagram in order to define design constraints and provide an accurate system specification when compared to a H.264/AVC encoder in the literature. Bruno Zatt, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
VLSI-SoC | 1 |
| 2009 | A real time H.264/AVC intra frame prediction hardware architecture for HDTV 1080P videoabstractThis work presents an intra frame prediction hardware architecture for H.264/AVC baseline/main profile encoder which performs real time processing of HDTV 1080p videos. It is achieved by exploring the parallelism of intra prediction and by reducing the latency for Intra 4times4 processing, which is the intra encoding bottleneck. Synthesis results on Xilinx Virtex-II Pro FPGA and TSMC 0.18 mum standard-cells indicate that this architecture is able to real time encode HDTV 1080p video operating at 110 MHz. Our architecture can encode HD1080p, 720p and SD video in real time at a frequency 25% lower when compared to similar works. Cláudio Machado Diniz, Bruno Zatt, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi |
ICME | 2 |
| 2008 | HP422-MoCHA: A H.264/AVC High Profile motion compensation architecture for HDTVabstractThis work presents the HP422-MoCHA, the first published hardware architecture that implements a full compliant H.264/AVC motion compensator for high profile 4:2:2. The hardware is composed by three main modules: Motion Vector Predictor, Memory Access and Sample Interpolator. The designed architecture was described in VHDL and mapped to a Xilinx Virtex-II PRO FPGA. This architecture reaches the throughput to decode HDTV Level 4 (1080p) @ 30fps. Bruno Zatt, Altamiro Amadeu Susin, Sergio Bampi, Luciano Volcan Agostini |
ISCAS | 1 |
| 2007 | MoCHA: a Bi-Predictive Motion Compensation Hardware for H.264/AVC Decoder Targeting HDTVabstractThis paper presents the MoCHA (motion compensation hardware architecture) design. MoCHA is an architectural design for bi-predictive motion compensation of the H.264/AVC decoder. The designed architecture features a memory hierarchy to reduce the memory bandwidth and the number of memory access cycles. The architecture uses a single datapath to process bi-predictive reference areas and it processes luma and chroma samples in parallel. The design was mapped to a Xilinx Virtex II Pro FPGA and it is able to run at 100MHz. The throughput is enough to support more than 30 bi-predictive HDTV frames per second. Arnaldo Azevedo, Bruno Zatt, Luciano Volcan Agostini, Sergio Bampi |
ISCAS | 2 |
| 2007 | Motion Compensation Hardware Accelerator Architecture for H.264/AVC
Bruno Zatt, Valter Ferreira, Luciano Volcan Agostini, Flávio Rech Wagner, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 1 |
| 2006 | Motion Compensation Decoder Architecture for H.264/AVC Main Profile Targeting HDTVabstractThis work presents the design, the validation and the prototyping of a motion compensation architecture for a H.264/AVC video decoder. The designed architecture supports the main profile level 4.0 and it targets high resolution applications, like HDTV. This design considers the sample processing of the motion compensation block, which includes quarter-pel interpolation, weighted prediction, average to bi-predictive processing and clipping. The architecture processes luma and chroma samples in parallel, with independent luma and chroma datapaths. The design uses a single interpolator to process bi-predictive macroblocks. The design was synthesized to FPGA and standard cell technologies. The synthesis results had indicated that this architecture reaches 100 MHz in both technologies, allowing real time to decode HDTV videos with 1920times1080 pixels. The prototype was targeted to a Xilinx Virtex-II PRO FPGA Arnaldo Azevedo, Bruno Zatt, Luciano Volcan Agostini, Sergio Bampi |
VLSI-SoC | 2 |