EDBT 2026 Demo / reviewers in the wild / expert
Guilherme Corrêa 0001
dblp:14/8181 · also Guilherme Ribeiro Corrêa
· DBLP profile ↗
46ranked-venue papers
12as first author
17since 2021 · last 2026
0000-0002-2739-6194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-author · 7 since 2021Systems, architecture and hardware · 17 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantitative assessment of inter-frame prediction in versatile video coding standardabstractAbstract Versatile Video Coding (VVC) is the latest video coding standard established by ISO and ITU-T, boasting a doubled coding efficiency compared to previous standards. However, this substantial enhancement in coding efficiency comes at the cost of an increase in computational demands, posing challenges for practical implementations. This paper conducts a thorough and quantitative assessment of VVC inter-frame prediction, the most computationally intensive tool in the VVC encoder. A comprehensive set of experiments on inter-frame prediction is presented, with a detailed evaluation of the primary bottlenecks of this tool. To the best of the authors’ knowledge, this work is the first in the literature to offer this in-depth assessment of VVC inter-frame prediction. Marta Loose, Ramiro Viana, Gustavo Sanchez, Guilherme Corrêa 0001, Luciano Volcan Agostini |
Multim. Tools Appl. | 4 |
| 2025 | Machine Learning-Driven Multiple Transform Selection for Low-Complexity VVC EncodingabstractThe H.266/VVC video coding standard achieves significant compression rates, but its deployment faces important challenges due to high computational demands. This is particularly evident in the Multiple Transform Selection (MTS) tool, which applies various combinations of transforms to enhance the energy concentration of the prediction residue. This work first presents an analysis of the MTS tool and shows that the distribution of transform mode decisions varies significantly. Next, a machine learning approach based on decision trees is employed to avoid the need for exhaustively testing all transform mode combinations in MTS. The predictive models are trained on features extracted directly from the encoding process, thus avoiding additional processing overhead. The proposed strategy reduces the overall encoding time by 30.32% in the All-Intra configuration, with only a minor increase of 1.17% in BD-rate, achieving a better balance between efficiency and complexity. Caroline Camargo, Bianca Silveira, Bruno Zatt, Guilherme Corrêa 0001 |
ISCAS | 4 |
| 2025 | Machine Learning-Based Block Partitioning for V-PCC Encoding AccelerationabstractPoint clouds capture spatial and attribute information about objects or environments. It has been widely used in applications like autonomous driving, augmented and virtual reality, and 3D scanning. However, the large volume of data involved in point cloud processing poses challenges in terms of storage, transmission, and real-time processing. The Video-based Point Cloud Compression (V-PCC) standard addresses these challenges by employing 2D video compression techniques to encode dynamic 3D point clouds. Despite its effectiveness, V-PCC demands a high computational cost, particularly in encoding geometry and attribute video sub-streams. This paper presents a machine learning-based approach to accelerate the V-PCC encoding, focusing on the block partitioning process. The proposed method utilizes decision tree models to predict the partitioning of Coding Tree Units (CTUs) and Prediction Units (PUs) during the encoding of geometry and attribute sub-streams. The proposed method significantly accelerates the V-PCC encoder by reducing the overall coding time by up to 60%, with an average decrease in the coding efficiency of 1.31% and 1.75% for the attribute (luma) and geometry (D2) sub-stream in Random Access configuration and negligible impact for All-Intra configuration. Gustavo Rehbein, Cristiano Santos, Guilherme Corrêa 0001, Marcelo Schiavon Porto |
VCIP | 3 |
| 2024 | A systematic literature review on video transcoding acceleration: challenges, solutions, and trends
Bruno Zatt, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
Multim. Tools Appl. | 4 |
| 2023 | H.264-to-AV1 Video Transcoding Acceleration Based on Lightweight Machine LearningabstractVideo streaming platforms have been using the H.264/AVC standard for a long time, even though it was released almost 20 years ago and much more efficient codecs are currently available. The AOMedia Video 1 (AV1) format is an alternative with significant coding efficiency gains in comparison to H.264/AVC, besides being a royalty-free format. However, migrating legacy content from older to newer formats is a costly task, which requires long processing times. This work presents a solution for accelerating the H.264-to-AV1 transcoder based on machine learning. Sixteen decision tree models trained with data gathered during the H.264/AVC decoding and the AV1 encoding processes are proposed and implemented in the libaom reference software, leading to a complexity reduction of 18.96% at the cost of coding efficiency losses of 2.85% on average. To the best of the authors' knowledge, this is the first H.264-to-AV1 transcoding acceleration solution published in the literature. Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001 |
ISCAS | 4 |
| 2023 | Fast Intra Mode Decision Using Machine Learning for the Versatile Video Coding StandardabstractThis paper presents a fast intra mode decision solution for the VVC standard using machine learning. The idea is to reorder the evaluation of modes performed by the Rate-Distortion Optimization (RDO) process according to the modes occurrence rate. Based on the new evaluation order, three Decision Tree models were trained to skip the modes less likely to be chosen. The results show that the proposed solution achieves time savings of up to 15.57% with coding efficiency degradation of only 0.41% on average. When compared with related works, the proposed solution shows competitive results. Adson Duarte, Bruno Zatt, Guilherme Corrêa 0001, Daniel Palomino 0001 |
ISCAS | 3 |
| 2023 | Multiversion Low-Power Hardware Accelerator for the AV1 Interpolation FiltersabstractOne of the new tools included in the AV1 video codec is the adaptive filtering scheme used in the sample interpolation process. This scheme includes three different filter families called Regular, Sharp and Smooth, offering high flexibility for motion estimation (ME) and motion compensation (MC). However, the high number of interpolation filters also leads to greater complexity and energy consumption, since the generation of samples at sub-pixel position is a costly process. This paper proposes a low-power and high-throughput hardware accelerator focused on the AV1 interpolation filters called Multiversion Interpolation Processor (MVIP). The accelerator includes the three AV1 interpolation filter families, with versions that employ operand isolation for power reduction in unused filters. The accelerator also includes a precise MVIP assuming the MC scenario, besides two approximate versions to reduce the cost on the ME scenario. The proposed design is able to process 8K video at 50fps in MC and 2,656.14 Msamples/sec in ME, with a power dissipation of 41.30mW. Daiane Freitas, Mateus Grellert, Cláudio Machado Diniz, Guilherme Corrêa 0001 |
ISCAS | 4 |
| 2023 | Learning-Based Fast VVC Affine Motion EstimationabstractThis paper presents a fast Affine Motion Estimation (AME) of Versatile Video Coding (VVC) Standard, based on Machine Learning and using Random Forest (RF) classification method. This encoding approach develops an RF model for each block size. The models were trained with information extracted during the VVC encoding process of the current, parent, and neighboring Coding Units (CU). Each model is applied to predict whether the Affine Motion Estimation (AME) will be skipped or not for that CU size. The proposed solution achieves a reduction of 20% on average in AME encoding time, with an insignificant impact of 0.07% on BD-BR. Fernando Sagrilo, Marta Loose, Ramiro Viana, Gustavo Sanchez, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 5 |
| 2023 | Learning-based bypass zone search algorithm for fast motion estimation
Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
Multim. Tools Appl. | 2 |
| 2022 | Low-Complexity Multi-Type Tree Partitioning for Versatile Video Coding Based on Machine LearningabstractThe Versatile Video Coding (VVC) standard introduces new types of frame partitioning structures, such as QuadTrees (QT) and Multi-Type Trees (MTT). To achieve the best compression efficiency, for each block of pixels the encoder performs a recursive search over the partitioning possibilities, which also impacts significantly on the encoding complexity and processing time. This work proposes a machine learning-based solution for quick block partitioning decisions. A set of fourteen Random Forests were trained using data gathered during the encoding process and the models were employed to decide whether vertical and horizontal partitions are required for each candidate block. The proposed solution leads to an average encoding time reduction of 34.76% at the cost of a compression efficiency loss of 1.03%. Matheus Lindino, Bruno Zatt, Mateus Grellert, Guilherme Corrêa 0001 |
ICIP | 4 |
| 2022 | Mode-Adaptive Subsampling of SAD/SSE Operations for Intra Prediction Cost ReductionabstractModern video encoders, such as the recently proposed AV1 and VVC, offer significant encoding gains at the cost of a corresponding increase of the computational effort. This is the case of the adopted intra prediction techniques, comprehending an increased number of prediction modes and range. To mitigate this computational cost, the presented work proposes a new mode-adaptive algorithm that significantly reduces the number of SAD/SSE operations during intra prediction, by generating an optimized subsampling pattern adaptive to each prediction mode. The method can be applied to any video codec and, when applied to AV1, it led to an encoding time reduction and BD-BR impact of 15.36% and 0.6%, respectively, or 7.97% and –0.02%, depending on the selected subsampling parameters. When implemented in hardware, the proposed technique provides an effective reduction as high as 75% of both the area and power on the modified distortion calculation module. Marcel Moscarelli Corrêa, Nuno Roma, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 4 |
| 2022 | Fast Affine Motion Estimation for VVC using Machine-Learning-Based Early Search TerminationabstractThe Affine Motion Estimation (AME) was introduced in the Versatile Video Coding (VVC) standard to allow for the detection of non-translational transformations during inter-frame prediction. Although providing important coding efficiency gains, this new tool represents 43% of the motion estimation (ME) complexity. However, an analysis over the AME step shows that the Affine motion vectors are often generated without resulting in the best ME prediction. This paper proposes a AME early search termination based on supervised machine learning. Six Random Forest models were trained with features obtained during the encoding process to accurately predict whether the AME step should be executed, partially executed or skipped, avoiding unnecessary calculations. As result, the proposed solution achieves an average time saving of 46.94% in the AME step with a coding efficiency loss of only 0.18%. Adson Duarte, Luciano Volcan Agostini, Bruno Zatt, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Daniel Palomino 0001 |
ISCAS | 5 |
| 2022 | Multiple Transform Selection Hardware Design for 4K@60fps Real-Time Versatile Video CodingabstractOne of the main innovations introduced in the Versatile Video Coding (VVC) standard is the possibility to employ and combine different types of transforms for residual coding through a tool named as Multiple Transform Selection (MTS). This improved flexibility leads to a high computational cost, requiring efficient hardware designs for the transform module to achieve real-time processing. This work presents a dedicated hardware design for the MTS module. The architecture is capable of processing several block sizes and it implements all the allowed transform combinations of the MTS tool. The obtained results show that the architecture is capable of processing up to 4K@60fps videos in real time with a frequency of 279 MHz and a power dissipation of 583 mW. Also, when compared with related works, the proposed solution shows competitive results. Bianca Silveira, Luiz Neto, Daniel Palomino 0001, Cláudio Machado Diniz, Guilherme Corrêa 0001 |
ISCAS | 5 |
| 2022 | Multi-Objective optimized Complexity Control for the AV1 Video EncoderabstractAOMedia Video 1 (AV1) is an open-source video encoding format launched in 2018 by Alliance for Open Media. To achieve high coding efficiency, AV1 brings significant innovations in comparison to its predecessor, the VP9 format. However, this came at the cost of a higher encoding complexity due to the newly introduced tools and block partitioning structures. To allow for wide deployment of AV1 in different multimedia devices, adjusting its encoding process according to the available computational resources is strongly desirable. Thus, this work presents an adaptive complexity controller for the AV1 encoder based on multi-objective optimization. Experimental results show that the proposed controller reduces encoding complexity in a target range from 10% to 40%, with satisfactory precision varying between 0.04 and 4.10 percentage points, and a BD-Rate between 0.22% and 5.24%. Gustavo Rehbein, Isis Bender, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
PCS | 3 |
| 2021 | Block-Based Inter-Frame Prediction For Dynamic Point Cloud CompressionabstractIn recent years, 3D point clouds have gained popularity thanks to technological advances such as the increased computational power and the availability of low-cost devices for acquisition of 3D information, like RGBD sensors. However, raw point clouds demand a large amount of data for their representation, and compression is mandatory to allow efficient transmission and storage. Inter-frame prediction is a widely used approach to achieve high compression rates in 2D video encoders, but the current literature still lacks solutions that efficiently exploit temporal redundancy for point cloud encoding. In this work, we propose a novel inter-frame prediction for 3D point cloud compression, which explores temporal redundancies in the 3D space. Moreover, a mode decision algorithm is also proposed to dynamically choose the best encoding mode between inter and intra prediction. The proposed method yields a bitrate reduction of 15.6% and 3.5% for geometry and luma information respectively, with no significant impact in objective quality when compared to the MPEG 3DG solution, called G-PCC. Cristiano Santos, Mateus M. Gonçalves, Guilherme Corrêa 0001, Marcelo Schiavon Porto |
ICIP | 3 |
| 2021 | Low-Power and High-Throughput Approximated Architecture for AV1 FME InterpolationabstractModern video encoders like the AOM Video 1 (AV1) implement several complex tools to allow the required high level of compression efficiency. The Fractional Motion Estimation (FME) is one of these tools and in AV1 the FME defines 90 different filters. To handle such complexity, hardware acceleration using approximate computing has become an alternative to be explored. This paper presents an approximate solution for the AV1 FME interpolation filters based on the approximation of the original filter coefficients intending to generate more hardware friendly coefficients. The approximated version was designed in hardware and can achieve real-time interpolation for UHD 8K videos at 30 frames per second, when synthesized using 40nm TSMC standard-cells technology. The designed architecture dissipates 26.79mW which represents more than 80% power reduction when compared to the original precise solution. The approximation implied in a small average coding efficiency degradation of 0.54% in BD-BR. When comparing with related works, this architecture reaches an expressive power reduction (2.1 to 4.8 times) even supporting more complex tools. Robson Domanski, William Kolodziejski, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 3 |
| 2021 | Complexity and Coding Efficiency Assessment of the Versatile Video Coding StandardabstractThe Versatile Video Coding standard was finalized by the Joint Video Exploration Team in July 2020 and is currently considered the state-of-the-art video compression technology. VVC significantly improves coding efficiency compared to HEVC thanks to several new features and tools that incur a large increase in computational cost. This paper presents a complexity and coding efficiency assessment of VVC divided into three analyses, focusing on: (1) the impact of using SIMD optimizations in the VVC Test Model software, (2) the impact of limiting partitioning structures when encoding, and (3) the computational cost associated to each encoding tool in VVC. Experimental results show that SIMD optimizations accelerate the encoding time by 40%, on average, and that limiting the available partitioning structures can decrease encoding time between 40% and 73%. The software profiling revealed that inter-frame prediction is responsible for almost half of the total encoding time. Finally, the paper also presents an analytic discussion on tools and partitioning possibilities that are rarely chosen in the mode decision process despite their high impact in coding complexity. Ícaro Siqueira, Guilherme Corrêa 0001, Mateus Grellert |
ISCAS | 2 |
| 2020 | RDE-MOGA: Automatic Selection of Rate-Distortion-Energy Control Points for Video Encoders Using Muti-Objetive Genetic AlgorithmabstractControlling energy consumption of video encoders is a complex multi-objective optimization problem of great importance. In this work we propose the RDE-MOGA, an multi-objective genetic algorithm capable of finding energetically efficient configurations for the HEVC encoder and replacing the current sensitivity analysis methodologies in the development of energy controllers. The utilization of our algorithm improved its efficiency in 60% whereas increasing the range of achievable reductions of the controller in at least 50%. Furthermore, the algorithm proved capable of sustaining 30% energy reduction at a cost of 3.45 BD-BR loss. Italo Machado, Marilton S. de Aguiar, Marcelo Schiavon Porto, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ICASSP | 4 |
| 2020 | Spatially Adaptive Intra Mode Pre-Selection for ERP 360 Video CodingabstractIn this work, we propose a spatially adaptive HEVC intra mode pre-selection for equirectangular (ERP) 360 video coding. The proposed technique exploits the spatial characteristics of 360 video in the ERP projection to reduce the complexity of intra prediction mode selection. The number of intra modes evaluated in Rate-Distortion Optimization is reduced based on a score technique that is adaptive to the frame region being encoded. Results show that the proposed technique achieves a complexity reduction of 16.5% with low coding efficiency penalties. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
ICASSP | 4 |
| 2020 | ASIC Solution for the Directional Intra Prediction of the AV1 Encoder Targeting UHD 4K VideosabstractAOMedia Video 1 (AV1), developed by the Alliance for Open Media consortium and released in 2018, is an open-source and royalty-free video format. It was designed to deliver substantial compression gains over its predecessor VP9 whilst keeping hardware feasibility and a practical decoding complexity. When compared to state-of-the-art formats, AV1 has more complex encoder tools, including the intra prediction which is the focus of this work. This paper presents a highly parallelized ASIC solution for the directional intra prediction module. It supports all 56 directional modes defined in AV1 and is able to process all combinations of block partitions. When synthesized to the TSMC 40nm technology with a target frequency of 1,296MHz, the proposed design used an area of 455.8K gates and showed a power dissipation and energy consumption per predicted sample of 40.92mW and 0.055pJ/sample, respectively. The reached throughput supports the processing of 60 frames per second for UHD 4K videos (3840×2160 pixels). No other work was found in the literature with a hardware design supporting the AV1 intra prediction directional modes. Marcel Moscarelli Corrêa, Luiz Neto, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 4 |
| 2020 | Low-Power and Memory-Aware Approximate Hardware Architecture for Fractional Motion Estimation Interpolation on HEVCabstractNowadays, current video coding standards like the High Efficiency Video Coding (HEVC) implement several complex coding tools, like the Fractional Motion Estimation (FME). An alternative to improve performance and save power is allying the hardware acceleration with approximate computing solutions, focusing on such complex tools. In this work, we present a low-power and memory-aware hardware architecture for the HEVC FME interpolator, proposing the development of two novel hardware designs for the interpolation filters, called Approximate Unified FME Filters (AUFF). These solutions exploit the usage of approximate computing at both algorithmic and data levels, leading to a reduction in dissipated power and memory bandwidth. The proposed design is capable of real-time interpolation of UHD (Ultra High Definition) 4K and 8K videos when synthesized using a 40 nm standard-cell library, with a power dissipation ranging from 22.04 to 62.06 mW. Wagner Penny, Guilherme Corrêa 0001, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Gabriel L. Nazar, Bruno Zatt |
ISCAS | 2 |
| 2020 | 4D-DCT Hardware Architecture for JPEG Pleno Light Field CodingabstractThis paper presents a 4D-DCT hardware architecture for Light Field Coding according to the JPEG Pleno standard. It is composed of two instances of 2D-DCT engines and a novel 4D Transposition Memory organization. Experimentally-defined fixed-point representation and LSB pruning techniques are employed do reduce hardware area and power dissipation. The proposed architecture operates over 4D-hypercubes of up to 8x8x8x8 samples and reaches performance to process 30 Lytro-like light fields per second dissipating 145.32mW at 825.75MHz. This is the first known 4D-DCT hardware architecture for light field coding and demonstrates the feasibility of such solutions on real-world systems. Matheus Jahnke, Jones Goebel, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 4 |
| 2020 | Complexity and compression efficiency assessment of 3D-HEVC encoder
Mário Saldanha, Ruhan A. Conceição, Vladimir Afonso, Giovanni Avila, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001, Luciano Volcan Agostini |
Multim. Tools Appl. | 8 |
| 2019 | Online Machine Learning for Fast Coding Unit Decisions in HEVCabstractThe High Efficiency Video Coding standard introduced a flexible frame partitioning process that increased significantly compression rates in comparison to previous standards at the cost of a high computational cost. To accelerate frame partitioning decisions, this paper proposes a method that replaces the usual Rate-Distortion Optimization employed in Coding Unit size decision by a set of simpler decision tree models, which are built during encoding time by the C5 machine learning algorithm. The algorithm and the set of attributes employed in the model training process were chosen based on an extensive analysis that compared several options in terms of decision accuracy and training complexity. Experimental results show that the proposed method is capable of building accurate models for each video sequence, decreasing the HEVC encoding complexity in 34.4% on average with a compression efficiency loss of only 0.2% in comparison to the original HEVC reference encoder. Guilherme Corrêa 0001, Pargles Dall'Oglio, Daniel Palomino 0001, Luciano Volcan Agostini |
DCC | 1 |
| 2019 | Coding Tree Early Termination for Fast HEVC Transrating Based on Random ForestsabstractVideo transrating has become an essential task in streaming service providers that need to transmit and deliver different versions of the same content for a multitude of users that operate under different network conditions. As the transrating operation is comprised of a decoding and an encoding step in sequence, a huge computational cost is required in such large-scale services, especially when considering the use of complex state-of-the-art codecs, such as the High Efficiency Video Coding (HEVC). This work proposes an early-termination method for complexity reduction of the HEVC transrating based on Random Forests, which use features obtained from the HEVC decoding process to accelerate the coding tree decisions during the re-encoding process. Experimental results show that the proposed method achieves an average transrating time reduction of 47.09% at the cost of a negligible bitrate increase of 0.292%. Thiago Luiz Alves Bubolz, Mateus Grellert, Bruno Zatt, Guilherme Corrêa 0001 |
ICASSP | 4 |
| 2019 | Fast Hevc-to-Av1 Transcoding Based On Coding Unit Depth InheritanceabstractWith the advent of the recently launched AOMedia Video 1 (AV1) bitstream specification, there is currently a need for converting legacy content encoded with the state-of-the-art High Efficiency Video Coding (HEVC) standard to the new format. However, transcoding is a complex task composed of a decoding and an encoding process in sequence, which requires long processing times and high energy consumption. This paper proposes the first HEVC-to-AV1 transcoding solution, which is based on the high correlation between block size decisions in HEVC and AV1. The solution allows the AV1 encoder to inherit Coding Unit (CU) depth information from the HEVC bitstream to constrain the AV1 re-encoding process. Experimental results show an average transcoding time reduction of 35.41% at the cost of a compression efficiency loss of 4.54%. Bruno Zatt, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 4 |
| 2019 | Encoding Efficiency and Computational Cost Assessment of State-Of-The-Art Point Cloud CodecsabstractPoint clouds have recently emerged as a suitable solution to generate and display 3D digital models due to their capacity of representing high resolution images and videos through multiple viewpoints. However, as they are usually made up of thousands up to billions of points, advanced techniques of data compression are essential to store and transmit this type of data. This paper compares the two state-of-the-art solutions for point cloud compression, the Point Cloud Codec (PCC) and the Test Model Category 2 (TMC2), in terms of compression efficiency and encoding time. Experimental results show that the compression efficiency for geometry information is highly dependent upon the available bitrate for both TMC2 and PCC. However, for texture compression TMC2 almost always achieves the best results. The experiments have also shown that TMC2 presents a computational cost from 22.2 to 26 times larger the observed in PCC. Mateus M. Gonçalves, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 5 |
| 2019 | High Throughput Hardware Design for AV1 Paeth and Smooth Intra ModesabstractDeveloped by AOMedia industry consortium and released in June 2018, AV1 is an open-source and royalty-free video coding format. The main goal of AV1 is to deliver substantial compression gains over state-of-the-art codecs such as VP9 and HEVC, while keeping a practical decoding complexity, hardware feasibility and its open and free status. This paper presents a high throughput hardware architecture for four important AV1 intra prediction coding modes: Paeth, Smooth, Smooth Vertical and Smooth Horizontal. The proposed architecture was designed to support all 19 block sizes specified by AV1 and to process every single combination of these blocks according to the 10-way partition tree, with a throughput of UHD 4K (3840×2160 pixels) videos at up to 30 frames per second. When synthesized to the TSMC 40nm cell library targeting a frequency of 648MHz, the proposed design used 109.57K gates and showed a power dissipation and an energy efficiency of 16.1mW and 1.23pJ/sample respectively. No other works were found in the literature describing hardware designs for AV1 intra prediction. Marcel Moscarelli Corrêa, Bianca Waskow, Bruno Zatt, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 5 |
| 2019 | Energy-Aware Motion and Disparity Estimation System for 3D-HEVC With Run-Time Adaptive Memory HierarchyabstractThe popularization of multimedia services has pushed forward the development of 2D/3D video-capable embedded mobile devices. Such devices require efficient energy/memory-management strategies to deal with severe memory/processing requirements and limited energy supply. Therefore, we propose a motion and disparity estimation (ME and DE) system—the most memory/processing demanding encoding steps—for the 3D High Efficiency Video Coding (3D-HEVC) standard. It was designed for low energy consumption, featuring a run-time adaptive memory hierarchy. The processing unit employs flexible coding order and optimizations to reduce the computational effort by exploring the inter-channel and inter-view redundancies. The memory hierarchy features window-based prefetching, data reuse, subsampling, and dynamic voltage scaling controlled by our depth-based dynamic search window resizing algorithm. Memory results demonstrate an average on-chip energy reduction of 79% in comparison to the widely used Level-C solution for a 45-nm technology. The proposed energy-aware ME and DE system dissipates 7.55 W while processing three HD 1080p views (video + depth) at 30 frames per second and presents a mean energy consumption of 0.107 J per access unit. To the best of our knowledge, this is the first work that proposes a real-time ME/DE system for the 3D-HEVC standard with an adaptive memory hierarchy. Vladimir Afonso, Ruhan A. Conceição, Mário Saldanha, Luciano Almeida Braatz, Murilo R. Perleberg, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt, Altamiro Amadeu Susin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Octagonal-Axis Raster Pattern for Improved Test Zone Search Motion EstimationabstractTest Zone Search (TZS) is considered the current state-of-the-art fast Motion Estimation algorithm because it presents the best tradeoff between compression efficiency and complexity in comparison to the Full Search strategy. However, it is still one of the most computationally-demanding tools of current video coding standards, such as the High Efficiency Video Coding (HEVC). This paper presents an analysis on the search area opportunities and best match distributions in TZS, which led to the proposal of a novel search pattern in its most complex step, the Raster Search (RS). The new pattern, named Octagonal-Axis Raster Pattern (OARP), allowed an average complexity reduction of 61 % in TZS, with a negligible BD-rate increase of 0.0371 % in comparison to the original algorithm. Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001 |
ICASSP | 5 |
| 2018 | Learning-Based Complexity Reduction and Scaling for HEVC EncodersabstractThis article proposes a fast Coding Unit (CU) partition decision for use in HEVC encoders based on Decision Tree classifiers. The trees are employed in a modified low-complexity encoder that implements a fast CU partition decision algorithm. Using the proposed method, an average complexity reduction of 47.8% is achieved with a Bjontegaard Delta bitrate (BD-BR) loss of 0.24% in the Random Access coding configuration, and a 42.8% complexity reduction with a 0.19% BD-BR loss in the Low Delay B configuration. A decision threshold analysis is also presented to assess the rate-distortion-complexity trade-off of the proposed method at different complexity points, varying the complexity reduction from 28% (with a 0.04% loss in BD-BR) up to 60% (with a 3.6% BD-BR loss) using the Random Access configuration. A comparison with related works shows that the proposed method outperforms competing solutions in terms of both rate-distortion efficiency and complexity reduction. Mateus Grellert, Sergio Bampi, Guilherme Corrêa 0001, Bruno Zatt, Luís Alberto da Silva Cruz |
ICASSP | 3 |
| 2018 | Hardware-Friendly Unidirectional Disparity-Search Algorithm for 3D-HEVCabstractThis paper presents a novel hardware-friendly Unidirectional Disparity-Search (UDS) algorithm for the 3D-HEVC. This algorithm explores the typical camera arrangements used in 3D-HEVC. UDS was evaluated in two operation points reaching a computational effort reduction from 32.7% to 61.8%, with a BD-Rate increase from 0.3123% to 0.4803%, when compared to TZS. Estimated hardware results showed memory-size and leakage-energy reductions of 66.7%, a dynamic-energy reduction from 34.5% to 62.1%, and an energy-consumption reduction from 32.7% to 61.8% in SAD calculations, when compared to TZS. To the best of the authors' knowledge, this is the first proposed DE algorithm that explores the 3D-HEVC typical camera arrangement. Vladimir Afonso, Altamiro Amadeu Susin, Murilo R. Perleberg, Ruhan A. Conceição, Guilherme Corrêa 0001, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2018 | OTED: Encoding Optimization Technique Targeting Energy-Efficient HEVC DecodingabstractThis work exploits the encoding for decoding concept by proposing an encoding optimization technique targeting energy-efficient HEVC decoding (called OTED). OTED changes the Rate-Distortion Optimization (RDO) calculation at the encoder side by adding the decoding energy estimation as a new variable to be considered. Along with this estimation, we use the Running Average Power Limit (RAPL) energy measurement tool to present real energy results and prove the efficiency of OTED. Experimental results show that the proposed algorithm can reach an energy reduction of up to 17.7% at the HEVC decoder with a small cost in coding efficiency. Besides being compliant with the standard HEVC decoder, the algorithm presents high levels of energy reduction for different encoding configurations. Furthermore, when compared with state-of-the-art works OTED presents the best relation between energy reduction consumption at the encoder side and coding efficiency. Douglas Corrêa, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ISCAS | 2 |
| 2017 | Multiple early-termination scheme for TZ search algorithm based on data mining and decision treesabstractThe latest video compression standards, such as the H.264/AVC and the High Efficiency Video Coding (HEVC), provide fast Motion Estimation (ME) algorithms in their reference software aiming at complexity reduction. Test Zone Search (TZS) is the state-of-the-art fast ME algorithm, currently deployed in the reference HEVC encoder due to its great coding efficiency. However, ME is still one of the main sources of complexity in HEVC. This paper proposes an early-termination scheme for TZS, called e-TZS, based on an extensive data mining process on ME attributes. The data mining process allowed identifying the most relevant information during the encoding process to build a set of decision tree models that terminate TZS in different steps of its execution. The e-TZS scheme was implemented in the HEVC reference software and achieved average decision precision of 94.2%. Experimental results showed an average complexity reduction of 62.53% in TZS, with a negligible BD-rate increase of only 0.49%, in comparison to the original algorithm. Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
MMSP | 2 |
| 2016 | Complexity reduction for 3D-HEVC depth map coding based on early Skip and early DIS schemeabstractThis paper presents a novel early Skip/DIS mode decision for 3D-HEVC depth encoding which aims at reducing the complexity effort of this process. The proposed solution is based on an adaptive threshold model, which takes into consideration the occurrence rate of both Skip and DIS modes. Occurrence analysis showed that the lower is the Skip and DIS Rate-Distortion cost, the higher is the probability of these modes being chosen. Furthermore, software evaluations showed that the proposed early Skip/DIS scheme is capable of reducing the depth coder complexity in 24.4% for a target hit rate of 99%, and in 33.7% for a target hit rate of 95%, leading to a negligible coding efficiency penalty in both scenarios. Ruhan A. Conceição, Giovanni Avila, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 3 |
| 2016 | Fast H.264/AVC to HEVC transcoder based on data mining and decision treesabstractHigh Efficiency Video Coding (HEVC) is gradually replacing its predecessor, the H.264/AVC standard, as the state-of-the-art technology for video compression. However, H.264/AVC has dominated the market for over a decade, so that there is an enormous amount of legacy content that must be migrated. This paper proposes a fast transcoder based on an extensive data mining process on H.264/AVC decoding attributes. The data mining allowed identifying relevant information from the H.264/AVC decoding process, which was conveyed to the C4.5 machine learning algorithm to build a set of decision trees that simplify the complex Coding Unit (CU) size decision in HEVC. Experimental results have shown an average reduction of 44% in the transcoding time, with a small bit rate increase of 1.67%. These results outperform any previous works available in the literature. Guilherme Corrêa 0001, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
ISCAS | 1 |
| 2016 | Pareto-Based Method for High Efficiency Video Coding With Limited Encoding TimeabstractSeveral different methods have been investigated in recent years, aiming at computational complexity reduction and scaling of High Efficiency Video Coding (HEVC) software implementations. However, maintaining the encoding time per frame or group of pictures (GOPs) below an adjustable upper bound is still an open research issue. A solution for this problem is devised in this paper based on a set of Pareto-efficient encoding configurations, identified through rate-distortion-complexity analysis. The proposed method combines a medium-granularity encoding time control with a fine-granularity encoding time control to accurately limit the HEVC encoding time below a predefined target for each GOP. It is shown that the encoding time can be kept below a desired target for a wide range of encoding time reductions, e.g., up to 90% in comparison with the original encoder. The results also show that compression efficiency loss (Bjøntegaard delta-rate) varies from negligible (0.16%) to moderate (9.83%) in the extreme case of 90% computational complexity reduction. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Encoding time control system for HEVC based on Rate-Distortion-Complexity analysisabstractThe improved compression efficiency of High Efficiency Video Coding (HEVC) comes with large increases in computational complexity, which has been lately dealt by researchers with complexity reduction and scaling methods. However, encoding time control at frame or Group of Pictures (GOP) level is still an open issue that must be investigated. In this work, a Rate-Distortion-Complexity analysis is performed upon a set of configurations that have been created based on findings and techniques of previous works. The proposed control system uses the best 27 configurations to adjust the encoding time per GOP and yields encoding time reductions of up to 84.5% with an average difference between target and encoding times of 4.1%. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
ISCAS | 1 |
| 2015 | Fast HEVC Encoding Decisions Using Data MiningabstractThe High Efficiency Video Coding standard provides improved compression ratio in comparison with its predecessors at the cost of large increases in the encoding computational complexity. An important share of this increase is due to the new flexible partitioning structures, namely the coding trees, the prediction units, and the residual quadtrees, with the best configurations decided through an exhaustive rate-distortion optimization (RDO) process. In this paper, we propose a set of procedures for deciding whether the partition structure optimization algorithm should be terminated early or run to the end of an exhaustive search for the best configuration. The proposed schemes are based on decision trees obtained through data mining techniques. By extracting intermediate data, such as encoding variables from a training set of video sequences, three sets of decision trees are built and implemented to avoid running the RDO algorithm to its full extent. When separately implemented, these schemes achieve average computational complexity reductions (CCRs) of up to 50% at a negligible cost of 0.56% in terms of Bjontegaard Delta (BD) rate increase. When the schemes are jointly implemented, an average CCR of up to 65% is achieved, with a small BD-rate increase of 1.36%. Extensive experiments and comparisons with similar works demonstrate that the proposed early termination schemes achieve the best rate-distortion-complexity tradeoffs among all the compared works. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Four-step algorithm for early termination in HEVC inter-frame prediction based on decision treesabstractThe flexible encoding structures of High Efficiency Video Coding (HEVC) are the main responsible for the improvements of the standard in terms of compression efficiency in comparison to its predecessors. However, the flexibility provided by these structures is accompanied by high levels of computational complexity, since more options are considered in a Rate-Distortion (R-D) optimization scheme. In this paper, we propose a four-step early-termination method, which decides whether the inter mode decision should be halted without testing all possibilities. The method employs a set of decision trees, which are trained offline once, using information from unconstrained HEVC encoding runs. The resulting trees present a mode decision accuracy ranging from 97.6% to 99.4% with a negligible computational overhead. The method is capable of achieving an average computational complexity decrease of 49% at the cost of a very small Bjontegaard Delta (BD)-rate increase (0.58%). Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
VCIP | 1 |
| 2013 | Coding Tree Depth Estimation for Complexity Reduction of HEVCabstractThe emerging HEVC standard introduces a number of tools which increase compression efficiency in comparison to its predecessors at the cost of greater computational complexity. This paper proposes a complexity control method for HEVC encoders based on dynamic adjustment of the newly proposed coding tree structures. The method improves a previous solution by adopting a strategy that takes into consideration both spatial and temporal correlation in order to decide the maximum coding tree depth allowed for each coding tree block. Complexity control capability is increased in comparison to a previous work, while compression losses are decreased by 70%. Experimental results show that the encoder computational complexity can be downscaled to 60% with an average bit rate increase around 1.3% and a PSNR decrease under 0.07 dB. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
DCC | 1 |
| 2013 | Constrained encoding structures for computational complexity scalability in HEVCabstractThe High Efficiency Video Coding standard shows improved compression efficiency in comparison to previous standards at the cost of higher computational complexity. In this paper, a complexity scalability method for HEVC encoders based on the dynamic adjustment of the number of constrained coding treeblocks is proposed. The method limits the Prediction Unit shapes and the maximum tree depth used in each Coding Treeblock in order to decrease the number of evaluations performed in the Rate-Distortion Optimization process. The encoder is capable of trading off computational complexity and compression efficiency while still maintaining the encoding time per Group of Pictures (GOP) under a pre-defined target. The encoding complexity can be decreased in up to 60% when compared to the original encoder at the cost of small Bjontegaard Delta (BD)-rate increases. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
PCS | 1 |
| 2012 | Motion compensated tree depth limitation for complexity control of HEVC encodingabstractThe recently introduced quadtree coding structures used in HEVC increase compression efficiency in comparison to previous standards at the cost of higher computational complexity levels. This paper proposes an evolution of a complexity control method for HEVC encoders based on the dynamic adjustment of these structures' maximum depth. The new method improves the previous solution by adopting a new control strategy and compensating the motion effect on a maximum tree depth map which is central to the complexity control strategy. The proposed method is capable of performing a more accurate complexity control than our previous strategy while still reducing compression efficiency losses in terms of image quality and bit rate. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
ICIP | 1 |
| 2012 | Adaptive coding tree for complexity control of high efficiency video encodersabstractThe emerging HEVC standard introduces several techniques which increase compression efficiency in comparison to its predecessors. However, such advances are accompanied by increases in computational complexity, limiting the encoder use in computational or power-constrained devices. This paper proposes a novel complexity control method for the future HEVC encoders based on a dynamic adjustment of the newly proposed coding tree structures. The relationship between coding tree depths and the encoding complexity is explored to selectively constrain encoding possibilities in order to not exceed a predefined complexity target. Experimental results show that the encoder computational complexity can be downscaled to 60% with a bit rate increase under 3.5% and a PSNR decrease under 0.1 dB. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
PCS | 1 |
| 2012 | Performance and Computational Complexity Assessment of High-Efficiency Video EncodersabstractThis paper presents a performance evaluation study of coding efficiency versus computational complexity for the forthcoming High Efficiency Video Coding (HEVC) standard. A thorough experimental investigation was carried out to identify the tools that most affect the encoding efficiency and computational complexity of the HEVC encoder. A set of 16 different encoding configurations was created to investigate the impact of each tool, varying the encoding parameter set and comparing the results with a baseline encoder. This paper shows that, even though the computational complexity increases monotonically from the baseline to the most complex configuration, the encoding efficiency saturates at some point. Moreover, the results of this paper provide relevant information for implementation of complexity-constrained encoders by taking into account the tradeoff between complexity and coding efficiency. It is shown that low-complexity encoding configurations, defined by careful selection of coding tools, achieve coding efficiency comparable to that of high-complexity configurations. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | SHBS: A heuristic for fast inter mode decision of H.264/AVC standard targeting VLSI designabstractIn the Rate-Distortion Optimization technique for H.264/AVC, the process of choosing the best mode is performed through exhaustive executions of the whole encoding process, which increases significantly the encoder complexity, sometimes even forbidding its use in real time video coding applications. In order to reduce the number of calculations necessary to determine the best inter-frame mode, this work proposes the SHBS (Stationarity, Heterogeneity and Border Strength) heuristic. The use of SHBS causes a reduction of 168 times in the encoding iterations, with a better PSNR, at the cost of a relatively small bit-rate increase. The SHBS heuristic was designed in hardware targeting FPGAs and this architecture achieved an operation frequency of 118 MHz, being able to process up to 438 HD 1080p frames per second. Guilherme Corrêa 0001, Daniel Palomino 0001, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
ICME | 1 |