VLDB 2026 Research / reviewers in the wild / expert
Luciano Volcan Agostini
dblp:80/5216 · also Luciano Agostini
· DBLP profile ↗
128ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-3421-5830ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 13 since 2021Systems, architecture and hardware · 53 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5Software engineering, systems software and programming languages · 3Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardware-Friendly Machine-Learning-Based Fast AV1 Overlapped Block Motion Compensation
William Kolodziejski, Leonardo Braga, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 4 |
| 2026 | Energy-Efficient Neural Video Coding via a High-Throughput Hardware Design for Pointwise Convolution and LUT-based WSiLU
Denis Maass, Vanessa Aldrighi, Ruhan A. Conceição, Wen-Hsiao Peng, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2026 | Quantitative assessment of inter-frame prediction in versatile video coding standardabstractAbstract Versatile Video Coding (VVC) is the latest video coding standard established by ISO and ITU-T, boasting a doubled coding efficiency compared to previous standards. However, this substantial enhancement in coding efficiency comes at the cost of an increase in computational demands, posing challenges for practical implementations. This paper conducts a thorough and quantitative assessment of VVC inter-frame prediction, the most computationally intensive tool in the VVC encoder. A comprehensive set of experiments on inter-frame prediction is presented, with a detailed evaluation of the primary bottlenecks of this tool. To the best of the authors’ knowledge, this work is the first in the literature to offer this in-depth assessment of VVC inter-frame prediction. Marta Loose, Ramiro Viana, Gustavo Sanchez, Guilherme Corrêa 0001, Luciano Volcan Agostini |
Multim. Tools Appl. | 5 |
| 2026 | DM-FIFS: A Dual-Model Machine-Learning Method for Fast Interpolation Filter Search in AV1 EncodingabstractAV1 is a video codec developed by leading technology companies to meet the increasing demands of modern video applications. Fractional Motion Estimation (FME), the focus of this work, is an important AV1 encoder tool. FME employs interpolation filters to generate sub-pixel predictions, thereby improving motion estimation accuracy. In AV1, FME uses sophisticated interpolation filters that can be combined in horizontal and vertical directions, with the optimal filter pair selected by the Interpolation Filter Search (IFS) process. The paper presents DM-FIFS, a dual-model, machine-learning-based approach designed to overcome prior limitations in filter prediction accuracy, which often led to suboptimal trade-offs between computational effort and coding efficiency. By splitting the decision space into two specialized models, DM-FIFS achieves more accurate filter predictions, thereby improving the balance between gains in computational effort and losses in coding efficiency compared to single-model approaches. The paper also presents a set of assessment and ablation experiments, a comprehensive discussion of key innovations in the AV1 encoder, and a detailed analysis of the interpolation filters used in AV1 FME. Experimental results show that DM-FIFS reduces IFS execution time by 51.40% with only a 0.11% increase in BD-BR, demonstrating a superior trade-off between computational effort and coding efficiency. To the best of our knowledge, DM-FIFS represents the most advanced machine-learning-based solution to reduce the computational complexity of AV1 IFS reported to date. William Kolodziejski, Leonardo Braga, Marcelo Rezende, Marcelo Schiavon Porto, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Cross-Platform Neural Video Coding: A Case StudyabstractIn this paper, we first show that current learning-based video codecs, specifically the SSF codec, are not suitable for real-world applications due to the mismatch between the encoder and decoder caused by floating-point round-off errors. To address this issue, we propose the static quantization of the hyper prior decoding path. The quantization parameters are determined through an exhaustive search of all possible combinations of observers and quantization schemes from PyTorch. For the SSF codec, when encoding and decoding on different machines, the proposed solution effectively mitigates the mismatch issue and enhances compression efficiency results by preventing severe image quality degradation. When encoding and decoding are performed on the same machine, it constrains the average BD-rate increase to 9.93% and 9.02% for UVG and HEVC-B sequences, respectively. Ruhan A. Conceição, Marcelo Schiavon Porto, Wen-Hsiao Peng, Luciano Volcan Agostini |
ISCAS | 4 |
| 2025 | A Power-Efficient Architecture for LES Solving and ∆MV Calculation of VVC Affine PredictionabstractThe increasing demand for video content has created a need for more efficient video compression techniques and the Versatile Video Coding (VVC) standard introduces several techniques to reach this goal. One key innovation in VVC is affine prediction, which gives more flexibility in inter-frame prediction by using multiple motion vectors for motion representation. However, its processing demands significant computational effort, boosting the use of dedicated hardware accelerators to enable real-time processing, mainly when focusing on mobile devices. This work introduces a pipelined and parallelized hardware design for the Linear Equation System (LES) solving and the ∆MV calculation, both essential steps in the search for affine motion vectors. The proposed architecture can generate new ∆MV values every three clock cycles and can process UHD 4K@60fps videos dissipating 47.32mW, with a negligible coding efficiency impact of 0.006%. These results outperform all related works in the literature, achieving twice the throughput, 84.2% lower power dissipation, and five times better coding efficiency. Denis Maass, Marcello M. Muñoz, Murilo R. Perleberg, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2025 | A Fast and Hardware-Friendly CTB-Based Approach for VVC Affine Motion Estimation
Denis Maass, Marcello M. Muñoz, Murilo R. Perleberg, Luciano Volcan Agostini, Marcelo Schiavon Porto |
VCIP | 4 |
| 2025 | FastGW: A Machine Learning-Based Early Skip for the AV1 Global Warped Motion CompensationabstractThe growing consumption of digital media, driven by technological advancements and exacerbated by the COVID-19 pandemic, has led to an increased demand for efficient video compression techniques. Among the various video encoders available, the AOMedia Video 1 (AV1) stands out since it was defined by the Alliance for Open Media (AOMedia), which is formed by big techs such as Google, Amazon, NetFlix, Meta, and Intel, among others. AV1 was launched in 2018 and it reaches high compression rates, especially for high-resolution videos. However, AV1 computational cost is significantly higher when compared to other current codecs. This paper is focused on one of the main novelties introduced by AV1: the Global Warped Motion Compensation (GWMC) tool. A computational effort reduction approach called Fast Global Warped (FastGW), using machine learning, is proposed to reduce the GWMC processing time. Then, a decision tree was trained to decide whether to skip the GWMC’s most computationally intensive step: the Refinement. This decision tree was implemented inside the AV1 encoder, resulting in an average time reduction of 23% at the GWMC, with a minimal impact on coding efficiency of 0.14% in BD-BR on average. To the best of the authors’ knowledge, this is the first work in the literature exploring machine learning to reduce the AV1 GWMC computational effort. William Kolodziejski, Robson Domanski, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | High-Throughput and Multiplierless Hardware Design for the AV1 Local Warped MC InterpolationabstractMost of the current video codecs support only translational motion models. However, real motion is often complex and cannot be precisely estimated using only translational models. To handle complex motions like panning, zooming, scaling, shearing and rotation, AOMedia AV1 encoder counts with two tools, called Global and Local Warped Motion Compensation (LWMC). This paper presents two dedicated hardware designs for the AV1 LWMC interpolation filters. The presented hardware can process up to UHD 8K videos at 60fps. The architecture was synthesized for 40nm TSMC standard cells, requiring 454.37K gates with a power dissipation of 189.35mW. To the best of the authors’ knowledge, this is the first work in the literature targeting a dedicated hardware design for LWMC AV1 tool. Robson Domanski, William Kolodziejski, Wagner Penny, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 6 |
| 2023 | High-Throughput Design for a Multi-Size DCT-II Targeting the AV1 EncoderabstractThis paper presents a dedicated multi-size hardware design for the Discrete Cosine Transform type II (DCT-II) of AV1 encoder. The DCT-II is one of four transform kernels supported by AV1; however, DCT-II is used in all configurations defined by AV1. Moreover, the 1D DCT-II can be applied for five different sizes ranging from 4-point up to 64-point. The 1D multi-size DCT-II was designed to process multiple transform sizes in parallel, always processing 64 samples in parallel for any size. The presented solution can process UHD 8K videos at 60 frames per second when running at 46.6 MHz, with a power dissipation of 44.48 mW and an area of 261.28 Kgates. To the best of authors' knowledge, this is the first work in the literature presenting a hardware design for the AV1 DCT-II transform. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 2 |
| 2023 | Learning-Based Fast VVC Affine Motion EstimationabstractThis paper presents a fast Affine Motion Estimation (AME) of Versatile Video Coding (VVC) Standard, based on Machine Learning and using Random Forest (RF) classification method. This encoding approach develops an RF model for each block size. The models were trained with information extracted during the VVC encoding process of the current, parent, and neighboring Coding Units (CU). Each model is applied to predict whether the Affine Motion Estimation (AME) will be skipped or not for that CU size. The proposed solution achieves a reduction of 20% on average in AME encoding time, with an insignificant impact of 0.07% on BD-BR. Fernando Sagrilo, Marta Loose, Ramiro Viana, Gustavo Sanchez, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 6 |
| 2023 | Learning-based bypass zone search algorithm for fast motion estimation
Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
Multim. Tools Appl. | 3 |
| 2023 | A High-Throughput Hardware Design for the AV1 Decoder IntrapredictionabstractThe Alliance for Open Media (AOMedia) (AV1) was released in 2018 as a royalty-free and open-source video codec. AV1 was developed by the AOMedia that is composed of many leading tech companies. AV1 has the goal to process ultrahigh definition (UHD) 8K (7680$\times4320$pixels) and 4K videos (3840$\times2160$pixels) and to achieve high coding efficiency, which leads to increased complexity when compared to other codecs in the market, such as VP9, HEVC, and H.264. This article presents the AV1 intraprediction decoder (AVID), a dedicated high-throughput hardware design for the AV1 decoder intraprediction supporting the AV1 68 prediction modes and 19 block sizes. The proposed architecture can decode UHD 4K videos at 120 frames/s in the worst case, requiring an operation frequency of 279.93 MHz and demanding a total area of 234.45 kgates with a power dissipation of 27.74 mW. The comparison with related works showed that AVID reached the smallest area and very competitive power results. To the best of the authors’ knowledge, this is the first article detailing the hardware design of a complete decoder for intraprediction targeting the AV1 codec. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | GM-RF: An AV1 Intra-Frame Fast Decision Based on Random ForestabstractThis paper presents the Grouping of Modes based on Random Forest (GM-RF), a fast decision algorithm for the AOMedia Video 1 (AV1) intra-frame prediction applying machine learning (ML). AV1 implements a wide variety of intra-frame prediction tools, significantly increasing the required computational effort. The GM-RF uses trained Random Forest (RF) models to reduce the number of intra-frame prediction modes evaluated for each encoded block. Experimental results show that the GM-RF achieves an average time savings of 50.19%, with a BD-BR of 7.41%. Compared with related works, GM-RF reached time savings from 5.6 to 10 times higher at a cost of a higher BDBR. To the best of the authors’ knowledge, this is the first solution in the literature using ML to reduce the AV1 intra-frame prediction computational effort. Pablo Rosa, Daniel Palomino 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 4 |
| 2022 | Mode-Adaptive Subsampling of SAD/SSE Operations for Intra Prediction Cost ReductionabstractModern video encoders, such as the recently proposed AV1 and VVC, offer significant encoding gains at the cost of a corresponding increase of the computational effort. This is the case of the adopted intra prediction techniques, comprehending an increased number of prediction modes and range. To mitigate this computational cost, the presented work proposes a new mode-adaptive algorithm that significantly reduces the number of SAD/SSE operations during intra prediction, by generating an optimized subsampling pattern adaptive to each prediction mode. The method can be applied to any video codec and, when applied to AV1, it led to an encoding time reduction and BD-BR impact of 15.36% and 0.6%, respectively, or 7.97% and –0.02%, depending on the selected subsampling parameters. When implemented in hardware, the proposed technique provides an effective reduction as high as 75% of both the area and power on the modified distortion calculation module. Marcel Moscarelli Corrêa, Nuno Roma, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 5 |
| 2022 | Fast Affine Motion Estimation for VVC using Machine-Learning-Based Early Search TerminationabstractThe Affine Motion Estimation (AME) was introduced in the Versatile Video Coding (VVC) standard to allow for the detection of non-translational transformations during inter-frame prediction. Although providing important coding efficiency gains, this new tool represents 43% of the motion estimation (ME) complexity. However, an analysis over the AME step shows that the Affine motion vectors are often generated without resulting in the best ME prediction. This paper proposes a AME early search termination based on supervised machine learning. Six Random Forest models were trained with features obtained during the encoding process to accurately predict whether the AME step should be executed, partially executed or skipped, avoiding unnecessary calculations. As result, the proposed solution achieves an average time saving of 46.94% in the AME step with a coding efficiency loss of only 0.18%. Adson Duarte, Luciano Volcan Agostini, Bruno Zatt, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Daniel Palomino 0001 |
ISCAS | 3 |
| 2022 | A High-Throughput Design for the H.266/VVC Low-Frequency Non-Separable TransformabstractThis paper presents a high throughput hardware design for the Low-Frequency Non-Separable Transform (LFNST) of the Versatile Video Coding (H.266/VVC) standard. The LFNST is a secondary transform used to transform the coefficients already transformed by the DCT-II as primary transform over the residues from the directional intra prediction. The LFNST architecture was designed to process Ultra-High Definition (UHD) videos with $4098 \times 2160$ pixels (4K) at 60 frames per second. Our solution presents an area utilization of 99.13 kgates and a power dissipation of 38.50 mW, when running at 186.62 MHz and considering the worst-case operation (processing the LFNST $4\times 4$ through TU size of $4\times 4$). Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 3 |
| 2022 | Fast Transform Decision Scheme for VVC Intra-Frame Prediction Using Decision TreesabstractThis paper presents a fast transform decision scheme for Versatile Video Coding (VVC) intra-frame prediction using decision trees. VVC introduces several novel coding tools to improve the coding efficiency of the intra-frame prediction at the cost of a high computational effort, including a new transform coding process using Multiple Transform Selection (MTS) for primary transform and Low-Frequency Non-Separable Transform (LFNST) for secondary transform. We developed an efficient complexity reduction scheme composed of two solutions based on decision tree classifiers to avoid the MTS and LFNST evaluations in the costly Rate-Distortion optimization (RDO) process. Experimental results showed that the proposed scheme provides 11% of encoding timesaving with a negligible impact on the coding efficiency. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 4 |
| 2022 | Multi-Objective optimized Complexity Control for the AV1 Video EncoderabstractAOMedia Video 1 (AV1) is an open-source video encoding format launched in 2018 by Alliance for Open Media. To achieve high coding efficiency, AV1 brings significant innovations in comparison to its predecessor, the VP9 format. However, this came at the cost of a higher encoding complexity due to the newly introduced tools and block partitioning structures. To allow for wide deployment of AV1 in different multimedia devices, adjusting its encoding process according to the available computational resources is strongly desirable. Thus, this work presents an adaptive complexity controller for the AV1 encoder based on multi-objective optimization. Experimental results show that the proposed controller reduces encoding complexity in a target range from 10% to 40%, with satisfactory precision varying between 0.04 and 4.10 percentage points, and a BD-Rate between 0.22% and 5.24%. Gustavo Rehbein, Isis Bender, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
PCS | 4 |
| 2022 | Configurable Fast Block Partitioning for VVC Intra Coding Using Light Gradient Boosting MachineabstractThis article presents a configurable fast block partitioning decision for Versatile Video Coding (VVC) intra-frame prediction using Light Gradient Boosting Machine (LGBM). VVC further improves the coding efficiency by introducing a Quadtree with nested Multi-Type Tree (QTMT), enabling five split types allowing square and rectangular Coding Unit (CU) sizes. However, this improvement in the coding efficiency comes at the cost of a high computational burden since several combinations of block sizes and prediction modes are evaluated through the costly Rate-Distortion Optimization (RDO) process. In this article, we propose a partitioning decision using LGBM classifiers to avoid the exhaustive RDO process and skip the evaluation of split types that are unlikely to be chosen as the best one. For this purpose, five classifiers (one for each split type) were offline trained with an efficient training process and using effective features of texture, coding, and context information. The proposed solution is highly configurable and can provide several operation points with different tradeoffs between timesaving and coding efficiency, according to the application requirements. Considering five operation points, the configurable solution can reduce the encoding time from 35.22% to 61.34%, with coding efficiency losses from 0.46% to 2.43%. Compared to the state-of-the-art, our solution is able to outperform the related works in terms of combined rate-distortion and timesaving. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | FastInter360: A Fast Inter Mode Decision for HEVC 360 Video CodingabstractThis paper presents FastInter360, a fast inter mode decision algorithm for accelerating the encoding of ERP 360 videos. The development of FastInter360 involves an in-depth and comprehensive set of evaluations performed to understand the differences in the encoder’s behavior when encoding 360 and conventional videos. These evaluations showed that due to the texture distortions resulting from projection, the encoder presents a specific behavior when encoding 360 videos, making it more likely to use a recurrent set of encoding modes when processing 360 videos. Besides, the coding efficiency is less sensible to approximations in some encoding steps depending on the frame region. FastInter360 is then proposed to reduce the encoding complexity by exploiting these differences. FastInter360 comprises three algorithms that accelerate the encoding by performing early decision by SKIP mode, reducing integer motion estimation search range, and adjusting fractional motion estimation precision. Furthermore, each of these algorithms behaves according to distortion intensity, performing greater complexity reduction in more distorted regions. When employed altogether, these algorithms compose FastInter360, which is able to achieve an average complexity reduction of 22.84% with a coding efficiency loss of 0.652% BD-BR, on average, making FastInter360 competitive with literature works. Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Low-Power and High-Throughput Approximated Architecture for AV1 FME InterpolationabstractModern video encoders like the AOM Video 1 (AV1) implement several complex tools to allow the required high level of compression efficiency. The Fractional Motion Estimation (FME) is one of these tools and in AV1 the FME defines 90 different filters. To handle such complexity, hardware acceleration using approximate computing has become an alternative to be explored. This paper presents an approximate solution for the AV1 FME interpolation filters based on the approximation of the original filter coefficients intending to generate more hardware friendly coefficients. The approximated version was designed in hardware and can achieve real-time interpolation for UHD 8K videos at 30 frames per second, when synthesized using 40nm TSMC standard-cells technology. The designed architecture dissipates 26.79mW which represents more than 80% power reduction when compared to the original precise solution. The approximation implied in a small average coding efficiency degradation of 0.54% in BD-BR. When comparing with related works, this architecture reaches an expressive power reduction (2.1 to 4.8 times) even supporting more complex tools. Robson Domanski, William Kolodziejski, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 6 |
| 2021 | SAD or SATD? How the Distortion Metric Impacts a Fractional Motion Estimation VLSI ArchitectureabstractVideo coding systems have to deal with a number of tradeoffs. The decision of adopting a specific distortion metric in the Fractional Motion Estimation (FME) step, for instance, presents a designer with a tradeoff between energy and coding efficiency. This paper analyzes such a tradeoff considering two of the most known and used distortion metrics, the Sum of Absolute Differences (SAD) and the Sum of Absolute Transformed Differences (SATD), within a High Efficiency Video Coding (HEVC)-compatible FME hardware architecture. We show that the SATD-based FME architecture is 1.94 times larger than the SAD-based one and consumes 2.07 times more energy. Ismael Seidel, Vanio Rodrigues Filho, Mateus Grellert, Luciano Volcan Agostini, José Luís Güntzel |
MMSP | 4 |
| 2021 | Analysis of VVC Intra Prediction Block Partitioning StructureabstractThis paper presents an encoding time and encoding efficiency analysis of the Quadtree with nested Multi-type Tree (QTMT) structure in the Versatile Video Coding (VVC) intra-frame prediction. The QTMT structure enables VVC to improve the compression performance compared to its predecessor standard at the cost of a higher encoding complexity. The intra-frame prediction time raised about 26 times compared to the HEVC reference software, and most of this time is related to the new block partitioning structure. Thus, this paper provides a detailed description of the VVC block partitioning structure and an in-depth analysis of the QTMT structure regarding coding time and coding efficiency. Based on the presented analyses, this paper can guide outcoming works focusing on the block partitioning of the VVC intra-frame prediction. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
VCIP | 4 |
| 2021 | Learning-Based Complexity Reduction Scheme for VVC Intra-Frame PredictionabstractThis paper presents a learning-based complexity reduction scheme for Versatile Video Coding (VVC) intra-frame prediction. VVC introduces several novel coding tools to improve the coding efficiency of the intra-frame prediction at the cost of a high computational effort. Thus, we developed an efficient complexity reduction scheme composed of three solutions based on machine learning and statistical analysis to reduce the number of intra prediction modes evaluated in the costly Rate-Distortion Optimization (RDO) process. Experimental results demonstrated that the proposed solution provides 18.32% encoding timesaving with a negligible impact on the coding efficiency. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
VCIP | 4 |
| 2021 | Using curved angular intra-frame prediction to improve video coding efficiency
Ramon Fernandes, Gustavo Sanchez, Rodrigo Cataldo, Luciano Volcan Agostini, César A. M. Marcon |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Performance analysis of VVC intra coding
Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Spatially Adaptive Intra Mode Pre-Selection for ERP 360 Video CodingabstractIn this work, we propose a spatially adaptive HEVC intra mode pre-selection for equirectangular (ERP) 360 video coding. The proposed technique exploits the spatial characteristics of 360 video in the ERP projection to reduce the complexity of intra prediction mode selection. The number of intra modes evaluated in Rate-Distortion Optimization is reduced based on a score technique that is adaptive to the frame region being encoded. Results show that the proposed technique achieves a complexity reduction of 16.5% with low coding efficiency penalties. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
ICASSP | 3 |
| 2020 | Memory Assessment Of Versatile Video CodingabstractThis paper presents a memory assessment of the next-generation Versatile Video Coding (VVC). The memory analyses are performed adopting as a baseline the state-of-the-art High-Efficiency Video Coding (HEVC). The goal is to offer insights and observations of how critical the memory requirements of VVC are aggravated, compared to HEVC. The adopted methodology consists of two sets of experiments: (1) an overall memory profiling and (2) an inter-prediction specific memory analysis. The results obtained in the memory profiling show that VVC access up to 13.4x more memory than HEVC. Moreover, the inter-prediction module remains (as in HEVC) the most resource-intensive operation in the encoder: 60%-90% of the memory requirements. The inter-prediction specific analysis demonstrates that VVC requires up to 5.3x more memory accesses than HEVC. Furthermore, our analysis indicates that up to 23% of such growth is due to VVC novel-CU sizes (larger than 64x64). Arthur Cerveira, Luciano Volcan Agostini, Bruno Zatt, Felipe Sampaio |
ICIP | 2 |
| 2020 | Complexity Analysis Of VVC Intra CodingabstractVersatile Video Coding (VVC) is the next-generation of video coding standards, which was developed to double the coding efficiency over its predecessor High-Efficiency Video Coding (HEVC). Several new coding tools have been investigated and adopted in the VVC Test Model (VTM), whose current version can improve the intra coding efficiency by 24% at the cost of a much higher coding complexity than the HEVC Test Model (HM). Thus, this paper provides a detailed VVC intra coding complexity analysis, which can support upcoming works for finding the most timeconsuming tool that could be simplified to achieve a real-time encoder design. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ICIP | 4 |
| 2020 | ASIC Solution for the Directional Intra Prediction of the AV1 Encoder Targeting UHD 4K VideosabstractAOMedia Video 1 (AV1), developed by the Alliance for Open Media consortium and released in 2018, is an open-source and royalty-free video format. It was designed to deliver substantial compression gains over its predecessor VP9 whilst keeping hardware feasibility and a practical decoding complexity. When compared to state-of-the-art formats, AV1 has more complex encoder tools, including the intra prediction which is the focus of this work. This paper presents a highly parallelized ASIC solution for the directional intra prediction module. It supports all 56 directional modes defined in AV1 and is able to process all combinations of block partitions. When synthesized to the TSMC 40nm technology with a target frequency of 1,296MHz, the proposed design used an area of 455.8K gates and showed a power dissipation and energy consumption per predicted sample of 40.92mW and 0.055pJ/sample, respectively. The reached throughput supports the processing of 60 frames per second for UHD 4K videos (3840×2160 pixels). No other work was found in the literature with a hardware design supporting the AV1 intra prediction directional modes. Marcel Moscarelli Corrêa, Luiz Neto, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 5 |
| 2020 | Low-Power and Memory-Aware Approximate Hardware Architecture for Fractional Motion Estimation Interpolation on HEVCabstractNowadays, current video coding standards like the High Efficiency Video Coding (HEVC) implement several complex coding tools, like the Fractional Motion Estimation (FME). An alternative to improve performance and save power is allying the hardware acceleration with approximate computing solutions, focusing on such complex tools. In this work, we present a low-power and memory-aware hardware architecture for the HEVC FME interpolator, proposing the development of two novel hardware designs for the interpolation filters, called Approximate Unified FME Filters (AUFF). These solutions exploit the usage of approximate computing at both algorithmic and data levels, leading to a reduction in dissipated power and memory bandwidth. The proposed design is capable of real-time interpolation of UHD (Ultra High Definition) 4K and 8K videos when synthesized using a 40 nm standard-cell library, with a power dissipation ranging from 22.04 to 62.06 mW. Wagner Penny, Guilherme Corrêa 0001, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Gabriel L. Nazar, Bruno Zatt |
ISCAS | 3 |
| 2020 | Fast Partitioning Decision Scheme for Versatile Video Coding Intra-Frame PredictionabstractThis paper presents a fast partitioning decision scheme for Versatile Video Coding (VVC) intra prediction. VVC adopts the tree coding block structure named Quadtree with nested Multi-Type (QTMT), which significantly improves the coding efficiency at the cost of a high computational effort, limiting its adoption for real applications. Thus, we developed a scheme that encompasses two strategies exploring the features of the current block and the encoding context through the selected intra prediction mode to skip unnecessary evaluations of binary and ternary partitions. Experimental results, obtained with VVC Test Model 5.0 and considering the Common Test Conditions (CTC) under All-Intra (AI) encoder configuration, demonstrated that the proposed scheme achieves 31.41% of coding time saving, on average, with negligible coding efficiency loss. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 4 |
| 2020 | Efficient Hardware Design for the AV1 CDEF Filter Targeting 4K UHD VideosabstractDeveloped by the AOMedia industry consortium, the AOM Video 1 (AV1) is an open-source and royalty-free video encoder released in June 2018. The Constrained Directional Enhancement Filter (CDEF) is one of the three AV1 in-loop filters and it is the focus of this work. The CDEF has the goal to reduce ringing artifacts generated with the encoding process, acting as a directional deringing filter. This paper presents a hardware design for the AV1 CDEF targeting real-time processing of 4K Ultra High Definition (UHD) videos. The architecture was synthesized to ASIC using the 40nm TSMC library, requiring 185 kgates and with a power dissipation of 43 mW when running at 93 MHz, reaching the frame rate of 60 frames per second (fps). To the best of the author's knowledge, there is no other work in the literature with dedicated hardware design for the AV1 CDEF. Eduardo Zummach, Roberta Palau, Jones Goebel, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2020 | ERP-Based CTU Splitting Early Termination for Intra Prediction of 360 videosabstractThis work presents an Equirectangular projection (ERP) based Coding Tree Unit (CTU) splitting early termination algorithm for the High Efficiency Video Coding (HEVC) intra prediction of 360-degree videos. The proposed algorithm adaptively employs early termination in the HEVC CTU splitting based on distortion properties of the ERP projection, that generate homogeneous regions at the top and bottom portion of a video frame. Experimental results show an average of 24% time saving with 0.11% coding efficiency loss, significantly reducing the encoding complexity with minor impacts in the encoding efficiency. Besides, solution presents the best results considering the relation between time saving and coding efficiency when compared with all related works. Bernardo Beling, Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
VCIP | 3 |
| 2020 | 4D-DCT Hardware Architecture for JPEG Pleno Light Field CodingabstractThis paper presents a 4D-DCT hardware architecture for Light Field Coding according to the JPEG Pleno standard. It is composed of two instances of 2D-DCT engines and a novel 4D Transposition Memory organization. Experimentally-defined fixed-point representation and LSB pruning techniques are employed do reduce hardware area and power dissipation. The proposed architecture operates over 4D-hypercubes of up to 8x8x8x8 samples and reaches performance to process 30 Lytro-like light fields per second dissipating 145.32mW at 825.75MHz. This is the first known 4D-DCT hardware architecture for light field coding and demonstrates the feasibility of such solutions on real-world systems. Matheus Jahnke, Jones Goebel, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 5 |
| 2020 | Complexity and compression efficiency assessment of 3D-HEVC encoder
Mário Saldanha, Ruhan A. Conceição, Vladimir Afonso, Giovanni Avila, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001, Luciano Volcan Agostini |
Multim. Tools Appl. | 9 |
| 2020 | Fast 3D-HEVC Depth Map Encoding Using Machine LearningabstractThis paper presents a fast depth map encoding for 3D-High Efficiency Video Coding (3D-HEVC) based on static decision trees. We used data mining and machine learning to correlate the encoder context attributes, building the static decision trees. Each decision tree defines that a depth map Coding Unit (CU) must be or not be split into smaller blocks, considering the encoding context through the evaluation of the encoder attributes. Specialized decision trees for I-frames, P-frames and B-frames define the partitioning of 64 × 64, 32 × 32, and 16 × 16 CUs. We trained the decision trees using data extracted from the 3D-HEVC Test Model considering all-intra and random-access configurations, and we evaluated the proposed approach considering the common test conditions. The experimental results demonstrated that this approach can halve the 3D-HEVC encoder computational effort with less than 0.24% of BD-rate increase on the average for all-intra configuration. When running on random-access configuration, our solution is able to reduce up to 58% the complete 3D-HEVC encoder computational effort with a BD-rate drop of only 0.13%. These results surpass all related works regarding computational effort reduction and BD-rate. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Online Machine Learning for Fast Coding Unit Decisions in HEVCabstractThe High Efficiency Video Coding standard introduced a flexible frame partitioning process that increased significantly compression rates in comparison to previous standards at the cost of a high computational cost. To accelerate frame partitioning decisions, this paper proposes a method that replaces the usual Rate-Distortion Optimization employed in Coding Unit size decision by a set of simpler decision tree models, which are built during encoding time by the C5 machine learning algorithm. The algorithm and the set of attributes employed in the model training process were chosen based on an extensive analysis that compared several options in terms of decision accuracy and training complexity. Experimental results show that the proposed method is capable of building accurate models for each video sequence, decreasing the HEVC encoding complexity in 34.4% on average with a compression efficiency loss of only 0.2% in comparison to the original HEVC reference encoder. Guilherme Corrêa 0001, Pargles Dall'Oglio, Daniel Palomino 0001, Luciano Volcan Agostini |
DCC | 4 |
| 2019 | FastIntra360: A Fast Intra-Prediction Technique for 360-Degrees Video Codingabstract360-degrees videos represent a whole sphere and enable the user to feel as if he is inside the scene. These videos demand more data than conventional videos to be represented, therefore they also must be compressed to be handled properly. However, current video coding standards only process rectangular videos, thus 360 videos must be represented in a flat fashion to be encoded. There are several projections to perform this and the currently most used one is the equirectangular projection (ERP), which transforms each parallel from the sphere into a row of the rectangle, resulting in a faithful representation of the equatorial area, and a stretched representation of the polar regions. This stretching in the polar regions tends to impact the behavior of intra-frame prediction, which is used to exploit the spatial redundancies in each frame. Therefore, this paper proposes FastIntra360 to accelerate the encoding of 360 videos. FastIntra360 is implemented in HEVC video coding standard [1], which is a recently established standard and poses high computational demand. During the development of FastIntra360, a set of videos were encoded and the behavior of the intra-prediction throughout the frame was extracted. Then, a statistical analysis was conducted over such data and it concluded that when encoding the polar regions of the frame, the prediction modes which exploit horizontal directions are selected more frequently than the remaining modes, whereas in the center of the frame all prediction modes present similar occurrence rates. FastIntra360 exploits this behavior to reduce the number of prediction modes evaluated in different regions of the frame to accelerate the encoding. FastIntra360 is developed in two variants: one considering three bands and other considering five bands, where each band is a horizontal stripe of the frame. Each band divides the frame samples into three or five stripes and performs the statistical analysis over these stripes individually. Both implementations were evaluated and compared against the HEVC Test Model version 16.16 (HM-16.16) according to time reduction and coding efficiency (considering BD-BR), where BD-BR represents the bitrate increase of the proposed technique. Experimental results showed that both implementations present good performance, reaching up to 16.5% complexity reduction with negligible BD-BR, that is, they present considerable complexity reduction whereas posing no harm to the video quality. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
DCC | 3 |
| 2019 | Encoding Efficiency and Computational Cost Assessment of State-Of-The-Art Point Cloud CodecsabstractPoint clouds have recently emerged as a suitable solution to generate and display 3D digital models due to their capacity of representing high resolution images and videos through multiple viewpoints. However, as they are usually made up of thousands up to billions of points, advanced techniques of data compression are essential to store and transmit this type of data. This paper compares the two state-of-the-art solutions for point cloud compression, the Point Cloud Codec (PCC) and the Test Model Category 2 (TMC2), in terms of compression efficiency and encoding time. Experimental results show that the compression efficiency for geometry information is highly dependent upon the available bitrate for both TMC2 and PCC. However, for texture compression TMC2 almost always achieves the best results. The experiments have also shown that TMC2 presents a computational cost from 22.2 to 26 times larger the observed in PCC. Mateus M. Gonçalves, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 2 |
| 2019 | High Throughput Hardware Design for AV1 Paeth and Smooth Intra ModesabstractDeveloped by AOMedia industry consortium and released in June 2018, AV1 is an open-source and royalty-free video coding format. The main goal of AV1 is to deliver substantial compression gains over state-of-the-art codecs such as VP9 and HEVC, while keeping a practical decoding complexity, hardware feasibility and its open and free status. This paper presents a high throughput hardware architecture for four important AV1 intra prediction coding modes: Paeth, Smooth, Smooth Vertical and Smooth Horizontal. The proposed architecture was designed to support all 19 block sizes specified by AV1 and to process every single combination of these blocks according to the 10-way partition tree, with a throughput of UHD 4K (3840×2160 pixels) videos at up to 30 frames per second. When synthesized to the TSMC 40nm cell library targeting a frequency of 648MHz, the proposed design used 109.57K gates and showed a power dissipation and an energy efficiency of 16.1mW and 1.23pJ/sample respectively. No other works were found in the literature describing hardware designs for AV1 intra prediction. Marcel Moscarelli Corrêa, Bianca Waskow, Bruno Zatt, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 6 |
| 2019 | TITAN: Tile Timing-Aware Balancing Algorithm for Speeding Up the 3D-HEVC Intra CodingabstractThis paper presents the Tile Timing-Aware balancing algorithm (TITAN) for speeding up the 3D-High Efficiency Video Coding (3D-HEVC) video encoding. The design of the TITAN algorithm is based on the premise that the encoding time of tiles partitioning of neighbor frames are similar. Therefore, based on the encoding time of the last-encoded frame, TITAN controls the tiles boundaries aiming to maximize the balance among tiles. Our software evaluation with TITAN implemented in 3D-HEVC Test Model 16.0 achieved an average of 6.2% higher speedup compared to the uniform-sized tiles. This is the first work in the literature proposing to speed up the 3D-HEVC video encoding by balancing the tiles workload. Mário Saldanha, Gustavo Sanchez, Bruno Zatt, César A. M. Marcon, Luciano Volcan Agostini |
ISCAS | 5 |
| 2019 | A Knapsack Methodology for Hardware-based DMR Protection against Soft Errors in Superscalar Out-of-Order ProcessorsabstractHigh-performance superscalar processors have been adopted to satisfy the rising demand for processing applications of ever-growing complexity. This extra complexity, added to the increasing vulnerability of transistors due to technology scaling, poses a great challenge since these effects have also been proven to affect ground-level safety-critical applications. To increase microarchitectural resilience, designers may adopt Dual Modular Redundancy (DMR), which offers full fault detection. However, given that DMR incurs in high area and energy overheads, we propose a design-time methodology aiming to achieve the best tradeoff between resilience and area overhead, decreasing DMR costs and maintaining acceptable detection levels for such a complex design. This is done by adopting the Knapsack Problem (KSP) as a heuristic to identify the optimal micro-architectural structures that should be duplicated to achieve target resilience with the smallest possible area overhead. By injecting over 800k faults in 12 significant micro-architectural structures of different versions of the complex Berkeley Out-of-Order Machine (BOOM) superscalar processor modeled with RTL accuracy, we compare this optimal strategy against a greedy one, showing that 90% of vulnerability reduction may be achieved with 50.6% and 107.8% area overheads for the optimal and greedy strategies, respectively. Rafael Billig Tonetto, Douglas Maciel Cardoso, Marcelo Brandalero, Luciano Volcan Agostini, Gabriel L. Nazar, José Rodrigo Azambuja, Antonio Carlos Schneider Beck |
VLSI-SoC | 4 |
| 2019 | Energy-Aware Motion and Disparity Estimation System for 3D-HEVC With Run-Time Adaptive Memory HierarchyabstractThe popularization of multimedia services has pushed forward the development of 2D/3D video-capable embedded mobile devices. Such devices require efficient energy/memory-management strategies to deal with severe memory/processing requirements and limited energy supply. Therefore, we propose a motion and disparity estimation (ME and DE) system—the most memory/processing demanding encoding steps—for the 3D High Efficiency Video Coding (3D-HEVC) standard. It was designed for low energy consumption, featuring a run-time adaptive memory hierarchy. The processing unit employs flexible coding order and optimizations to reduce the computational effort by exploring the inter-channel and inter-view redundancies. The memory hierarchy features window-based prefetching, data reuse, subsampling, and dynamic voltage scaling controlled by our depth-based dynamic search window resizing algorithm. Memory results demonstrate an average on-chip energy reduction of 79% in comparison to the widely used Level-C solution for a 45-nm technology. The proposed energy-aware ME and DE system dissipates 7.55 W while processing three HD 1080p views (video + depth) at 30 frames per second and presents a mean energy consumption of 0.107 J per access unit. To the best of our knowledge, this is the first work that proposes a real-time ME/DE system for the 3D-HEVC standard with an adaptive memory hierarchy. Vladimir Afonso, Ruhan A. Conceição, Mário Saldanha, Luciano Almeida Braatz, Murilo R. Perleberg, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt, Altamiro Amadeu Susin |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2019 | Performance Analysis of Depth Intra-Coding in 3D-HEVCabstractThe depth maps intra-frame prediction of 3D High Efficiency Video Coding (3D-HEVC) inherits all texture encoding techniques provided by HEVC and provides new coding tools for depth map predictions. These tools comprise algorithms, such as bipartition modes, intra-picture skip, and DC-only. This paper details these tools and shows how they work together with the original HEVC algorithms in the depth map intra-frame prediction for allowing high-efficiency encoding. Besides, this paper analyzes the encoding time and the encoding mode distribution of the intra-frame prediction tools over different quantization scenarios. We aim to provide support for upcoming works on depth map encoding, including complexity reduction and control, real-time embedded systems implementations, and even the development of improved tools to encode depth maps. Gustavo Sanchez, Jarbas Silveira, Luciano Volcan Agostini, César A. M. Marcon |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Intravenous Electromedical Equipment: A Proposal to Improve Accuracy in Generating AlertsabstractThe incidence of false alerts in a hospital environment undermines the health professionals' tasks, since: (i) they stressful the teams, their caregivers and the patients themselves; (ii) promote additional costs of time and attention; (iii) may cause risky situations, some of which have severe health implications in patients. Considering this scenario, the objectives of this work are to discuss the ocurrency of alerts in intravenous systems and to contribute to the reduction of the emission of false alerts in electromedical equipments, in particular in infusion pumps. To this purpose, the developed proposal, called BIRB, explores the use of Bayesian Networks to minimize the occurrence of false alerts when occurs a change in the flow rate caused by occlusions in infusion pumps. The occlusion is the procedure with the highest index of false alerts in this type of equipment. The achieved results with the BIRB proposal are promising, reaching 85% of accuracy, based on data from actual infusion pumps. Fabrício Neitzke Ferreira, Leonardo da Rosa Silveira João, João Ladislau Lopes, Adenauer C. Yamin, Luciano Volcan Agostini |
CLEI | 5 |
| 2018 | Octagonal-Axis Raster Pattern for Improved Test Zone Search Motion EstimationabstractTest Zone Search (TZS) is considered the current state-of-the-art fast Motion Estimation algorithm because it presents the best tradeoff between compression efficiency and complexity in comparison to the Full Search strategy. However, it is still one of the most computationally-demanding tools of current video coding standards, such as the High Efficiency Video Coding (HEVC). This paper presents an analysis on the search area opportunities and best match distributions in TZS, which led to the proposal of a novel search pattern in its most complex step, the Raster Search (RS). The new pattern, named Octagonal-Axis Raster Pattern (OARP), allowed an average complexity reduction of 61 % in TZS, with a negligible BD-rate increase of 0.0371 % in comparison to the original algorithm. Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001 |
ICASSP | 4 |
| 2018 | Fast 3D-Hevc Depth Maps Intra-Frame Prediction Using Data MiningabstractThis paper presents a fast 3D-High Efficiency Video Coding (3D-HEVC) depth maps intra-frame prediction based on static Coding Unit (CU) splitting decisions trees. This coding approach uses data mining to extract the correlation among the encoder context attributes and to define a split decision tree for each CU level of the depth maps encoding. The decision trees were trained using the information extracted from 3D-HEVC Test Model (3D-HTM) and using the Common Test Conditions (CTC). Each decision tree defines if the current CU must be split into smaller sizes, considering the encoding context through the evaluation of some current encoder attributes. The proposed solution reaches a complexity reduction of 59.0% for depth maps coding with a negligible impact of 0.18% in the encoding efficiency of synthesized views. Mário Saldanha, Gustavo Sanchez, César A. M. Marcon, Luciano Volcan Agostini |
ICASSP | 4 |
| 2018 | LF-CAE: Context-Adaptive Encoding for Lenslet Light Fields Using HEVCabstractLight fields can outperform the representability of current imaging technologies by considering the angle of light-rays striking in the camera in addition to the color information. The additional information brings a set of challenges related to the amount of data required to represent light fields, arising the need for efficient compression schemes. This work proposes a novel and efficient scheme to encode lenslet light fields, called Light Fields Context-Adaptive Encoding(LF-CAE). LF -CAE is a video-based compression solution that defines a flexible and dynamic scheme to create intermediate video sequences from light fields in order to efficiently exploit the HEVC inter-frame encoder structure. This scheme reduces the inter-frame prediction residue leading to compression efficiency gains in light-field coding. LF -CAE reaches an average compression rate of 99.56% in relation to the uncompressed light field, with an average PSNR of 37.61dB. LF-CAE surpasses all related works that decompose the light field into an intermediate video sequence, reaching the highest PSNR and with BD-Rate gains ranging from 6.07% up to 20.50%. Ruhan A. Conceição, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 4 |
| 2018 | DCDM-Intra: Dynamically Configurable 3D-HEVC Depth Maps Intra-Frame Prediction AlgorithmabstractThis work proposes the Dynamically Configurable 3D-HEVC Depth Maps Intra-Frame Prediction (DCDM-Intra), which explores the fact that Rough Mode Decision (RMD) was inherited from texture coding without taking advantages of the depth maps simplicity. DCDM-Intra classifies the HEVC Intra-frame Prediction Modes (IPMs) according to their BD-rate impact when encoding depth maps. Then, the application or the user can dynamically define the number of IPMs supported by the depth maps intra-prediction, according to the system status or application requirements, removing the IPMs that less affect the encoding efficiency. DCDM-Intra allows 35 distinct operation points with different encoding efficiency impacts. Experimental results demonstrate that the exclusion of 20% of IPMs causes a BD-rate increase of 0.03%, and the removal of almost 51% of IPMs rises 0.11% the BD-rate. Gustavo Sanchez, Ramon Fernandes, Luciano Volcan Agostini, César A. M. Marcon |
ICIP | 3 |
| 2018 | Hardware-Friendly Unidirectional Disparity-Search Algorithm for 3D-HEVCabstractThis paper presents a novel hardware-friendly Unidirectional Disparity-Search (UDS) algorithm for the 3D-HEVC. This algorithm explores the typical camera arrangements used in 3D-HEVC. UDS was evaluated in two operation points reaching a computational effort reduction from 32.7% to 61.8%, with a BD-Rate increase from 0.3123% to 0.4803%, when compared to TZS. Estimated hardware results showed memory-size and leakage-energy reductions of 66.7%, a dynamic-energy reduction from 34.5% to 62.1%, and an energy-consumption reduction from 32.7% to 61.8% in SAD calculations, when compared to TZS. To the best of the authors' knowledge, this is the first proposed DE algorithm that explores the 3D-HEVC typical camera arrangement. Vladimir Afonso, Altamiro Amadeu Susin, Murilo R. Perleberg, Ruhan A. Conceição, Guilherme Corrêa 0001, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 6 |
| 2018 | High-Throughput and Low-Power Integrated Direct/Inverse HEVC Quantization Hardware DesignabstractThis paper presents a high-throughput and low-power integrated HEVC direct/inverse quantization hardware design. The main focus of this design is to allow the evaluation of multiple coding modes during the residual encoding process of the HEVC for real-time Ultra-High Definition (UHD) video processing. The ASIC synthesis results, for a Nangate 45nm standard cell library, presented a maximum operational frequency of 1679.51MHz and a processing rate of 53.74 Gsps (giga samples per second). This throughput allows processing of real-time up to 72 coding modes for UHD 4K@60fps or up to nine coding modes for the UHD 8K@120fps while dissipating 369.37mW. Luciano Almeida Braatz, Bruno Zatt, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2018 | Configurable Cache Memory Architecture for Low-Energy Motion EstimationabstractThe popularization of mobile devices and the increased demand for video applications from these devices necessitates the design of efficient video encoders such as HEVC. Since the Motion Estimation (ME) is the most processing and memory intensive unit in a video encoder, our focus is in the communication between external memory and the ME unit. The TZS algorithm is widely used in video encoders and has an unpredictable behavior, which leads to an unknown pattern of memory accesses, making SPMs ineffective solutions, for example. Therefore, this work proposes a configurable cache memory architecture for fast ME algorithms. This cache has settings that suit different video encoding scenarios. Six optimal cache configurations were defined based on our evaluation considering 23 video sequences, 4 QPs, and 32 different cache settings. External memory bandwidth savings of up to 96.84% were reached, representing a reduction from 25.48GB/s to 548.53MB/s in the best case. When compared to Level-C SPM and to a static 16KB 8-way associative cache, the proposed configurable cache achieves energy savings of up to 86.91% and 78.09%, respectively. Anderson Martins, Wagner Penny, Matheus Weber, Luciano Volcan Agostini, Marcelo Schiavon Porto, Daniel Palomino 0001, Júlio C. B. de Mattos, Bruno Zatt |
ISCAS | 4 |
| 2018 | High Efficient Architecture for 3D-HEVC DMM-1 Decoder Targeting 1080p VideosabstractThis paper presents an efficient hardware design for the Depth Modeling Mode 1 (DMM-1) decoder of the 3D-High Efficiency Video Coding (3D-HEVC). The designed architecture uses a lossless wedgelet memory compression technique to reduce the used memory, and a well-balanced parallelism level to allow the desired throughput at the minimum possible power consumption and area usage. The architecture was synthesized for the 65nm ST standard cells technology, using 4,047 gates and consuming 0.95mW. It is capable of processing 1080p videos at 30 frames per second, decoding all allowed block sizes and wedgelets patterns. Besides, the proposed architecture saves 10.7% of area and 29.1% of power when compared with a version without memory compression. At the best of the author's knowledge, this is the first work with a dedicated hardware design targeting the DMM-1 decoder. Gustavo Sanchez, Luciano Volcan Agostini, César A. M. Marcon |
ISCAS | 2 |
| 2018 | Coding- and Energy-Efficient FME Hardware DesignabstractHybrid video standards rely on encoding prediction residues. To improve coding efficiency of inter-frame prediction, interpolated samples may be generated in fractional positions i.e., between neighbor pixels in the reference frame. However, performing Fractional Motion Estimation (FME) increases the overall encoder complexity. Since portable mobile devices are increasingly used to capture and reproduce videos, energy-efficient FME hardware accelerators are of utmost importance. In this work, we propose and evaluate a coding- and energy-efficient hardware design strategy for FME. Such strategy addresses the main weaknesses of the architectures found in the literature. The architecture designed as case study can achieve 2160p@120fps for the HEVC 8×8 FME. We also provide an insightful area and power breakdown of the synthesized design, to drive the design of FME hardware towards further energy improvements. Ismael Seidel, Vanio Rodrigues Filho, Luciano Volcan Agostini, José Luís Güntzel |
ISCAS | 3 |
| 2018 | A reduced computational effort mode-level scheme for 3D-HEVC depth maps intra-frame prediction
Gustavo Sanchez, Luciano Volcan Agostini, César A. M. Marcon |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Low-power and high-throughput hardware design for the 3D-HEVC depth intra skipabstractThis paper presents a low-power and high-throughput hardware design for the 3D-HEVC (Three Dimensional High Efficiency Video Coding) Depth Intra Skip coding tool. A strategy to reduce the computational effort was employed based on an analysis using the 3D-HEVC reference software. The proposed strategy consists of replacing the SVDC (Synthesized View Distortion Change) for the SAD (Sum of Absolute Differences) as the similarity criterion. This way, the number of arithmetic operations related with the similarity criterion is reduced over 71%, and a rendering process is avoided at the cost of only 0.21% increase in the BD-Rate. The hardware was described in VHDL and synthesized for ASIC technology. The synthesis results for the 45nm Nangate standard cells demonstrate that the architecture can process 60 UHD 2160p frames per second (five views) with a power dissipation of 19.57mW. Vladimir Afonso, Altamiro Amadeu Susin, Luan Audibert, Mário Saldanha, Ruhan A. Conceição, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 8 |
| 2017 | A multiplierless parallel HEVC quantization hardware for real-time UHD 8K video codingabstractOne step required several times for current video encoders is the residual coding loop, composed of the direct transformation, direct quantization, inverse quantization, and inverse transformation. These operations demand high throughput and low latency since their outputs must be processed by other steps of the coder. This paper proposes a high-throughput parallel and multiplierless hardware architecture for the HEVC direct quantization targeting real-time processing of Ultra-High Definition 8K videos. The proposed architecture support frequency dependent quantization steps. The binary multiplications were replaced by multiple constant multiplications in order to improve the throughput and to reduce the area and power dissipation. The developed design is able to process 32 samples in parallel, which represents one line of the biggest HEVC transform block. The ASIC synthesis results, obtained with Nangate 45nm standard cells library, show that the proposed architecture is able to quantize about 8 billion coefficients per second, when running at 186.6 MHz, with a gate count of 168,330. This throughput is enough to process UHD 8K videos at 120 fps. Luciano Almeida Braatz, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 2 |
| 2017 | High-throughput HEVC intrapicture prediction hardware design targeting UHD 8K videosabstractThis paper presents a high-throughput hardware architecture for the HEVC intrapicture prediction targeting the processing of UHD 8K (7680×4320 pixels) videos at 120 frames per second. The proposed design supports all intra prediction modes and all block sizes. It implements an internal mode decision algorithm that is more hardware friendly than RMD at the cost of a negligible 0.17% BD-rate impact (intra-only). When synthesized to the NanGate 45nm 0.95v cell library targeting a frequency of 529MHz, the proposed design used 4952K gates and showed a power dissipation and an energy efficiency of 363mW and 32.02pJ/sample respectively. Marcel Moscarelli Corrêa, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 4 |
| 2017 | Complexity reduction by modes reduction in RD-list for intra-frame prediction in 3D-HEVC depth mapsabstractThis paper presents a complexity reduction technique for 3D High Efficiency Video Coding (3D-HEVC) using Rate-Distortion (RD) list reduction in the depth maps intra-frame prediction. In the 3D-HEVC standardization, the HEVC intra-frame prediction was inherited from texture to depth maps. However, considering depth maps simpler behavior than texture, the intra-frame prediction contains an unnecessary complexity since there is a high number of modes evaluated by their RD-cost. This paper evaluates the usage of fewer best-ranked modes in RD-list. As a case study, 23.9% of complexity reduction was achieved by reducing the Rough Mode Decision (RMD) selection and the Most Probable Modes (MPM) algorithm to the selected configuration with small impact in BD-rate. Gustavo Sanchez, Luciano Volcan Agostini, César A. M. Marcon |
ISCAS | 2 |
| 2017 | Multiple early-termination scheme for TZ search algorithm based on data mining and decision treesabstractThe latest video compression standards, such as the H.264/AVC and the High Efficiency Video Coding (HEVC), provide fast Motion Estimation (ME) algorithms in their reference software aiming at complexity reduction. Test Zone Search (TZS) is the state-of-the-art fast ME algorithm, currently deployed in the reference HEVC encoder due to its great coding efficiency. However, ME is still one of the main sources of complexity in HEVC. This paper proposes an early-termination scheme for TZS, called e-TZS, based on an extensive data mining process on ME attributes. The data mining process allowed identifying the most relevant information during the encoding process to build a set of decision tree models that terminate TZS in different steps of its execution. The e-TZS scheme was implemented in the HEVC reference software and achieved average decision precision of 94.2%. Experimental results showed an average complexity reduction of 62.53% in TZS, with a negligible BD-rate increase of only 0.49%, in comparison to the original algorithm. Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
MMSP | 5 |
| 2017 | Energy-efficient motion estimation with approximate arithmeticabstractEnergy efficiency has become a primary concern in the design of multimedia digital systems, particularly when targeting mobile devices. Approximate computing is a highly promising approach to address this challenge. This paper presents an architectural exploration in a variable block size motion estimation (VBSME) architecture using imprecise Lower-Part-OR Adders (LOA). These adders were applied to Sum of Absolute Differences units (SAD) in order to reduce the energy consumption while introducing a minimum impact on the coding efficiency. Three VBSME architectures with LOA operators were developed by considering different imprecision levels. The conducted evaluations, performed using the High-Efficiency Video Coding standard (HEVC) reference software, showed that this technique introduces a negligible impact on the coding efficiency (between 0.6% and 2.5% increase of the BD-Rate). Nevertheless, when the designed architectures were synthesized for a 45nm standard cells technology, significant power savings were observed (between 7% and 11.5%, depending on the used LOA version), demonstrating the viability and significant gains of the proposed approach. Roger Endrigo Carvalho Porto, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto, Nuno Roma, Leonel Sousa |
MMSP | 2 |
| 2016 | Complexity reduction for 3D-HEVC depth map coding based on early Skip and early DIS schemeabstractThis paper presents a novel early Skip/DIS mode decision for 3D-HEVC depth encoding which aims at reducing the complexity effort of this process. The proposed solution is based on an adaptive threshold model, which takes into consideration the occurrence rate of both Skip and DIS modes. Occurrence analysis showed that the lower is the Skip and DIS Rate-Distortion cost, the higher is the probability of these modes being chosen. Furthermore, software evaluations showed that the proposed early Skip/DIS scheme is capable of reducing the depth coder complexity in 24.4% for a target hit rate of 99%, and in 33.7% for a target hit rate of 95%, leading to a negligible coding efficiency penalty in both scenarios. Ruhan A. Conceição, Giovanni Avila, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 6 |
| 2016 | High-throughput and memory-aware hardware of a sub-pixel interpolator for multiple video coding standardsabstractReal-time operation and low-power dissipation in video coding systems have become important research challenges, especially in mobile devices with limited battery and computational resources. There are many video coding standards coexisting in the market nowadays, so it is important for current devices to support different video coding standards. This paper presents a multi-standard luminance sub-samples interpolator hardware design for the Motion Compensation (MC) and Fractional Motion Estimation (FME), with support to MPEG-2/4, H.264/AVC, HEVC, and AVS/2 video coding standards. Our design is able to save hardware resources through an optimized filter organization, totally compliant with the focused standards and capable to interpolate samples for UHD 4320p@60fps at real time. The 45nm standard-cell library implementation dissipates 10mW, when processing according MPEG-2 standard, up to 46.4mW when processing AVS2. Guilherme Paim, Jones Goebel, Wagner Penny, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 6 |
| 2016 | An efficient sub-sample interpolator hardware for VP9-10 standardsabstractThis paper presents a hardware design for the sub-sample interpolator used in FME (Fractional Motion Estimation) and MC (Motion Compensation) stages according to the VP9 and VP10 video-coding standards. The proposed architecture is able to save hardware resources through an optimized-filter organization whereas reaching high-throughput and low-power dissipation. The hardware design was described in Verilog and synthesized for ASIC technology. The synthesis results were generated for 45nm Nangate standard cells and demonstrate that the developed architecture is able to process 2160p@60fps videos with a power dissipation of 2.34mW focusing on a VP9-10 decoder. Guilherme Paim, Wagner Penny, Jones Goebel, Vladimir Afonso, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 8 |
| 2016 | Pareto-based energy control for the HEVC encoderabstractThe current state-of-art video coding standard, the High Efficiency Video Coding (HEVC), brings many innovations as a way to improve the coding performance. However, the improvement on performance also brought higher computational effort and energy consumption. Since most of devices that handle digital videos are battery powered, the energy consumption became an important issue that demands efficient solutions. This way, controlling energy consumption is strongly desirable to adapt the encoding process to the energy availability. This goal is a hard task due the heterogeneous dynamic behavior of HEVC encoder. This work presents the development of a Pareto-based dynamic energy controller for the HEVC encoder, reaching up to 70% energy saving with small losses on coding efficiency for most of the cases. Wagner Penny, Italo Machado, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt |
ICIP | 4 |
| 2016 | Rate-constrained successive elimination of Hadamard-based SATDsabstractThe efficiency improvements achieved by new video coding standards come at the cost of a huge increase in the encoder computational complexity. Paradoxically, such increasing complexity is commonly addressed by methods that have an adverse effect on coding efficiency. In this work, we propose a method to reduce the complexity of HEVC Hadamard ME, without compromising coding efficiency. Our method relies on a new metric named Absolute First Difference (AFD), which is able to eliminate impossible candidates using only a few operations. Despite its simplicity, AFD filters an average of 13.52% of SATDs in the HEVC reference model (HM), without reducing coding efficiency. Moreover, our method is able to filter up to 37.9% of SATDs for video conferencing and up to 64% for screen content. Ismael Seidel, Luiz Henrique Cancellier, José Luís Güntzel, Luciano Volcan Agostini |
ICIP | 4 |
| 2016 | Speedup-aware history-based tiling algorithm for the HEVC standardabstractThis paper proposes a history-based tiling algorithm aiming at the increase of speedup when using Tiles. The algorithm is composed of two independent steps that use workload history information to define the vertical and horizontal boundaries of the Tiles. The workload distribution of previous frames are used as reference to perform the tiling of the current frame exploiting the temporal similarity between neighboring frames. Experimental results show that the proposed algorithm outperforms the speedup when compared to uniform tiling by 6.85% on average for tested sequences, besides, there is no significant complexity increase and similar coding efficiency results. When compared to related works the proposed solution also presents better speedup results. Iago Storch, Daniel Palomino 0001, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 4 |
| 2016 | Fast H.264/AVC to HEVC transcoder based on data mining and decision treesabstractHigh Efficiency Video Coding (HEVC) is gradually replacing its predecessor, the H.264/AVC standard, as the state-of-the-art technology for video compression. However, H.264/AVC has dominated the market for over a decade, so that there is an enormous amount of legacy content that must be migrated. This paper proposes a fast transcoder based on an extensive data mining process on H.264/AVC decoding attributes. The data mining allowed identifying relevant information from the H.264/AVC decoding process, which was conveyed to the C4.5 machine learning algorithm to build a set of decision trees that simplify the complex Coding Unit (CU) size decision in HEVC. Experimental results have shown an average reduction of 44% in the transcoding time, with a small bit rate increase of 1.67%. These results outperform any previous works available in the literature. Guilherme Corrêa 0001, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
ISCAS | 2 |
| 2016 | An HEVC multi-size DCT hardware with constant throughput and supporting heterogeneous CUsabstractThis paper presents an efficient hardware design for the Discrete Cosine Transform (DCT) of High Efficiency Video Coding standard (HEVC). This hardware supports all HEVC transform sizes: 4×4, 8×8, 16×16, and 32×32 including any combination of the Transform Unit (TU) sizes. The proposed DCT architecture has a constant throughput of 32 coefficients per cycle, independently of the transform sizes combination. The architecture was synthesized for a Nangate 45nm standard-cell library and the power analysis was made considering real input vectors. The synthesis results show a very good tradeoff between area, power dissipation and processing rates. The architecture is able to process 1.6G coeff/s when running at 50MHz dissipating 24.2 mW. These results allow a processing rate of 30 HD 1080p frames per second when evaluating 17 HEVC prediction modes. Jones Goebel, Guilherme Paim, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 3 |
| 2016 | Energy-efficient SATD for beyond HEVCabstractState-of-the-art video coding standards adopt large block and transform sizes. Moreover, as resolutions keep growing, there is a trend in adopting even larger structures in future video encoders, resulting in higher complexity. Therefore, the design of energy-efficient architectures for variable block size distortion metrics are key to keep the energy requirements of battery devices under a reasonable budget. The Hadamard-based Sum of Absolute Transformed Differences (SATD) is used as distortion metric in several steps of encoding, increasing the overall encoding efficiency at the cost of rising complexity. In this work, we propose two main approaches for SATD calculation of N × N block sizes. One using a Transpose Buffer (TB) and another one using a Linear Buffer (LB). Furthermore, we synthesized four sizes (4 × 4 up to 32 × 32) of each main SATD architecture to evaluate their area and energy estimates. The results show a large increase in area for TB-SATD, while LB-SATD area increases in a much smaller pace. On the other hand, the TB-SATD synthesized for Low-Vdd/High-Vt show up to be the most energy-efficient architectures for all sizes. Ismael Seidel, André Beims Bräscher, José Luís Güntzel, Luciano Volcan Agostini |
ISCAS | 4 |
| 2016 | Pareto-Based Method for High Efficiency Video Coding With Limited Encoding TimeabstractSeveral different methods have been investigated in recent years, aiming at computational complexity reduction and scaling of High Efficiency Video Coding (HEVC) software implementations. However, maintaining the encoding time per frame or group of pictures (GOPs) below an adjustable upper bound is still an open research issue. A solution for this problem is devised in this paper based on a set of Pareto-efficient encoding configurations, identified through rate-distortion-complexity analysis. The proposed method combines a medium-granularity encoding time control with a fine-granularity encoding time control to accurately limit the HEVC encoding time below a predefined target for each GOP. It is shown that the encoding time can be kept below a desired target for a wide range of encoding time reductions, e.g., up to 90% in comparison with the original encoder. The results also show that compression efficiency loss (Bjøntegaard delta-rate) varies from negligible (0.16%) to moderate (9.83%) in the extreme case of 90% computational complexity reduction. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | A multi-standard interpolation hardware solution for H.264 and HEVCabstractAttending real-time constraints in video coding systems represents a big challenge for nowadays systems, especially for high definition videos at mobile systems. The Fractional Motion Estimation (FME) and Motion Compensation (MC) are responsible for a large share of processing effort in both state-of-the-art video coding standards, the High Efficiency Video Coding (HEVC), and its predecessor, the H.264. This work proposes a multi-standard hardware solution for the fractional sample interpolation used in FME/MC processing of the HEVC and H.264 standards. The hardware design is composed of four IP (Intellectual Property) cores able to process 1080p@60fps videos independently. The whole architecture can process 2160p@60fps with 80.69mW, considering bi-prediction. Henrique Maich, Guilherme Paim, Vladimir Afonso, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ICIP | 4 |
| 2015 | Fast mode selection algorithm based on texture analysis for 3D-HEVC intra predictionabstractThe increasing availability of 3D video systems and applications has attracted more consumers for 3D viewing experiences and, consequently, the demand for storage and transmission of 3D video content is growing. An interesting alternative for this need is the transmission of 3D video based on the Multiview Video plus Depth (MVD) format. The upcoming 3D High Efficiency Video Coding (3D-HEVC) standard will adopt this format, which associates a depth map to each texture frame. This paper presents a fast mode decision algorithm, which analyses the texture frames and depth maps to detect the edge orientation of the prediction units (PUs), optimizing the intra prediction process and reducing the 3D-HEVC computational complexity. The experimental results show that the proposed algorithm achieves an average processing time reduction of 26.2%, with a small degradation in encoding efficiency (BD-rate increase of 0.3% on average). Thaísa Leal da Silva, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
ICME | 2 |
| 2015 | Encoding time control system for HEVC based on Rate-Distortion-Complexity analysisabstractThe improved compression efficiency of High Efficiency Video Coding (HEVC) comes with large increases in computational complexity, which has been lately dealt by researchers with complexity reduction and scaling methods. However, encoding time control at frame or Group of Pictures (GOP) level is still an open issue that must be investigated. In this work, a Rate-Distortion-Complexity analysis is performed upon a set of configurations that have been created based on findings and techniques of previous works. The proposed control system uses the best 27 configurations to adjust the encoding time per GOP and yields encoding time reductions of up to 84.5% with an average difference between target and encoding times of 4.1%. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
ISCAS | 4 |
| 2015 | Complexity reduction for the 3D-HEVC depth maps codingabstractThis paper presents a qualitative discussion of the depth maps properties that can be considered to achieve complexity reduction for 3D-High Efficiency Video Coding (3D-HEVC) depth maps coding. Both intra and inter-frame predictions are considered in this discussion that conduced to the proposition of two simple complexity reduction techniques: the Simplified Edge Detector (SED) and the Diamond Search (DS) simplified inter-prediction. The SED anticipates the blocks that are likely to be better predicted by the HEVC intra-prediction, avoiding evaluations of Depth Modeling Modes (DMM). The DS and SED were compared to anchor results and experimental analysis showed that the proposed algorithms are able to achieve a time saving of 11.3% encoding time reduction, with acceptable impact on the BD-Rate of the synthesized views of 0.6%. Mário Saldanha, Gustavo Sanchez, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 5 |
| 2015 | A real-time architecture for reference frame compression for high definition video codersabstractCurrent battery-powered devices that manipulate digital videos must consider the energy consumption of this process as an important issue, especially when high or ultra-high definition videos are handled. In this scenario, this paper proposes a solution to reduce the energy consumption in video coding systems by reducing the external memory communication during the motion estimation. The scheme presented in this paper is called Differential Reference Frame Coder and it implements an algorithm that combines two techniques to reduce the memory bandwidth: a differential coding based on a simplified intra-prediction process, to reduce the spatial redundancy of the reconstructed samples, and a semi-fixed length coding applied in the residues generated by the differential coding step. This solution reaches an average lossless compression ratio higher than 57% for the evaluated HD 1080p video sequences whereas supporting random access to reference frame blocks. The proposed hardware architectures (Coder and Decoder) were described in VHDL and synthesized targeting ASIC for 65nm and 180nm TSMC standard-cell libraries. The results show that with 65nm, the architectures are able to process UHD 2160p (3840×2160 samples) at 30 fps or HD 1080p (1920×1080 samples) at 120 fps with a power dissipation of 0.885mW. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2015 | Fast HEVC Encoding Decisions Using Data MiningabstractThe High Efficiency Video Coding standard provides improved compression ratio in comparison with its predecessors at the cost of large increases in the encoding computational complexity. An important share of this increase is due to the new flexible partitioning structures, namely the coding trees, the prediction units, and the residual quadtrees, with the best configurations decided through an exhaustive rate-distortion optimization (RDO) process. In this paper, we propose a set of procedures for deciding whether the partition structure optimization algorithm should be terminated early or run to the end of an exhaustive search for the best configuration. The proposed schemes are based on decision trees obtained through data mining techniques. By extracting intermediate data, such as encoding variables from a training set of video sequences, three sets of decision trees are built and implemented to avoid running the RDO algorithm to its full extent. When separately implemented, these schemes achieve average computational complexity reductions (CCRs) of up to 50% at a negligible cost of 0.56% in terms of Bjontegaard Delta (BD) rate increase. When the schemes are jointly implemented, an average CCR of up to 65% is achieved, with a small BD-rate increase of 1.36%. Extensive experiments and comparisons with similar works demonstrate that the proposed early termination schemes achieve the best rate-distortion-complexity tradeoffs among all the compared works. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | A low-complexity and lossless reference frame encoder algorithm for video codingabstractThis paper presents a lossless coding solution to reduce the large overhead of external memory communication during the motion estimation process in current video coders. Our solution is called Differential Reference Frame Coder (DRFC), and uses two techniques together to compress the reference frame: a differential coding based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples, and a simple VLC applied to differential coding residues. The proposed solution reaches an average compression rate higher than 45% for the evaluated HD 1080p video sequences. This is a lossless and low-complexity solution, and could easily be implemented in hardware. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ICASSP | 5 |
| 2014 | Complexity reduction for 3D-HEVC depth maps intra-frame prediction using simplified edge detector algorithmabstractThis paper presents a new mode decision for the depth maps intra-frame prediction in 3D-HEVC. The proposed technique decides if the traditional High Efficiency Video Coding-based (HEVC) intra-frame prediction should be performed or skipped. This technique is inspired by the fact that traditional intra-frame prediction may generate artifacts in the synthesized views when an edge is encoded. The Simplified Edge Detector (SED) algorithm has been proposed to classify if a block contains an edge or a nearly constant region demanding a minimum processing overhead. Through software evaluations, SED algorithm was capable to obtain an average complexity reduction of 23.8% for depth maps coding with no quality losses. Gustavo Sanchez, Mário Saldanha, Gabriel Balota, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 6 |
| 2014 | A new differential and lossless Reference Frame Variable-Length Coder: An approach for high definition video codersabstractThis paper presents a novel solution for external memory bandwidth reduction in video coding systems. The approach is based on reference frame compression, using a differential coding and a hardware-aware adaptation of the traditional Huffman algorithm, besides, it is a lossless solution fully compliant to state-of-art video coding standards, as H.264/AVC and HEVC. This solution is called DRFVLC (Differential Reference Frame Variable-Length Coder) and it uses differential coding to concentrate the samples values distribution. With the samples concentrated, an efficient static Huffman coding is applied to represent them in fewer bits. The DRFVLC reaches an average compression rate higher than 60% for the evaluated HD 1080p video sequences. This compression rate also indicates the external memory bandwidth reduction achieved with our technique. This solution can be easily implemented in hardware demanding one differentiator and a simple variable-length coder. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ICIP | 5 |
| 2014 | Power efficient and high troughtput multi-size IDCT targeting UHD HEVC decodersabstractThis paper is focused on the inverse transforms defined in the HEVC (High Efficiency Video Coding) standard. The HEVC standard allows the use of four transform sizes, including novel transforms applied over bigger block sizes (16×16 and 32×32). The hardware architecture presented in this paper was planned to reach real-time processing (at 30 frames per second) for ultra-higher solution videos, exploiting high level of parallelism. As a secondary goal, the architecture was also planned to reach low cost in terms of hardware consumption and power dissipation. Thus, the architecture was designed in a purely combinational way, using a multiplierless approach and employing an optimization algorithm through operations reuse and sub-expressions sharing. The synthesis targeted an Altera Stratix V FPGA and ASIC 90nm standard-cells technology. The synthesis results show that the designed architecture has the best performance results among all related works, being able to achieve real-time decoding for UHD videos (7680×4320 pixels) with a power consumption from 33.8mW to 339.2 mW. Ruhan A. Conceição, J. Claudio de Souza, Ricardo Jeske, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 6 |
| 2014 | Memory bandwidth reduction for H.264 and HEVC encoders using lossless reference frame codingabstractThis paper presents a hardware-efficient algorithm for external memory bandwidth reduction focusing on the state-of-the-art video encoders, like H.264/AVC and HEVC. The proposed approach is a lossless solution based on an adaptation of the traditional Huffman algorithm. This solution is entitled RFCAVLC8T (Reference Frame Context Adaptive Variable-Length Coder with 8 Tables) and is based on the use of off-line defined static Huffman tables. The RFCAVLC8T is a hardware-efficient version of the Huffman algorithm that employs eight static tables to avoid the cost of the on-the-fly Huffman statistical analysis. The best table to encode a block is selected at run time using a context evaluation, resulting in a context-adaptive configuration. The use of RFCAVLC8T reaches an average compression rate higher than 35% for the evaluated video sequences, with computational cost of a single VLC. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2014 | Four-step algorithm for early termination in HEVC inter-frame prediction based on decision treesabstractThe flexible encoding structures of High Efficiency Video Coding (HEVC) are the main responsible for the improvements of the standard in terms of compression efficiency in comparison to its predecessors. However, the flexibility provided by these structures is accompanied by high levels of computational complexity, since more options are considered in a Rate-Distortion (R-D) optimization scheme. In this paper, we propose a four-step early-termination method, which decides whether the inter mode decision should be halted without testing all possibilities. The method employs a set of decision trees, which are trained offline once, using information from unconstrained HEVC encoding runs. The resulting trees present a mode decision accuracy ranging from 97.6% to 99.4% with a negligible computational overhead. The method is capable of achieving an average computational complexity decrease of 49% at the cost of a very small Bjontegaard Delta (BD)-rate increase (0.58%). Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
VCIP | 3 |
| 2014 | Sample adaptive offset filter hardware design for HEVC encoderabstractThis work presents a hardware design for the Sample Adaptive Offset filter, which is an innovation brought by the new video coding standard HEVC. The architectures focus on the encoder side and include both classification methods used in SAO, the Band Offset and Edge Offset, and also the statistical calculations for the offset generation. The proposed architectures feature two sample buffers, classification units for both SAO types and the statistical collection unit. The architectures were described in VHDL and synthesized to an Altera Stratix V FPGA. The synthesis results show that the proposed architectures achieve 364MHz and are capable to process 44 QFHD (3840×2160) frames per second using 8,040 ALUTs of the target device hardware resources. Fabiane Rediess, Ruhan A. Conceição, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
VCIP | 5 |
| 2014 | A complexity reduction algorithm for depth maps intra prediction on the 3D-HEVCabstractThis paper proposes a complexity reduction algorithm for the depth maps intra prediction of the emerging 3D High Efficiency Video Coding standard (3D-HEVC). The 3D-HEVC introduces a new set of specific tools for the depth map coding that includes four Depth Modeling Modes (DMM) and these new features have inserted extra effort on the intra prediction. This extra effort is undesired and contributes to increasing the power consumption, which is a huge problem especially for embedded-systems. For this reason, this paper proposes a complexity reduction algorithm for the DMM 1, called Gradient-Based Mode One Filter (GMOF). This algorithm applies a filter to the borders of the encoded block and determines the best positions to evaluate the DMM 1, reducing the computational effort of DMM 1 process. Experimental analysis showed that GMOF is capable to achieve, in average, a complexity reduction of 9.8% on depth maps prediction, when evaluating under Common Test Conditions (CTC), with minor impacts on the quality of the synthesized views. Gustavo Sanchez, Mário Saldanha, Gabriel Balota, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
VCIP | 6 |
| 2014 | Complexity reduction of depth intra coding for 3D video extension of HEVCabstractThree dimensional (3D) video technology, systems and applications such as 3D television and freeviewpoint television (FTV) broadcasts require efficient encoding of video information. To fill that need a 3D video extension of High Efficiency Video Coding standard, called 3D-HEVC, is being developed. This extension is based on Multiview Video plus Depth (MVD) format, which associates a depth information to each texture frame of each view. This paper presents a method to accelerate the intra coding of these depth maps to reduce the 3D-HEVC computational complexity. The proposed algorithm exploits the edge orientation of the depth blocks to reduce the number of modes to be evaluated in the intra mode decision. In addition, the correlation between the Planar mode choice and the most probable modes (MPMs) selected is also exploited, to accelerate the depth intra coding. The experimental results show that the proposed algorithm achieves an average complexity reduction of 15% on depth information encoding, with a small degradation in encoding efficiency (BD-rate increase of 0.16% on average). Thaísa Leal da Silva, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
VCIP | 2 |
| 2013 | Energy-efficient memory hierarchy for motion and disparity estimation in multiview video codingabstractThis work presents an energy-efficient memory hierarchy for Motion and Disparity Estimation on Multiview Video Coding employing a Reference Frames-Centered Data Reuse (RCDR) scheme. In RCDR the reference search window becomes the center of the motion/disparity estimation processing flow and calls for processing all blocks requesting its data. By doing so, RCDR avoids multiple search window retransmissions leading to reduced number of external memory accesses, thus memory energy reduction. To deal with out-of-order processing and further reduce external memory traffic, a statistics-based partial results compressor is developed. The on-chip video memory energy is reduced by employing a statistical power gating scheme and candidate blocks reordering. Experimental results show that our reference-centered memory hierarchy outperforms the state-of-the-art [7][13] by providing reduction of up to 71% for external memory energy, 88% on-chip memory static energy, and 65% on-chip memory dynamic energy. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel |
DATE | 4 |
| 2013 | Simplified HEVC FME Interpolation Unit Targeting a Low Cost and High Throughput Hardware DesignabstractSummary form only given. The new demands for high resolution digital video applications are pushing the development of new techniques in the video coding area. This paper presents a simplified version of the original Fractional Motion Estimation (FME) algorithm defined by the HEVC emerging video coding standard targeting a low cost and high throughput hardware design. Based on evaluations using the HEVC Model (HM), the HEVC reference software, a simplification strategy was defined to be used in the hardware design, drastically reducing the HEVC complexity, but with some losses in terms of compression rates and quality. The used strategy considered the use of only the most used PU size in the Motion Estimation process, avoiding the evaluation of the 24 PU sizes defined in the HEVC and avoiding also the RDO decision process. This expressively reduces the ME complexity and causes a bit-rate loss lower than 13.18% and a quality loss lower than 0.45dB. Even with the proposed simplification, the proposed solution is fully compliant with the current version of the HEVC standard. The FME interpolation was also simplified targeting the hardware design through some algebraic manipulations, converting multiplications in shift-adds and sharing sub-expressions. The simplified FME interpolator was designed in hardware and the results showed a low use of hardware resources and a processing rate high enough to process QFHD videos (3840x2160 pixels) in real time. Vladimir Afonso, Henrique Maich, Luciano Volcan Agostini, Denis Franco |
DCC | 3 |
| 2013 | Coding Tree Depth Estimation for Complexity Reduction of HEVCabstractThe emerging HEVC standard introduces a number of tools which increase compression efficiency in comparison to its predecessors at the cost of greater computational complexity. This paper proposes a complexity control method for HEVC encoders based on dynamic adjustment of the newly proposed coding tree structures. The method improves a previous solution by adopting a strategy that takes into consideration both spatial and temporal correlation in order to decide the maximum coding tree depth allowed for each coding tree block. Complexity control capability is increased in comparison to a previous work, while compression losses are decreased by 70%. Experimental results show that the encoder computational complexity can be downscaled to 60% with an average bit rate increase around 1.3% and a PSNR decrease under 0.07 dB. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
DCC | 3 |
| 2013 | An adaptive workload management scheme for HEVC encodingabstractManaging the complexity of the emerging HEVC standard is a matter of academic and industrial research since its earlier versions. The sophisticated and computation-intensive tools involved in the encoding process must be leveraged if real-time applications are considered. In this paper, we propose a workload management scheme for dynamically controlling the computational complexity of HEVC, under user-defined operation frequency and target FPS. Our scheme receives these two parameters as input and aims to meet the target FPS by adjusting different encoding parameters during execution time. Experiments demonstrate that our scheme successfully meets the target FPS while introducing negligible rate-distortion losses. A comparison with state-of-the-art shows that our scheme is capable of achieving a time reduction of up to 43% for Full HD sequences, with a maximum loss of 0.03 dB in Y-PSNR and a 3.5% increase in bitrate. Mateus Grellert, Muhammad Shafique 0001, Muhammad Usman Karim Khan, Luciano Volcan Agostini, Júlio C. B. de Mattos, Jörg Henkel |
ICIP | 4 |
| 2013 | Content-adaptive reference frame compression based on intra-frame prediction for multiview video codingabstractThis paper presents a content-adaptive reference frame compression scheme to alleviate the large overhead of external memory communication during the Motion and Disparity Estimation process in Multiview Video Coding (MVC). Our scheme is based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples. The intra-prediction residue is compressed by a path composed of non-linear quantization and Huffman-based entropy encoder. Four different quantization strengths and Huffman tables were statistically defined. They are dynamically selected according to a content adaptation strategy, which classifies the original blocks based on their spatial homogeneity. Experimental results show that the proposed content-adaptive compression scheme is able to reduce the external memory accesses by up to 63% along with negligible losses in the MVC encoder rate-distortion performance. Compared to the best available related work [12] our content-adaptive reference frame compression achieves 39% reduced external memory accesses, while still providing a BD-PSNR increase of 0.03dB. Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Jörg Henkel, Sergio Bampi |
ICIP | 4 |
| 2013 | A hardware friedly motion estimation algorithm for the emergent HEVC standard and its low power hardware designabstractIn this paper the hardware friendly Multi Point Diamond Search (MPDS) motion estimation (ME) algorithm is evaluated in the High Efficiency Video Coding Standard (HEVC). The MPDS algorithm was implemented in HEVC reference software and its efficiency is compared with the standard HEVC fast algorithm, the Enhanced Predictive Zonal Search (EPZS). The evaluation result, in average, shows loses of only 1.7% in the compression rate and 0.05% in PSNR. However, the main advantage of the MPDS algorithm is its hardware friendly aspect and this paper also presents its hardware design focused on real time processing HD 1080p videos. The designed architecture is capable to process HD 1080p videos in real time when synthesized for TSCM 90nm technology. The MPDS algorithm has obtained a good tradeoff among quality, compression rate and costs for its hardware implementation. Gustavo Sanchez, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 3 |
| 2013 | A lossless approach for external memory bandwidth reduction in video coding systems and its VLSI architectureabstractThis paper presents the Reference Frame Context Adaptive Variable-Length Coder (RFCAVLC), which is a lossless solution to external memory bandwidth reduction in current video coding systems. The proposed approach is based on an adaptation of the traditional Huffman algorithm, and it uses eight static tables to avoid the cost of the on-the-fly statistical analysis. The best table to encode a block is defined using a context evaluation, resulting in a context-adaptive configuration. The use of RFCAVLC reached an average compression rate higher than 31% for the evaluated video sequences. The architectures that implement the RFCAVLC encoder and decoder were designed and synthesized to an FPGA device. The RFCAVLC design is able to reach real-time encoding for WQSXGA (3200 × 2048 pixels) at 30 fps. The synthesis results show that this solution can be easily coupled to a complete video encoder system with negligible hardware overhead and without compromising the throughput for real-time high-definition multimedia applications. Dieison Silveira, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICME | 3 |
| 2013 | Constrained encoding structures for computational complexity scalability in HEVCabstractThe High Efficiency Video Coding standard shows improved compression efficiency in comparison to previous standards at the cost of higher computational complexity. In this paper, a complexity scalability method for HEVC encoders based on the dynamic adjustment of the number of constrained coding treeblocks is proposed. The method limits the Prediction Unit shapes and the maximum tree depth used in each Coding Treeblock in order to decrease the number of evaluations performed in the Rate-Distortion Optimization process. The encoder is capable of trading off computational complexity and compression efficiency while still maintaining the encoding time per Group of Pictures (GOP) under a pre-defined target. The encoding complexity can be decreased in up to 60% when compared to the original encoder at the cost of small Bjontegaard Delta (BD)-rate increases. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
PCS | 4 |
| 2013 | Fast HEVC intra mode decision algorithm based on new evaluation order in the Coding Tree BlockabstractThis paper presents a fast mode decision algorithm for the HEVC intra prediction. A new evaluation order in the Coding Tree Block (CTB) allows the use of modes from low level PUs to be used as reference to the current PU decision. In this paper we use this idea to develop a fast intra mode decision algorithm that can be configured to run in two different complexity modes, relaxed and aggressive. Experimental results have shown that our algorithm achieved encoding time savings of almost 60% with negligible loss in the compression efficiency when compared to the full RDO based decision. Besides, our mode decision algorithm presented the best result in terms of time saving per compression efficiency when compared with all related works. Daniel Palomino 0001, Eduardo Cavichioli, Altamiro Amadeu Susin, Luciano Volcan Agostini, Muhammad Shafique 0001, Jörg Henkel |
PCS | 4 |
| 2013 | HEVC intra mode decision acceleration based on tree depth levels relationshipabstractThe new High Efficiency Video Coding (HEVC) standard is achieving higher encoding efficiency when compared to its predecessors such as H.264/AVC. One of the factors responsible to this improvement is the intra prediction method, which introduces a larger number of prediction directions resulting in an enhanced rate-distortion (RD) performance at the cost of a higher computational complexity. This paper proposes an algorithm to accelerate the intra mode decision, reducing the complexity of intra coding. The acceleration procedure takes into account the edge direction information and explores the correlation of intra modes across levels of the HEVC hierarchical tree structure. Experimental results show that the proposed algorithm provides a decrease of up to 40.82% in the HEVC intra prediction processing time, with a small degradation in encoding efficiency (BD-PSNR loss of 0.1 dB on average). Thaísa Leal da Silva, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
PCS | 3 |
| 2013 | Iterative random search: a new local minima resistant algorithm for motion estimation in high-definition videos
Marcelo Schiavon Porto, Cassio Cristani, Pargles Dall'Oglio, Mateus Grellert, Júlio C. B. de Mattos, Sergio Bampi, Luciano Volcan Agostini |
Multim. Tools Appl. | 7 |
| 2012 | Motion compensated tree depth limitation for complexity control of HEVC encodingabstractThe recently introduced quadtree coding structures used in HEVC increase compression efficiency in comparison to previous standards at the cost of higher computational complexity levels. This paper proposes an evolution of a complexity control method for HEVC encoders based on the dynamic adjustment of these structures' maximum depth. The new method improves the previous solution by adopting a new control strategy and compensating the motion effect on a maximum tree depth map which is central to the complexity control strategy. The proposed method is capable of performing a more accurate complexity control than our previous strategy while still reducing compression efficiency losses in terms of image quality and bit rate. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
ICIP | 3 |
| 2012 | A memory aware and multiplierless VLSI architecture for the complete Intra Prediction of the HEVC emerging standardabstractThis work proposes a hardware architecture for the Intra Frame Prediction of the emerging High Efficiency Video Coding (HEVC) standard. The architecture was designed considering all innovative features of the Intra Prediction included in the HEVC, i.e. all modes and all Prediction Units (PU) sizes. Performance and memory accesses are a problem in the HEVC intra prediction and hardware architecture designs are good alternative to solve these issues, especially when energy-efficient solutions are targeted. Buffers and internal memories were used in the designed architecture to decrease the number of external memory accesses. Two independent data paths processing eight samples in parallel and a deep and multiplierless pipeline were designed to increase the throughput. The architecture was synthesized using an IBM 65nm CMOS technology. The results have shown that the architecture is able to process 30 HD720p frames per second and 13 HD1080p frames per second when running at 500 MHz, reducing in 95% the accesses to the external memory. Daniel Palomino 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin |
ICIP | 3 |
| 2012 | High performance hardware architectures for the inverse Rotational Transform of the emerging HEVC standardabstractThis paper presents a dedicated hardware architecture for the Rotational Transform (ROT), which is one of the novel tools proposed for the emergent HEVC video coding standard. The main goal of this coding tool is to achieve higher energy compaction of the main transform coefficient matrix, minimizing the quantization error and improving the efficiency of the entropy encoding. Five versions of this architecture were implemented, using either a fully combinational structure or a pipeline with nine stages. The designed architectures were described in VHDL and synthesized for an Altera Stratix III FPGA. The synthesis results show that all versions can process very high resolution videos, such as QFHD, in real time. The version with the highest processing rate achieved a maximum operation frequency of 260.15 MHz. This architecture reaches a processing rate of 2.08 billion samples per second, allowing it to process UHDTV videos in real time. Henrique Avila Vianna, Gustavo Sanchez, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 4 |
| 2012 | Motion Vectors Merging: Low Complexity Prediction Unit Decision Heuristic for the Inter-prediction of HEVC EncodersabstractThis paper presents the Motion Vectors Merging (MVM) heuristic, which is a method to reduce the HEVC inter-prediction complexity targeting the PU partition size decision. In the HM test model of the emerging HEVC standard, computational complexity is mostly concentrated in the inter-frame prediction step (up to 96% of the total encoder execution time, considering common test conditions). The goal of this work is to avoid several Motion Estimation (ME) calls during the PU inter-prediction decision in order to reduce the execution time in the overall encoding process. The MVM algorithm is based on merging NxN PU partitions in order to compose larger ones. After the best PU partition is decided, ME is called to produce the best possible rate-distortion results for the selected partitions. The proposed method was implemented in the HM test model version 3.4 and provides an execution time reduction of up to 34% with insignificant rate-distortion losses (0.08 dB drop and 1.9% bitrate increase in the worst case). Besides, there is no related work in the literature that proposes PU-level decision optimizations. When compared with works that target CU-level fast decision methods, the MVM shows itself competitive, achieving results as good as those works. Felipe Sampaio, Sergio Bampi, Mateus Grellert, Luciano Volcan Agostini, Júlio C. B. de Mattos |
ICME | 4 |
| 2012 | Spread and Iterative Search: A High Quality Motion Estimation Algorithm for High Definition Videos and Its VLSI DesignabstractThis paper presents the Spread and Iterative Search (S&IS) motion estimation algorithm, which uses a random spread evaluation together with a central iterative evaluation to avoid local minima falls and to increase the image quality for high definition videos. Considering Full HD videos, S&IS reached an average PSNR gain of 1.41dB when compared to Diamond Search (DS), with an increase of about four times in the number of evaluated blocks. When compared to Full Search (FS), the S&IS achieved an average PSNR loss of 1.56 dB, evaluating 73 times less blocks than FS. An efficient architecture for the S&IS algorithm is also presented in this paper. The architecture was designed targeting in real time processing (30 frames per seconds) for QFHD videos (3840×2160 pixels). The architecture was described in VHDL and synthesized for and Altera Stratix 4 FPGA and for ST90nm standard cells technology. Booth syntheses show that the architecture is able to process QFHD frames in real time. The standard cells version is able to reach also a good trade-off among area, memory and power consumption, processing QFHD videos with 62.2 mW. Gustavo Sanchez, Luciano Volcan Agostini, Felipe Sampaio, Marcelo Schiavon Porto, Sergio Bampi |
ICME | 2 |
| 2012 | Adaptive coding tree for complexity control of high efficiency video encodersabstractThe emerging HEVC standard introduces several techniques which increase compression efficiency in comparison to its predecessors. However, such advances are accompanied by increases in computational complexity, limiting the encoder use in computational or power-constrained devices. This paper proposes a novel complexity control method for the future HEVC encoders based on a dynamic adjustment of the newly proposed coding tree structures. The relationship between coding tree depths and the encoding complexity is explored to selectively constrain encoding possibilities in order to not exceed a predefined complexity target. Experimental results show that the encoder computational complexity can be downscaled to 60% with a bit rate increase under 3.5% and a PSNR decrease under 0.1 dB. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luciano Volcan Agostini |
PCS | 4 |
| 2012 | Evaluating two implementations of the component responsible for decoding video and audio in the Brazilian digital TV middleware
Tiago Henrique Trojahn, Juliano Lucas Gonçalves, Júlio C. B. de Mattos, Luciano Volcan Agostini, Leomar S. da Rosa Jr. |
Multim. Tools Appl. | 4 |
| 2012 | Performance and Computational Complexity Assessment of High-Efficiency Video EncodersabstractThis paper presents a performance evaluation study of coding efficiency versus computational complexity for the forthcoming High Efficiency Video Coding (HEVC) standard. A thorough experimental investigation was carried out to identify the tools that most affect the encoding efficiency and computational complexity of the HEVC encoder. A set of 16 different encoding configurations was created to investigate the impact of each tool, varying the encoding parameter set and comparing the results with a baseline encoder. This paper shows that, even though the computational complexity increases monotonically from the baseline to the most complex configuration, the encoding efficiency saturates at some point. Moreover, the results of this paper provide relevant information for implementation of complexity-constrained encoders by taking into account the tradeoff between complexity and coding efficiency. It is shown that low-complexity encoding configurations, defined by careful selection of coding tools, achieve coding efficiency comparable to that of high-complexity configurations. Guilherme Corrêa 0001, Pedro A. Amado Assunção, Luciano Volcan Agostini, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Run-time adaptive energy-aware motion and disparity estimation in multiview video codingabstractThis paper presents a novel run-time adaptive energy-aware Motion and Disparity Estimation (ME, DE) architecture for Multiview Video Coding (MVC). It incorporates efficient memory access and data prefetching techniques for jointly reducing the on/off-chip memory energy consumption. A dynamically expanding search window is constructed at run time to reduce the off-chip memory accesses. Considering the multi-stage processing nature of advanced fast ME/DE schemes, a reduced-sized multi-bank on-chip memory is employed which can be power-gated depending upon the video properties. As a result, when tested for various video sequence, our approach provides a dynamic energy reduction of 82--96% for the off-chip memory and a leakage energy reduction of 57--75% for the on-chip memory compared to the Level-C and Level-C+ [7] prefetching techniques (which are the prominent data reuse and prefetching techniques in ME for video coding). The proposed ME/DE architecture is synthesized using a 65nm IBM low power technology. Compared to state-of-the-art MVC ME/DE hardware [14], our architecture provides 66% and 72% reduction in the area and power consumption, respectively. Moreover, our scheme achieves 30fps ME/DE 4-view HD1080p encoding with a power consumption of 74mW. Bruno Zatt, Muhammad Shafique 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel |
DAC | 4 |
| 2011 | SHBS: A heuristic for fast inter mode decision of H.264/AVC standard targeting VLSI designabstractIn the Rate-Distortion Optimization technique for H.264/AVC, the process of choosing the best mode is performed through exhaustive executions of the whole encoding process, which increases significantly the encoder complexity, sometimes even forbidding its use in real time video coding applications. In order to reduce the number of calculations necessary to determine the best inter-frame mode, this work proposes the SHBS (Stationarity, Heterogeneity and Border Strength) heuristic. The use of SHBS causes a reduction of 168 times in the encoding iterations, with a better PSNR, at the cost of a relatively small bit-rate increase. The SHBS heuristic was designed in hardware targeting FPGAs and this architecture achieved an operation frequency of 118 MHz, being able to process up to 438 HD 1080p frames per second. Guilherme Corrêa 0001, Daniel Palomino 0001, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
ICME | 4 |
| 2011 | A high throughput H.264/AVC intra-frame encoding loop architecture for HD1080pabstractIn this work we present a high throughput hardware architecture for the H.264/AVC intra-frame encoder exploiting the parallelism of intra prediction, forward and inverse transforms and quantization. Since there is a strong data dependency between the intra prediction and the image reconstruction loop, the latency of this path is a key design issue in order to provide high performance coding. Considering that 77% of the total intra-encoding computation is spent in these modules, our architecture handles a 4-pixel wide intra prediction module and a 16-pixel wide reconstruction loop. Compared to the state-of-the-art our approach reduces by 47% the number of cycles to process a macroblock. Running at 150 MHz our architecture guarantees encoding of 61 HD1080p frames per second. The developed architecture requires 73.4 MHz to real-time encode HD1080p, which is a 46% reduction of the frequency requirement compared to the state-of-the-art. Cláudio Machado Diniz, Bruno Zatt, Cristiano Thiele, Altamiro Amadeu Susin, Sergio Bampi, Felipe Sampaio, Daniel Palomino 0001, Luciano Volcan Agostini |
ISCAS | 8 |
| 2011 | A multilevel data reuse scheme for Motion Estimation and its VLSI designabstractMotion Estimation (ME) in video coding is a vital component that excels not only in computational complexity, but off-chip memory bandwidth as well. These two issues are considered critical constraints in terms of High Definition (HD) video coding, since a large volume of data must be processed. The multilevel data reuse scheme proposed in this paper is able to reduce the off-chip memory bandwidth, with direct impact in throughput and energy consumption. This scheme explores the concept of overlapped Search Windows (SW) in more than one level and poses no harm to video quality. Comparisons with related works show that this solution provides the best tradeoff between the use of on-chip memory and reduction of the off-chip memory bandwidth. The data reuse scheme was applied in a ME architecture and the synthesis results show that this solution presented the lowest use of hardware resources and the highest operation frequency among related works. The proposed architecture is able to process 1080p videos at 25 fps, and the reduction ratio of off-chip memory access achieved by the architecture is greater than 95% when compared to the traditional method. Mateus Grellert, Felipe Sampaio, Júlio C. B. de Mattos, Luciano Volcan Agostini |
ISCAS | 4 |
| 2010 | Timing and interface communication analysis of H.264/AVC encoder using SystemC modelabstractThis work presents a detailed timing and communication analysis for an H.264/AVC video encoder architecture using a SystemC model. The model was described using different abstraction levels in order to evaluate specific characteristics of each component module. The target encoder is defined to be able for H.264/AVC real-time encoding for 1080p video sequences at 30 fps and was modeled as a two-stage macro-pipeline system composed by eight component modules: Macroblock buffer, Intra- and Inter-Frame Predictors, Mode Decision, Forward and Inverse Transforms and Quantization, Reference Memory Write and Entropy Encoder (CAVLC). The bandwidth of each internal connection and of external memory interface was evaluated. The timing behavior and the data dependencies were characterized and summarized in a timing diagram in order to define design constraints and provide an accurate system specification when compared to a H.264/AVC encoder in the literature. Bruno Zatt, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
VLSI-SoC | 3 |
| 2009 | Low latency and high throughput dedicated loop of transforms and quantization focusing in the H.264/AVC Intra PredictionabstractThis paper presents an efficient architectural design for a dedicated transforms and quantization loop. This design targeted the Intra Prediction of the H.264/AVC standard. The architecture was designed intending to achieve the best possible relation between throughput, latency and hardware resources consumption. The latency and throughput of this loop are extremely important to define the intra prediction performance. The use of hardware was reduced through the reuse of the same datapath for different calculations. The architecture was synthesized to Altera Stratix III FPGA and to the TSMC 0.18 ¿m standard-cells technology. The architecture, when mapped to standard-cells, reaches a processing rate of 114 HDTV frames per second, attending the intra prediction restrictions. Daniel Palomino 0001, Felipe Sampaio, Robson Dornelles, Luciano Volcan Agostini |
ICIP | 4 |
| 2009 | A real time H.264/AVC intra frame prediction hardware architecture for HDTV 1080P videoabstractThis work presents an intra frame prediction hardware architecture for H.264/AVC baseline/main profile encoder which performs real time processing of HDTV 1080p videos. It is achieved by exploring the parallelism of intra prediction and by reducing the latency for Intra 4times4 processing, which is the intra encoding bottleneck. Synthesis results on Xilinx Virtex-II Pro FPGA and TSMC 0.18 mum standard-cells indicate that this architecture is able to real time encode HDTV 1080p video operating at 110 MHz. Our architecture can encode HD1080p, 720p and SD video in real time at a frequency 25% lower when compared to similar works. Cláudio Machado Diniz, Bruno Zatt, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi |
ICME | 3 |
| 2008 | A high throughput and low cost diamond search architecture for HDTV motion estimationabstractThis paper presents a high throughput and low cost architecture for motion estimation using a sub-sampled diamond search algorithm (SDS). The quality of SDS was compared with full search through software implementations and the results are presented. The designed hardware considered a search area of 100times100 samples, with blocks of 16times16 pixels. The architecture was described in VHDL and mapped to a Xilinx Virtex-4 FPGA. Synthesis results indicate that SDS is able to run at 185.7 MHz, using only 3541 LUTs. This architecture can reach real time for HDTV (1920times1080 pixels) in the worst case, and it can process 120 HDTV frames per second in the average case. Marcelo Schiavon Porto, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin |
ICME | 2 |
| 2008 | HP422-MoCHA: A H.264/AVC High Profile motion compensation architecture for HDTVabstractThis work presents the HP422-MoCHA, the first published hardware architecture that implements a full compliant H.264/AVC motion compensator for high profile 4:2:2. The hardware is composed by three main modules: Motion Vector Predictor, Memory Access and Sample Interpolator. The designed architecture was described in VHDL and mapped to a Xilinx Virtex-II PRO FPGA. This architecture reaches the throughput to decode HDTV Level 4 (1080p) @ 30fps. Bruno Zatt, Altamiro Amadeu Susin, Sergio Bampi, Luciano Volcan Agostini |
ISCAS | 4 |
| 2007 | RIC Fast Adder and its Set Tolerant Implementation in FPGAsabstractFPGA is currently a very important design technology to implement electronic systems due to its high logic density, its fast time-to-market and its low cost. But in order to provide high logic density FPGA devices are fabricated with nanometer CMOS technology that is becoming susceptible to radiation-induced soft errors. Among these errors, single-event transients (SETs) are those that are induced in the user's programmable logic. This paper presents a new fast adder, called RIC (Re-computing the Inverse Carry-in) and shows how this new adder architecture may be used to build SET-tolerant fast adders. Results considering FPGA-based implementation are presented. Eduardo Mesquita, Helen Franck, Luciano Volcan Agostini, José Luís Güntzel |
FPL | 3 |
| 2007 | MoCHA: a Bi-Predictive Motion Compensation Hardware for H.264/AVC Decoder Targeting HDTVabstractThis paper presents the MoCHA (motion compensation hardware architecture) design. MoCHA is an architectural design for bi-predictive motion compensation of the H.264/AVC decoder. The designed architecture features a memory hierarchy to reduce the memory bandwidth and the number of memory access cycles. The architecture uses a single datapath to process bi-predictive reference areas and it processes luma and chroma samples in parallel. The design was mapped to a Xilinx Virtex II Pro FPGA and it is able to run at 100MHz. The throughput is enough to support more than 30 bi-predictive HDTV frames per second. Arnaldo Azevedo, Bruno Zatt, Luciano Volcan Agostini, Sergio Bampi |
ISCAS | 3 |
| 2007 | High Throughput Hardware Architecture for Motion Estimation with 4: 1 Pel Subsampling Targeting Digital Television Applications
Marcelo Schiavon Porto, Luciano Volcan Agostini, Leandro Rosa, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 2 |
| 2007 | A Pipelined 8x8 2-D Forward DCT Hardware Architecture for H.264/AVC High Profile Encoder
Thaísa Leal da Silva, Cláudio Machado Diniz, João Alberto Vortmann, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 4 |
| 2007 | Motion Compensation Hardware Accelerator Architecture for H.264/AVC
Bruno Zatt, Valter Ferreira, Luciano Volcan Agostini, Flávio Rech Wagner, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 3 |
| 2006 | FPGA Design of A H.264/AVC Main Profile Decoder for HDTVabstractThis paper presents the architecture, design, validation, and prototyping of inverse transforms and quantization, intra prediction, motion compensation and loop filter, for a main profile H.264/AVC decoder. These architectures were designed to reach high throughputs and to be easily integrated with the other H.264/AVC modules. The architectures, all fully H.264/AVC compliant, were completely described in VHDL and further validated through simulations down to prototyping. The architectures were prototyped using a Digilent XUP V2P board, containing a Virtex-II Pro XC2VP30 Xilinx FPGA. The post place-and-route synthesis results indicate that the designed architectures are able to process 114 million of samples per second and, in the worst case, they are able to process 64 HDTV frames (1080×1920) per second, allowing their use in H.264/AVC decoders targeting real time HDTV applications. Luciano Volcan Agostini, Arnaldo Azevedo, Vagner Santos Da Rosa, Eduardo A. Berriel, Tatiana Gadelha Serra dos Santos, Sergio Bampi, Altamiro Amadeu Susin |
FPL | 1 |
| 2006 | FPGA Based Architectures for H. 264/AVC Video Compression StandardabstractThe H.264/AVC (as known as MPEG-4 part 10) [1, 2] is a video coding standard that has been developed to achieve significant improvements, in the compression performance, over the existing standards. The main blocks of a H.264/AVC encoder are the motion estimation, the motion compensation, the intra prediction, the loop filter, the entropy coder, the forward and inverse quantization and the forward and inverse transforms. The H.264/AVC decoder is formed by entropy decoder, motion compensation, intra prediction, loop filter, inverse quantization and inverse transforms [1]. This work focuses on the design of high performance architectures for the H.264/AVC standard. Luciano Volcan Agostini, Sergio Bampi |
FPL | 1 |
| 2006 | High throughput architecture for H.264/AVC forward transforms blockabstractThis paper presents a high throughput hardware for the complete H.264/AVC forward transforms block. There are three different transform inside this block and the presented architecture synchronizes these transforms, generating a constant processing rate in its outputs. This is an important characteristic of this architecture that was designed to be easily integrated to the other H.264/AVC blocks. The architecture does not use memory bits and the transforms in two dimensions are calculated directly, without the use of the separability property. The architecture was described in VHDL and was validated and prototyped using a Xilinx Virtex II Pro FPGA. The synthesis was directed to a VP30 FPGA and to a TSMC 0.35μm standard-cell technology. The throughputs of the T block architecture for these two different technologies reaches a processing rate higher than 120 million of samples per second, allowing its use in H.264/AVC codecs directed to HDTV. Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, Sergio Bampi, Leandro Rosa, José Luís Güntzel, Ivan Saraiva Silva |
ACM Great Lakes Symposium on VLSI | 1 |
| 2006 | High throughput multitransform and multiparallelism IP for H.264/AVC video compression standardabstractThis paper presents the design of a high throughput multitransform and multiparallelism IP for H.264/AVC standard. This solution supports the five H.264/AVC transforms and it supports five different levels of parallelism. The proposed architecture were described in VHDL and synthesized to Altera Stratix and Xilinx Virtex-II Pro FPGAs and to TSMC 0.35/spl mu/m standard cells. The multitransform and multiparallelism architecture mapped to FPGAs could process from 124 millions to 3.2 billions of samples per second, depending on the parallelism level selected. The standard cells version could process from 218.7 millions to 3.5 billions of samples per second. These results indicate that the proposed solution presents a high flexibility and that this solution is able to be used in various H.264/AVC codecs with different performance requirements. The performance results of all experiments realized indicated that this architecture is able to be used in high definition applications, like HDTV. Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, José Luís Güntzel, Ivan Saraiva Silva, Sergio Bampi |
ISCAS | 1 |
| 2006 | Motion Compensation Decoder Architecture for H.264/AVC Main Profile Targeting HDTVabstractThis work presents the design, the validation and the prototyping of a motion compensation architecture for a H.264/AVC video decoder. The designed architecture supports the main profile level 4.0 and it targets high resolution applications, like HDTV. This design considers the sample processing of the motion compensation block, which includes quarter-pel interpolation, weighted prediction, average to bi-predictive processing and clipping. The architecture processes luma and chroma samples in parallel, with independent luma and chroma datapaths. The design uses a single interpolator to process bi-predictive macroblocks. The design was synthesized to FPGA and standard cell technologies. The synthesis results had indicated that this architecture reaches 100 MHz in both technologies, allowing real time to decode HDTV videos with 1920times1080 pixels. The prototype was targeted to a Xilinx Virtex-II PRO FPGA Arnaldo Azevedo, Bruno Zatt, Luciano Volcan Agostini, Sergio Bampi |
VLSI-SoC | 3 |
| 2005 | A FPGA Based Design of a Multiplierless and Fully Pipelined JPEG CompressorabstractThis paper presents the design and implementation of a multiplierless JPEG compressor for gray scale images. The modules of this architecture were fully pipelined and targeted to FPGA device implementation. The designed architectures are detailed in this paper and they were described in VHDL, simulated and physically mapped to Altera Flex10KE FPGAs. The JPEG compressor pipeline has a minimum latency of 238 clock cycles, given the full modular pipeline depth. The minimum compressor period is 26.6ns and the compressor is able to process 37.6 millions of pixels per second. For example, the compressor can process a 640x480 pixels still image in 8.2 ms, reaching a maximum processing rate of 122.4 frames per second. Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, Sergio Bampi, Ivan Saraiva Silva |
DSD | 1 |
| 2004 | Project Space Exploration on the 2-D DCT Architecture of a JPEG Compressor Directed to FPGA ImplementationabstractThis paper presents a project space exploration on the baseline JPEG compressor proposed and implemented in previous works. This exploration took as basis the substitution of the operators used in the 2-D DCT calculation architecture of the compressor and the consequent evaluation of impact in terms of performance and resources utilization. This substitution was made with main focus in the carry lookahead, hierarchical carry lookahead and carry select architectures, with the objective to increase the JPEG compressor performance. As the compressor architecture was designed in an hierarchical mode the operators substitution was an activity quite simple, because it has not involved the other hierarchy levels. The operators were described in VHDL, synthesized and validated. They were inserted in the 2-D DCT architecture for synthesis in the whole module. The 2-D DCT was synthesized for an altera FPGA. With this project space exploration, the highest performance obtained for the 2-D DCT was 23% higher than the original, using 11% more logic cells. Roger Endrigo Carvalho Porto, Luciano Volcan Agostini |
DATE | 2 |