VLDB 2026 Research / reviewers in the wild / expert
Marcelo Schiavon Porto
dblp:44/1063
· DBLP profile ↗
65ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0003-3827-3023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 3 first-author · 10 since 2021Systems, architecture and hardware · 26 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardware-Friendly Machine-Learning-Based Fast AV1 Overlapped Block Motion Compensation
William Kolodziejski, Leonardo Braga, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 3 |
| 2026 | Energy-Efficient Neural Video Coding via a High-Throughput Hardware Design for Pointwise Convolution and LUT-based WSiLU
Denis Maass, Vanessa Aldrighi, Ruhan A. Conceição, Wen-Hsiao Peng, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 6 |
| 2026 | DM-FIFS: A Dual-Model Machine-Learning Method for Fast Interpolation Filter Search in AV1 EncodingabstractAV1 is a video codec developed by leading technology companies to meet the increasing demands of modern video applications. Fractional Motion Estimation (FME), the focus of this work, is an important AV1 encoder tool. FME employs interpolation filters to generate sub-pixel predictions, thereby improving motion estimation accuracy. In AV1, FME uses sophisticated interpolation filters that can be combined in horizontal and vertical directions, with the optimal filter pair selected by the Interpolation Filter Search (IFS) process. The paper presents DM-FIFS, a dual-model, machine-learning-based approach designed to overcome prior limitations in filter prediction accuracy, which often led to suboptimal trade-offs between computational effort and coding efficiency. By splitting the decision space into two specialized models, DM-FIFS achieves more accurate filter predictions, thereby improving the balance between gains in computational effort and losses in coding efficiency compared to single-model approaches. The paper also presents a set of assessment and ablation experiments, a comprehensive discussion of key innovations in the AV1 encoder, and a detailed analysis of the interpolation filters used in AV1 FME. Experimental results show that DM-FIFS reduces IFS execution time by 51.40% with only a 0.11% increase in BD-BR, demonstrating a superior trade-off between computational effort and coding efficiency. To the best of our knowledge, DM-FIFS represents the most advanced machine-learning-based solution to reduce the computational complexity of AV1 IFS reported to date. William Kolodziejski, Leonardo Braga, Marcelo Rezende, Marcelo Schiavon Porto, Luciano Volcan Agostini |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Cross-Platform Neural Video Coding: A Case StudyabstractIn this paper, we first show that current learning-based video codecs, specifically the SSF codec, are not suitable for real-world applications due to the mismatch between the encoder and decoder caused by floating-point round-off errors. To address this issue, we propose the static quantization of the hyper prior decoding path. The quantization parameters are determined through an exhaustive search of all possible combinations of observers and quantization schemes from PyTorch. For the SSF codec, when encoding and decoding on different machines, the proposed solution effectively mitigates the mismatch issue and enhances compression efficiency results by preventing severe image quality degradation. When encoding and decoding are performed on the same machine, it constrains the average BD-rate increase to 9.93% and 9.02% for UVG and HEVC-B sequences, respectively. Ruhan A. Conceição, Marcelo Schiavon Porto, Wen-Hsiao Peng, Luciano Volcan Agostini |
ISCAS | 2 |
| 2025 | A Power-Efficient Architecture for LES Solving and ∆MV Calculation of VVC Affine PredictionabstractThe increasing demand for video content has created a need for more efficient video compression techniques and the Versatile Video Coding (VVC) standard introduces several techniques to reach this goal. One key innovation in VVC is affine prediction, which gives more flexibility in inter-frame prediction by using multiple motion vectors for motion representation. However, its processing demands significant computational effort, boosting the use of dedicated hardware accelerators to enable real-time processing, mainly when focusing on mobile devices. This work introduces a pipelined and parallelized hardware design for the Linear Equation System (LES) solving and the ∆MV calculation, both essential steps in the search for affine motion vectors. The proposed architecture can generate new ∆MV values every three clock cycles and can process UHD 4K@60fps videos dissipating 47.32mW, with a negligible coding efficiency impact of 0.006%. These results outperform all related works in the literature, achieving twice the throughput, 84.2% lower power dissipation, and five times better coding efficiency. Denis Maass, Marcello M. Muñoz, Murilo R. Perleberg, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2025 | A Fast and Hardware-Friendly CTB-Based Approach for VVC Affine Motion Estimation
Denis Maass, Marcello M. Muñoz, Murilo R. Perleberg, Luciano Volcan Agostini, Marcelo Schiavon Porto |
VCIP | 5 |
| 2025 | Machine Learning-Based Block Partitioning for V-PCC Encoding AccelerationabstractPoint clouds capture spatial and attribute information about objects or environments. It has been widely used in applications like autonomous driving, augmented and virtual reality, and 3D scanning. However, the large volume of data involved in point cloud processing poses challenges in terms of storage, transmission, and real-time processing. The Video-based Point Cloud Compression (V-PCC) standard addresses these challenges by employing 2D video compression techniques to encode dynamic 3D point clouds. Despite its effectiveness, V-PCC demands a high computational cost, particularly in encoding geometry and attribute video sub-streams. This paper presents a machine learning-based approach to accelerate the V-PCC encoding, focusing on the block partitioning process. The proposed method utilizes decision tree models to predict the partitioning of Coding Tree Units (CTUs) and Prediction Units (PUs) during the encoding of geometry and attribute sub-streams. The proposed method significantly accelerates the V-PCC encoder by reducing the overall coding time by up to 60%, with an average decrease in the coding efficiency of 1.31% and 1.75% for the attribute (luma) and geometry (D2) sub-stream in Random Access configuration and negligible impact for All-Intra configuration. Gustavo Rehbein, Cristiano Santos, Guilherme Corrêa 0001, Marcelo Schiavon Porto |
VCIP | 4 |
| 2025 | Analysis of Real-Time Hardware-based HEVC Encoders on GPU and Mobile PlatformsabstractRecent hardware encoders in GPUs and mobile SoCs enable real-time high-resolution video processing by constraining the supported encoding tools available in video coding standards. These constraints are also used to meet strict power, area, and memory limits of mobile platforms. This paper presents an analysis of the hardware-based High Efficiency Video Coding (HEVC) present in the high-performance NVIDIA NVENC within the RTX 4070Ti GPU, and the power-efficient encoder present in the Snapdragon 8 Gen 2 chip within the Samsung Galaxy S23+ smartphone. The analysis is performed in two perspectives: (1) the tool set constraints employed by each implementation are identified by a bitstream analysis on UHD encoded videos, and (2) a compression-efficiency evaluation of both encoders through a rate vs. distortion and Bjontegaard-Delta Rate (BD-Rate) analysis against the HEVC Test Model reference software. The results reveal the different design trade-offs between the platforms, offering valuable insights for hardware designers by highlighting the implementation choices of major industry players like NVIDIA and Qualcomm. Allan Schuch, Daniel Palomino 0001, Marcelo Schiavon Porto |
VCIP | 4 |
| 2024 | A systematic literature review on video transcoding acceleration: challenges, solutions, and trends
Bruno Zatt, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
Multim. Tools Appl. | 3 |
| 2023 | High-Throughput and Multiplierless Hardware Design for the AV1 Local Warped MC InterpolationabstractMost of the current video codecs support only translational motion models. However, real motion is often complex and cannot be precisely estimated using only translational models. To handle complex motions like panning, zooming, scaling, shearing and rotation, AOMedia AV1 encoder counts with two tools, called Global and Local Warped Motion Compensation (LWMC). This paper presents two dedicated hardware designs for the AV1 LWMC interpolation filters. The presented hardware can process up to UHD 8K videos at 60fps. The architecture was synthesized for 40nm TSMC standard cells, requiring 454.37K gates with a power dissipation of 189.35mW. To the best of the authors’ knowledge, this is the first work in the literature targeting a dedicated hardware design for LWMC AV1 tool. Robson Domanski, William Kolodziejski, Wagner Penny, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 4 |
| 2023 | H.264-to-AV1 Video Transcoding Acceleration Based on Lightweight Machine LearningabstractVideo streaming platforms have been using the H.264/AVC standard for a long time, even though it was released almost 20 years ago and much more efficient codecs are currently available. The AOMedia Video 1 (AV1) format is an alternative with significant coding efficiency gains in comparison to H.264/AVC, besides being a royalty-free format. However, migrating legacy content from older to newer formats is a costly task, which requires long processing times. This work presents a solution for accelerating the H.264-to-AV1 transcoder based on machine learning. Sixteen decision tree models trained with data gathered during the H.264/AVC decoding and the AV1 encoding processes are proposed and implemented in the libaom reference software, leading to a complexity reduction of 18.96% at the cost of coding efficiency losses of 2.85% on average. To the best of the authors' knowledge, this is the first H.264-to-AV1 transcoding acceleration solution published in the literature. Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001 |
ISCAS | 2 |
| 2023 | High-Throughput Design for a Multi-Size DCT-II Targeting the AV1 EncoderabstractThis paper presents a dedicated multi-size hardware design for the Discrete Cosine Transform type II (DCT-II) of AV1 encoder. The DCT-II is one of four transform kernels supported by AV1; however, DCT-II is used in all configurations defined by AV1. Moreover, the 1D DCT-II can be applied for five different sizes ranging from 4-point up to 64-point. The 1D multi-size DCT-II was designed to process multiple transform sizes in parallel, always processing 64 samples in parallel for any size. The presented solution can process UHD 8K videos at 60 frames per second when running at 46.6 MHz, with a power dissipation of 44.48 mW and an area of 261.28 Kgates. To the best of authors' knowledge, this is the first work in the literature presenting a hardware design for the AV1 DCT-II transform. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2023 | Learning-based bypass zone search algorithm for fast motion estimation
Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
Multim. Tools Appl. | 4 |
| 2023 | A High-Throughput Hardware Design for the AV1 Decoder IntrapredictionabstractThe Alliance for Open Media (AOMedia) (AV1) was released in 2018 as a royalty-free and open-source video codec. AV1 was developed by the AOMedia that is composed of many leading tech companies. AV1 has the goal to process ultrahigh definition (UHD) 8K (7680$\times4320$pixels) and 4K videos (3840$\times2160$pixels) and to achieve high coding efficiency, which leads to increased complexity when compared to other codecs in the market, such as VP9, HEVC, and H.264. This article presents the AV1 intraprediction decoder (AVID), a dedicated high-throughput hardware design for the AV1 decoder intraprediction supporting the AV1 68 prediction modes and 19 block sizes. The proposed architecture can decode UHD 4K videos at 120 frames/s in the worst case, requiring an operation frequency of 279.93 MHz and demanding a total area of 234.45 kgates with a power dissipation of 27.74 mW. The comparison with related works showed that AVID reached the smallest area and very competitive power results. To the best of the authors’ knowledge, this is the first article detailing the hardware design of a complete decoder for intraprediction targeting the AV1 codec. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | GM-RF: An AV1 Intra-Frame Fast Decision Based on Random ForestabstractThis paper presents the Grouping of Modes based on Random Forest (GM-RF), a fast decision algorithm for the AOMedia Video 1 (AV1) intra-frame prediction applying machine learning (ML). AV1 implements a wide variety of intra-frame prediction tools, significantly increasing the required computational effort. The GM-RF uses trained Random Forest (RF) models to reduce the number of intra-frame prediction modes evaluated for each encoded block. Experimental results show that the GM-RF achieves an average time savings of 50.19%, with a BD-BR of 7.41%. Compared with related works, GM-RF reached time savings from 5.6 to 10 times higher at a cost of a higher BDBR. To the best of the authors’ knowledge, this is the first solution in the literature using ML to reduce the AV1 intra-frame prediction computational effort. Pablo Rosa, Daniel Palomino 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 3 |
| 2022 | Fast Affine Motion Estimation for VVC using Machine-Learning-Based Early Search TerminationabstractThe Affine Motion Estimation (AME) was introduced in the Versatile Video Coding (VVC) standard to allow for the detection of non-translational transformations during inter-frame prediction. Although providing important coding efficiency gains, this new tool represents 43% of the motion estimation (ME) complexity. However, an analysis over the AME step shows that the Affine motion vectors are often generated without resulting in the best ME prediction. This paper proposes a AME early search termination based on supervised machine learning. Six Random Forest models were trained with features obtained during the encoding process to accurately predict whether the AME step should be executed, partially executed or skipped, avoiding unnecessary calculations. As result, the proposed solution achieves an average time saving of 46.94% in the AME step with a coding efficiency loss of only 0.18%. Adson Duarte, Luciano Volcan Agostini, Bruno Zatt, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Daniel Palomino 0001 |
ISCAS | 6 |
| 2022 | A High-Throughput Design for the H.266/VVC Low-Frequency Non-Separable TransformabstractThis paper presents a high throughput hardware design for the Low-Frequency Non-Separable Transform (LFNST) of the Versatile Video Coding (H.266/VVC) standard. The LFNST is a secondary transform used to transform the coefficients already transformed by the DCT-II as primary transform over the residues from the directional intra prediction. The LFNST architecture was designed to process Ultra-High Definition (UHD) videos with $4098 \times 2160$ pixels (4K) at 60 frames per second. Our solution presents an area utilization of 99.13 kgates and a power dissipation of 38.50 mW, when running at 186.62 MHz and considering the worst-case operation (processing the LFNST $4\times 4$ through TU size of $4\times 4$). Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2022 | Standard Cell and Supergates Designs: An Electrical Comparison on 4-Input Logic FunctionsabstractThis paper presents an electrical study on logic functions with up to 4 inputs designed with a standard cell mapping and two automatically generated supergates methodologies. The results indicate that supergate-based designs reduce the average power in 84.4% of the studied cases while reducing area by 12.9%. Despite the supergate design increasing in average the circuit critical delay by 5.8%, it achieves better power-delay-product in 2823 (70.9%) of the 3982 studied logic functions. The reduction of logic levels is the main factor for gains obtained with supergates due to the glitch power reduction. Henrique Kessler, Marcelo Schiavon Porto, Leomar S. da Rosa Jr., Vinícius V. Camargo |
ISCAS | 2 |
| 2022 | Multi-Objective optimized Complexity Control for the AV1 Video EncoderabstractAOMedia Video 1 (AV1) is an open-source video encoding format launched in 2018 by Alliance for Open Media. To achieve high coding efficiency, AV1 brings significant innovations in comparison to its predecessor, the VP9 format. However, this came at the cost of a higher encoding complexity due to the newly introduced tools and block partitioning structures. To allow for wide deployment of AV1 in different multimedia devices, adjusting its encoding process according to the available computational resources is strongly desirable. Thus, this work presents an adaptive complexity controller for the AV1 encoder based on multi-objective optimization. Experimental results show that the proposed controller reduces encoding complexity in a target range from 10% to 40%, with satisfactory precision varying between 0.04 and 4.10 percentage points, and a BD-Rate between 0.22% and 5.24%. Gustavo Rehbein, Isis Bender, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
PCS | 5 |
| 2021 | Block-Based Inter-Frame Prediction For Dynamic Point Cloud CompressionabstractIn recent years, 3D point clouds have gained popularity thanks to technological advances such as the increased computational power and the availability of low-cost devices for acquisition of 3D information, like RGBD sensors. However, raw point clouds demand a large amount of data for their representation, and compression is mandatory to allow efficient transmission and storage. Inter-frame prediction is a widely used approach to achieve high compression rates in 2D video encoders, but the current literature still lacks solutions that efficiently exploit temporal redundancy for point cloud encoding. In this work, we propose a novel inter-frame prediction for 3D point cloud compression, which explores temporal redundancies in the 3D space. Moreover, a mode decision algorithm is also proposed to dynamically choose the best encoding mode between inter and intra prediction. The proposed method yields a bitrate reduction of 15.6% and 3.5% for geometry and luma information respectively, with no significant impact in objective quality when compared to the MPEG 3DG solution, called G-PCC. Cristiano Santos, Mateus M. Gonçalves, Guilherme Corrêa 0001, Marcelo Schiavon Porto |
ICIP | 4 |
| 2021 | Low-Power and High-Throughput Approximated Architecture for AV1 FME InterpolationabstractModern video encoders like the AOM Video 1 (AV1) implement several complex tools to allow the required high level of compression efficiency. The Fractional Motion Estimation (FME) is one of these tools and in AV1 the FME defines 90 different filters. To handle such complexity, hardware acceleration using approximate computing has become an alternative to be explored. This paper presents an approximate solution for the AV1 FME interpolation filters based on the approximation of the original filter coefficients intending to generate more hardware friendly coefficients. The approximated version was designed in hardware and can achieve real-time interpolation for UHD 8K videos at 30 frames per second, when synthesized using 40nm TSMC standard-cells technology. The designed architecture dissipates 26.79mW which represents more than 80% power reduction when compared to the original precise solution. The approximation implied in a small average coding efficiency degradation of 0.54% in BD-BR. When comparing with related works, this architecture reaches an expressive power reduction (2.1 to 4.8 times) even supporting more complex tools. Robson Domanski, William Kolodziejski, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 4 |
| 2021 | Energy-Throughput Configurable Design for Video Processing Binary Arithmetic EncoderabstractVideo encoding draws high research interest, due to the enormous demand for video traffic and real-time encoding for transmission. In video encoding standards such as HEVC (High-Efficiency Video Coding), the final step of the encoding stage is the CABAC (Context-Adaptive Binary Arithmetic Coding). The coding efficiency of the CABAC comes at the cost of increased computational complexity, especially for parallelization purposes, being the BAE (Binary Arithmetic Encoder) the critical part of CABAC. Thus, an important goal is to balance the real-time throughput requirements and the power/energy consumption in the design of BAE dedicated hardware. This work introduces a novel configurable high-throughput BAE design, named ET-BAE, utilizing a combination of a new modified Multiple-Bypass Bins Scheme (MBBS) and a power-saving approach into a single ASIC design with two-mode configuration. Synthesis and power-analysis results show that the configurable BAE design, the first of its kind with this feature, is more energy-efficient and less area consuming than utilizing non-configurable versions. The ET-BAE is able to accomplish the same real-time requirements of its competitors. Fábio Luís Livi Ramos, Bruno Zatt, Marcelo Schiavon Porto, Sergio Bampi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | RDE-MOGA: Automatic Selection of Rate-Distortion-Energy Control Points for Video Encoders Using Muti-Objetive Genetic AlgorithmabstractControlling energy consumption of video encoders is a complex multi-objective optimization problem of great importance. In this work we propose the RDE-MOGA, an multi-objective genetic algorithm capable of finding energetically efficient configurations for the HEVC encoder and replacing the current sensitivity analysis methodologies in the development of energy controllers. The utilization of our algorithm improved its efficiency in 60% whereas increasing the range of achievable reductions of the controller in at least 50%. Furthermore, the algorithm proved capable of sustaining 30% energy reduction at a cost of 3.45 BD-BR loss. Italo Machado, Marilton S. de Aguiar, Marcelo Schiavon Porto, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ICASSP | 3 |
| 2020 | Low-Power and Memory-Aware Approximate Hardware Architecture for Fractional Motion Estimation Interpolation on HEVCabstractNowadays, current video coding standards like the High Efficiency Video Coding (HEVC) implement several complex coding tools, like the Fractional Motion Estimation (FME). An alternative to improve performance and save power is allying the hardware acceleration with approximate computing solutions, focusing on such complex tools. In this work, we present a low-power and memory-aware hardware architecture for the HEVC FME interpolator, proposing the development of two novel hardware designs for the interpolation filters, called Approximate Unified FME Filters (AUFF). These solutions exploit the usage of approximate computing at both algorithmic and data levels, leading to a reduction in dissipated power and memory bandwidth. The proposed design is capable of real-time interpolation of UHD (Ultra High Definition) 4K and 8K videos when synthesized using a 40 nm standard-cell library, with a power dissipation ranging from 22.04 to 62.06 mW. Wagner Penny, Guilherme Corrêa 0001, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Gabriel L. Nazar, Bruno Zatt |
ISCAS | 5 |
| 2020 | Efficient Hardware Design for the AV1 CDEF Filter Targeting 4K UHD VideosabstractDeveloped by the AOMedia industry consortium, the AOM Video 1 (AV1) is an open-source and royalty-free video encoder released in June 2018. The Constrained Directional Enhancement Filter (CDEF) is one of the three AV1 in-loop filters and it is the focus of this work. The CDEF has the goal to reduce ringing artifacts generated with the encoding process, acting as a directional deringing filter. This paper presents a hardware design for the AV1 CDEF targeting real-time processing of 4K Ultra High Definition (UHD) videos. The architecture was synthesized to ASIC using the 40nm TSMC library, requiring 185 kgates and with a power dissipation of 43 mW when running at 93 MHz, reaching the frame rate of 60 frames per second (fps). To the best of the author's knowledge, there is no other work in the literature with dedicated hardware design for the AV1 CDEF. Eduardo Zummach, Roberta Palau, Jones Goebel, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 6 |
| 2020 | 4D-DCT Hardware Architecture for JPEG Pleno Light Field CodingabstractThis paper presents a 4D-DCT hardware architecture for Light Field Coding according to the JPEG Pleno standard. It is composed of two instances of 2D-DCT engines and a novel 4D Transposition Memory organization. Experimentally-defined fixed-point representation and LSB pruning techniques are employed do reduce hardware area and power dissipation. The proposed architecture operates over 4D-hypercubes of up to 8x8x8x8 samples and reaches performance to process 30 Lytro-like light fields per second dissipating 145.32mW at 825.75MHz. This is the first known 4D-DCT hardware architecture for light field coding and demonstrates the feasibility of such solutions on real-world systems. Matheus Jahnke, Jones Goebel, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 6 |
| 2020 | Power/QoS-Adaptive HEVC FME Hardware using Machine Learning-Based Approximation ControlabstractThis paper presents a machine learning-based adaptive approximate hardware design targeting the fractional motion estimation (FME) of HEVC encoder. Hardware designs targeting multiple levels of approximation are proposed, by changing FME filters coefficients and/or discarding taps. The level of approximation is defined by a decision tree, generated taking into account the behavior of several parameters of the encoding in order to predict homogeneous blocks, more suitable for more aggressive approximation without significant losses on quality of service (QoS). Instead of applying a specific level of approximation over the full video, different approximate FME accelerators are dynamically selected. Such a strategy is able to provide up to 50.54% of power reduction while keeping the QoS losses at 1.18% BD-BR. Wagner Penny, Daniel Palomino 0001, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 3 |
| 2020 | Complexity and compression efficiency assessment of 3D-HEVC encoder
Mário Saldanha, Ruhan A. Conceição, Vladimir Afonso, Giovanni Avila, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Guilherme Corrêa 0001, Luciano Volcan Agostini |
Multim. Tools Appl. | 6 |
| 2019 | Fast Hevc-to-Av1 Transcoding Based On Coding Unit Depth InheritanceabstractWith the advent of the recently launched AOMedia Video 1 (AV1) bitstream specification, there is currently a need for converting legacy content encoded with the state-of-the-art High Efficiency Video Coding (HEVC) standard to the new format. However, transcoding is a complex task composed of a decoding and an encoding process in sequence, which requires long processing times and high energy consumption. This paper proposes the first HEVC-to-AV1 transcoding solution, which is based on the high correlation between block size decisions in HEVC and AV1. The solution allows the AV1 encoder to inherit Coding Unit (CU) depth information from the HEVC bitstream to constrain the AV1 re-encoding process. Experimental results show an average transcoding time reduction of 35.41% at the cost of a compression efficiency loss of 4.54%. Bruno Zatt, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 3 |
| 2019 | Encoding Efficiency and Computational Cost Assessment of State-Of-The-Art Point Cloud CodecsabstractPoint clouds have recently emerged as a suitable solution to generate and display 3D digital models due to their capacity of representing high resolution images and videos through multiple viewpoints. However, as they are usually made up of thousands up to billions of points, advanced techniques of data compression are essential to store and transmit this type of data. This paper compares the two state-of-the-art solutions for point cloud compression, the Point Cloud Codec (PCC) and the Test Model Category 2 (TMC2), in terms of compression efficiency and encoding time. Experimental results show that the compression efficiency for geometry information is highly dependent upon the available bitrate for both TMC2 and PCC. However, for texture compression TMC2 almost always achieves the best results. The experiments have also shown that TMC2 presents a computational cost from 22.2 to 26 times larger the observed in PCC. Mateus M. Gonçalves, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 4 |
| 2019 | Energy-Aware Motion and Disparity Estimation System for 3D-HEVC With Run-Time Adaptive Memory HierarchyabstractThe popularization of multimedia services has pushed forward the development of 2D/3D video-capable embedded mobile devices. Such devices require efficient energy/memory-management strategies to deal with severe memory/processing requirements and limited energy supply. Therefore, we propose a motion and disparity estimation (ME and DE) system—the most memory/processing demanding encoding steps—for the 3D High Efficiency Video Coding (3D-HEVC) standard. It was designed for low energy consumption, featuring a run-time adaptive memory hierarchy. The processing unit employs flexible coding order and optimizations to reduce the computational effort by exploring the inter-channel and inter-view redundancies. The memory hierarchy features window-based prefetching, data reuse, subsampling, and dynamic voltage scaling controlled by our depth-based dynamic search window resizing algorithm. Memory results demonstrate an average on-chip energy reduction of 79% in comparison to the widely used Level-C solution for a 45-nm technology. The proposed energy-aware ME and DE system dissipates 7.55 W while processing three HD 1080p views (video + depth) at 30 frames per second and presents a mean energy consumption of 0.107 J per access unit. To the best of our knowledge, this is the first work that proposes a real-time ME/DE system for the 3D-HEVC standard with an adaptive memory hierarchy. Vladimir Afonso, Ruhan A. Conceição, Mário Saldanha, Luciano Almeida Braatz, Murilo R. Perleberg, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt, Altamiro Amadeu Susin |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2018 | Octagonal-Axis Raster Pattern for Improved Test Zone Search Motion EstimationabstractTest Zone Search (TZS) is considered the current state-of-the-art fast Motion Estimation algorithm because it presents the best tradeoff between compression efficiency and complexity in comparison to the Full Search strategy. However, it is still one of the most computationally-demanding tools of current video coding standards, such as the High Efficiency Video Coding (HEVC). This paper presents an analysis on the search area opportunities and best match distributions in TZS, which led to the proposal of a novel search pattern in its most complex step, the Raster Search (RS). The new pattern, named Octagonal-Axis Raster Pattern (OARP), allowed an average complexity reduction of 61 % in TZS, with a negligible BD-rate increase of 0.0371 % in comparison to the original algorithm. Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001 |
ICASSP | 2 |
| 2018 | LF-CAE: Context-Adaptive Encoding for Lenslet Light Fields Using HEVCabstractLight fields can outperform the representability of current imaging technologies by considering the angle of light-rays striking in the camera in addition to the color information. The additional information brings a set of challenges related to the amount of data required to represent light fields, arising the need for efficient compression schemes. This work proposes a novel and efficient scheme to encode lenslet light fields, called Light Fields Context-Adaptive Encoding(LF-CAE). LF -CAE is a video-based compression solution that defines a flexible and dynamic scheme to create intermediate video sequences from light fields in order to efficiently exploit the HEVC inter-frame encoder structure. This scheme reduces the inter-frame prediction residue leading to compression efficiency gains in light-field coding. LF -CAE reaches an average compression rate of 99.56% in relation to the uncompressed light field, with an average PSNR of 37.61dB. LF-CAE surpasses all related works that decompose the light field into an intermediate video sequence, reaching the highest PSNR and with BD-Rate gains ranging from 6.07% up to 20.50%. Ruhan A. Conceição, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 2 |
| 2018 | Hardware-Friendly Unidirectional Disparity-Search Algorithm for 3D-HEVCabstractThis paper presents a novel hardware-friendly Unidirectional Disparity-Search (UDS) algorithm for the 3D-HEVC. This algorithm explores the typical camera arrangements used in 3D-HEVC. UDS was evaluated in two operation points reaching a computational effort reduction from 32.7% to 61.8%, with a BD-Rate increase from 0.3123% to 0.4803%, when compared to TZS. Estimated hardware results showed memory-size and leakage-energy reductions of 66.7%, a dynamic-energy reduction from 34.5% to 62.1%, and an energy-consumption reduction from 32.7% to 61.8% in SAD calculations, when compared to TZS. To the best of the authors' knowledge, this is the first proposed DE algorithm that explores the 3D-HEVC typical camera arrangement. Vladimir Afonso, Altamiro Amadeu Susin, Murilo R. Perleberg, Ruhan A. Conceição, Guilherme Corrêa 0001, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 8 |
| 2018 | High-Throughput and Low-Power Integrated Direct/Inverse HEVC Quantization Hardware DesignabstractThis paper presents a high-throughput and low-power integrated HEVC direct/inverse quantization hardware design. The main focus of this design is to allow the evaluation of multiple coding modes during the residual encoding process of the HEVC for real-time Ultra-High Definition (UHD) video processing. The ASIC synthesis results, for a Nangate 45nm standard cell library, presented a maximum operational frequency of 1679.51MHz and a processing rate of 53.74 Gsps (giga samples per second). This throughput allows processing of real-time up to 72 coding modes for UHD 4K@60fps or up to nine coding modes for the UHD 8K@120fps while dissipating 369.37mW. Luciano Almeida Braatz, Bruno Zatt, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2018 | Configurable Cache Memory Architecture for Low-Energy Motion EstimationabstractThe popularization of mobile devices and the increased demand for video applications from these devices necessitates the design of efficient video encoders such as HEVC. Since the Motion Estimation (ME) is the most processing and memory intensive unit in a video encoder, our focus is in the communication between external memory and the ME unit. The TZS algorithm is widely used in video encoders and has an unpredictable behavior, which leads to an unknown pattern of memory accesses, making SPMs ineffective solutions, for example. Therefore, this work proposes a configurable cache memory architecture for fast ME algorithms. This cache has settings that suit different video encoding scenarios. Six optimal cache configurations were defined based on our evaluation considering 23 video sequences, 4 QPs, and 32 different cache settings. External memory bandwidth savings of up to 96.84% were reached, representing a reduction from 25.48GB/s to 548.53MB/s in the best case. When compared to Level-C SPM and to a static 16KB 8-way associative cache, the proposed configurable cache achieves energy savings of up to 86.91% and 78.09%, respectively. Anderson Martins, Wagner Penny, Matheus Weber, Luciano Volcan Agostini, Marcelo Schiavon Porto, Daniel Palomino 0001, Júlio C. B. de Mattos, Bruno Zatt |
ISCAS | 5 |
| 2018 | High-Throughput Binary Arithmetic Encoder using Multiple-Bypass Bins Processing for HEVC CABACabstractThe advance of massive video processing applications, devices and resolutions has led to new challenges in video encoding. The HEVC (High Efficiency Video Coding) standard emerges as one alternative in order to address the new video processing requirements. The HEVC allows only one type of entropy encoding algorithm, which is the CABAC (Context-Adaptive Binary Arithmetic Coding). The compression gains achieved by CABAC algorithm come at the cost of increasing complexity for implementation, due to intense data dependencies. The BAE (Binary Arithmetic Encoder) is the CABAC critical sub-block, in which the main part of the algorithm is executed. The present work proposes an 8-stage pipeline BAE architectural solution, named MB-BAE, with the addition of multiple-bypass bins processing, in order to increase throughput without compromising the critical path of the architecture. As a result, an average of 4.94 bins/cycle and around 2.6-Gbin/s of throughput are achieved in our BAE. This is the highest throughput found among related works in the literature, and being able to process 8K UHD videos with the lowest frequency when compared to the same related works. Fábio Luís Livi Ramos, Bruno Zatt, Marcelo Schiavon Porto, Sergio Bampi |
ISCAS | 3 |
| 2017 | Low-power and high-throughput hardware design for the 3D-HEVC depth intra skipabstractThis paper presents a low-power and high-throughput hardware design for the 3D-HEVC (Three Dimensional High Efficiency Video Coding) Depth Intra Skip coding tool. A strategy to reduce the computational effort was employed based on an analysis using the 3D-HEVC reference software. The proposed strategy consists of replacing the SVDC (Synthesized View Distortion Change) for the SAD (Sum of Absolute Differences) as the similarity criterion. This way, the number of arithmetic operations related with the similarity criterion is reduced over 71%, and a rendering process is avoided at the cost of only 0.21% increase in the BD-Rate. The hardware was described in VHDL and synthesized for ASIC technology. The synthesis results for the 45nm Nangate standard cells demonstrate that the architecture can process 60 UHD 2160p frames per second (five views) with a power dissipation of 19.57mW. Vladimir Afonso, Altamiro Amadeu Susin, Luan Audibert, Mário Saldanha, Ruhan A. Conceição, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 6 |
| 2017 | A multiplierless parallel HEVC quantization hardware for real-time UHD 8K video codingabstractOne step required several times for current video encoders is the residual coding loop, composed of the direct transformation, direct quantization, inverse quantization, and inverse transformation. These operations demand high throughput and low latency since their outputs must be processed by other steps of the coder. This paper proposes a high-throughput parallel and multiplierless hardware architecture for the HEVC direct quantization targeting real-time processing of Ultra-High Definition 8K videos. The proposed architecture support frequency dependent quantization steps. The binary multiplications were replaced by multiple constant multiplications in order to improve the throughput and to reduce the area and power dissipation. The developed design is able to process 32 samples in parallel, which represents one line of the biggest HEVC transform block. The ASIC synthesis results, obtained with Nangate 45nm standard cells library, show that the proposed architecture is able to quantize about 8 billion coefficients per second, when running at 186.6 MHz, with a gate count of 168,330. This throughput is enough to process UHD 8K videos at 120 fps. Luciano Almeida Braatz, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2017 | High-throughput HEVC intrapicture prediction hardware design targeting UHD 8K videosabstractThis paper presents a high-throughput hardware architecture for the HEVC intrapicture prediction targeting the processing of UHD 8K (7680×4320 pixels) videos at 120 frames per second. The proposed design supports all intra prediction modes and all block sizes. It implements an internal mode decision algorithm that is more hardware friendly than RMD at the cost of a negligible 0.17% BD-rate impact (intra-only). When synthesized to the NanGate 45nm 0.95v cell library targeting a frequency of 529MHz, the proposed design used 4952K gates and showed a power dissipation and an energy efficiency of 363mW and 32.02pJ/sample respectively. Marcel Moscarelli Corrêa, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 3 |
| 2017 | Multiple early-termination scheme for TZ search algorithm based on data mining and decision treesabstractThe latest video compression standards, such as the H.264/AVC and the High Efficiency Video Coding (HEVC), provide fast Motion Estimation (ME) algorithms in their reference software aiming at complexity reduction. Test Zone Search (TZS) is the state-of-the-art fast ME algorithm, currently deployed in the reference HEVC encoder due to its great coding efficiency. However, ME is still one of the main sources of complexity in HEVC. This paper proposes an early-termination scheme for TZS, called e-TZS, based on an extensive data mining process on ME attributes. The data mining process allowed identifying the most relevant information during the encoding process to build a set of decision tree models that terminate TZS in different steps of its execution. The e-TZS scheme was implemented in the HEVC reference software and achieved average decision precision of 94.2%. Experimental results showed an average complexity reduction of 62.53% in TZS, with a negligible BD-rate increase of only 0.49%, in comparison to the original algorithm. Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
MMSP | 3 |
| 2017 | Energy-efficient motion estimation with approximate arithmeticabstractEnergy efficiency has become a primary concern in the design of multimedia digital systems, particularly when targeting mobile devices. Approximate computing is a highly promising approach to address this challenge. This paper presents an architectural exploration in a variable block size motion estimation (VBSME) architecture using imprecise Lower-Part-OR Adders (LOA). These adders were applied to Sum of Absolute Differences units (SAD) in order to reduce the energy consumption while introducing a minimum impact on the coding efficiency. Three VBSME architectures with LOA operators were developed by considering different imprecision levels. The conducted evaluations, performed using the High-Efficiency Video Coding standard (HEVC) reference software, showed that this technique introduces a negligible impact on the coding efficiency (between 0.6% and 2.5% increase of the BD-Rate). Nevertheless, when the designed architectures were synthesized for a 45nm standard cells technology, significant power savings were observed (between 7% and 11.5%, depending on the used LOA version), demonstrating the viability and significant gains of the proposed approach. Roger Endrigo Carvalho Porto, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto, Nuno Roma, Leonel Sousa |
MMSP | 4 |
| 2016 | Complexity reduction for 3D-HEVC depth map coding based on early Skip and early DIS schemeabstractThis paper presents a novel early Skip/DIS mode decision for 3D-HEVC depth encoding which aims at reducing the complexity effort of this process. The proposed solution is based on an adaptive threshold model, which takes into consideration the occurrence rate of both Skip and DIS modes. Occurrence analysis showed that the lower is the Skip and DIS Rate-Distortion cost, the higher is the probability of these modes being chosen. Furthermore, software evaluations showed that the proposed early Skip/DIS scheme is capable of reducing the depth coder complexity in 24.4% for a target hit rate of 99%, and in 33.7% for a target hit rate of 95%, leading to a negligible coding efficiency penalty in both scenarios. Ruhan A. Conceição, Giovanni Avila, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 4 |
| 2016 | High-throughput and memory-aware hardware of a sub-pixel interpolator for multiple video coding standardsabstractReal-time operation and low-power dissipation in video coding systems have become important research challenges, especially in mobile devices with limited battery and computational resources. There are many video coding standards coexisting in the market nowadays, so it is important for current devices to support different video coding standards. This paper presents a multi-standard luminance sub-samples interpolator hardware design for the Motion Compensation (MC) and Fractional Motion Estimation (FME), with support to MPEG-2/4, H.264/AVC, HEVC, and AVS/2 video coding standards. Our design is able to save hardware resources through an optimized filter organization, totally compliant with the focused standards and capable to interpolate samples for UHD 4320p@60fps at real time. The 45nm standard-cell library implementation dissipates 10mW, when processing according MPEG-2 standard, up to 46.4mW when processing AVS2. Guilherme Paim, Jones Goebel, Wagner Penny, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 5 |
| 2016 | An efficient sub-sample interpolator hardware for VP9-10 standardsabstractThis paper presents a hardware design for the sub-sample interpolator used in FME (Fractional Motion Estimation) and MC (Motion Compensation) stages according to the VP9 and VP10 video-coding standards. The proposed architecture is able to save hardware resources through an optimized-filter organization whereas reaching high-throughput and low-power dissipation. The hardware design was described in Verilog and synthesized for ASIC technology. The synthesis results were generated for 45nm Nangate standard cells and demonstrate that the developed architecture is able to process 2160p@60fps videos with a power dissipation of 2.34mW focusing on a VP9-10 decoder. Guilherme Paim, Wagner Penny, Jones Goebel, Vladimir Afonso, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 6 |
| 2016 | Pareto-based energy control for the HEVC encoderabstractThe current state-of-art video coding standard, the High Efficiency Video Coding (HEVC), brings many innovations as a way to improve the coding performance. However, the improvement on performance also brought higher computational effort and energy consumption. Since most of devices that handle digital videos are battery powered, the energy consumption became an important issue that demands efficient solutions. This way, controlling energy consumption is strongly desirable to adapt the encoding process to the energy availability. This goal is a hard task due the heterogeneous dynamic behavior of HEVC encoder. This work presents the development of a Pareto-based dynamic energy controller for the HEVC encoder, reaching up to 70% energy saving with small losses on coding efficiency for most of the cases. Wagner Penny, Italo Machado, Marcelo Schiavon Porto, Luciano Volcan Agostini, Bruno Zatt |
ICIP | 3 |
| 2016 | An HEVC multi-size DCT hardware with constant throughput and supporting heterogeneous CUsabstractThis paper presents an efficient hardware design for the Discrete Cosine Transform (DCT) of High Efficiency Video Coding standard (HEVC). This hardware supports all HEVC transform sizes: 4×4, 8×8, 16×16, and 32×32 including any combination of the Transform Unit (TU) sizes. The proposed DCT architecture has a constant throughput of 32 coefficients per cycle, independently of the transform sizes combination. The architecture was synthesized for a Nangate 45nm standard-cell library and the power analysis was made considering real input vectors. The synthesis results show a very good tradeoff between area, power dissipation and processing rates. The architecture is able to process 1.6G coeff/s when running at 50MHz dissipating 24.2 mW. These results allow a processing rate of 30 HD 1080p frames per second when evaluating 17 HEVC prediction modes. Jones Goebel, Guilherme Paim, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 5 |
| 2015 | A multi-standard interpolation hardware solution for H.264 and HEVCabstractAttending real-time constraints in video coding systems represents a big challenge for nowadays systems, especially for high definition videos at mobile systems. The Fractional Motion Estimation (FME) and Motion Compensation (MC) are responsible for a large share of processing effort in both state-of-the-art video coding standards, the High Efficiency Video Coding (HEVC), and its predecessor, the H.264. This work proposes a multi-standard hardware solution for the fractional sample interpolation used in FME/MC processing of the HEVC and H.264 standards. The hardware design is composed of four IP (Intellectual Property) cores able to process 1080p@60fps videos independently. The whole architecture can process 2160p@60fps with 80.69mW, considering bi-prediction. Henrique Maich, Guilherme Paim, Vladimir Afonso, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ICIP | 6 |
| 2015 | Complexity reduction for the 3D-HEVC depth maps codingabstractThis paper presents a qualitative discussion of the depth maps properties that can be considered to achieve complexity reduction for 3D-High Efficiency Video Coding (3D-HEVC) depth maps coding. Both intra and inter-frame predictions are considered in this discussion that conduced to the proposition of two simple complexity reduction techniques: the Simplified Edge Detector (SED) and the Diamond Search (DS) simplified inter-prediction. The SED anticipates the blocks that are likely to be better predicted by the HEVC intra-prediction, avoiding evaluations of Depth Modeling Modes (DMM). The DS and SED were compared to anchor results and experimental analysis showed that the proposed algorithms are able to achieve a time saving of 11.3% encoding time reduction, with acceptable impact on the BD-Rate of the synthesized views of 0.6%. Mário Saldanha, Gustavo Sanchez, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ISCAS | 4 |
| 2015 | A real-time architecture for reference frame compression for high definition video codersabstractCurrent battery-powered devices that manipulate digital videos must consider the energy consumption of this process as an important issue, especially when high or ultra-high definition videos are handled. In this scenario, this paper proposes a solution to reduce the energy consumption in video coding systems by reducing the external memory communication during the motion estimation. The scheme presented in this paper is called Differential Reference Frame Coder and it implements an algorithm that combines two techniques to reduce the memory bandwidth: a differential coding based on a simplified intra-prediction process, to reduce the spatial redundancy of the reconstructed samples, and a semi-fixed length coding applied in the residues generated by the differential coding step. This solution reaches an average lossless compression ratio higher than 57% for the evaluated HD 1080p video sequences whereas supporting random access to reference frame blocks. The proposed hardware architectures (Coder and Decoder) were described in VHDL and synthesized targeting ASIC for 65nm and 180nm TSMC standard-cell libraries. The results show that with 65nm, the architectures are able to process UHD 2160p (3840×2160 samples) at 30 fps or HD 1080p (1920×1080 samples) at 120 fps with a power dissipation of 0.885mW. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 6 |
| 2014 | A low-complexity and lossless reference frame encoder algorithm for video codingabstractThis paper presents a lossless coding solution to reduce the large overhead of external memory communication during the motion estimation process in current video coders. Our solution is called Differential Reference Frame Coder (DRFC), and uses two techniques together to compress the reference frame: a differential coding based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples, and a simple VLC applied to differential coding residues. The proposed solution reaches an average compression rate higher than 45% for the evaluated HD 1080p video sequences. This is a lossless and low-complexity solution, and could easily be implemented in hardware. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ICASSP | 6 |
| 2014 | Complexity reduction for 3D-HEVC depth maps intra-frame prediction using simplified edge detector algorithmabstractThis paper presents a new mode decision for the depth maps intra-frame prediction in 3D-HEVC. The proposed technique decides if the traditional High Efficiency Video Coding-based (HEVC) intra-frame prediction should be performed or skipped. This technique is inspired by the fact that traditional intra-frame prediction may generate artifacts in the synthesized views when an edge is encoded. The Simplified Edge Detector (SED) algorithm has been proposed to classify if a block contains an edge or a nearly constant region demanding a minimum processing overhead. Through software evaluations, SED algorithm was capable to obtain an average complexity reduction of 23.8% for depth maps coding with no quality losses. Gustavo Sanchez, Mário Saldanha, Gabriel Balota, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 5 |
| 2014 | A new differential and lossless Reference Frame Variable-Length Coder: An approach for high definition video codersabstractThis paper presents a novel solution for external memory bandwidth reduction in video coding systems. The approach is based on reference frame compression, using a differential coding and a hardware-aware adaptation of the traditional Huffman algorithm, besides, it is a lossless solution fully compliant to state-of-art video coding standards, as H.264/AVC and HEVC. This solution is called DRFVLC (Differential Reference Frame Variable-Length Coder) and it uses differential coding to concentrate the samples values distribution. With the samples concentrated, an efficient static Huffman coding is applied to represent them in fewer bits. The DRFVLC reaches an average compression rate higher than 60% for the evaluated HD 1080p video sequences. This compression rate also indicates the external memory bandwidth reduction achieved with our technique. This solution can be easily implemented in hardware demanding one differentiator and a simple variable-length coder. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ICIP | 6 |
| 2014 | Power efficient and high troughtput multi-size IDCT targeting UHD HEVC decodersabstractThis paper is focused on the inverse transforms defined in the HEVC (High Efficiency Video Coding) standard. The HEVC standard allows the use of four transform sizes, including novel transforms applied over bigger block sizes (16×16 and 32×32). The hardware architecture presented in this paper was planned to reach real-time processing (at 30 frames per second) for ultra-higher solution videos, exploiting high level of parallelism. As a secondary goal, the architecture was also planned to reach low cost in terms of hardware consumption and power dissipation. Thus, the architecture was designed in a purely combinational way, using a multiplierless approach and employing an optimization algorithm through operations reuse and sub-expressions sharing. The synthesis targeted an Altera Stratix V FPGA and ASIC 90nm standard-cells technology. The synthesis results show that the designed architecture has the best performance results among all related works, being able to achieve real-time decoding for UHD videos (7680×4320 pixels) with a power consumption from 33.8mW to 339.2 mW. Ruhan A. Conceição, J. Claudio de Souza, Ricardo Jeske, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ISCAS | 4 |
| 2014 | Memory bandwidth reduction for H.264 and HEVC encoders using lossless reference frame codingabstractThis paper presents a hardware-efficient algorithm for external memory bandwidth reduction focusing on the state-of-the-art video encoders, like H.264/AVC and HEVC. The proposed approach is a lossless solution based on an adaptation of the traditional Huffman algorithm. This solution is entitled RFCAVLC8T (Reference Frame Context Adaptive Variable-Length Coder with 8 Tables) and is based on the use of off-line defined static Huffman tables. The RFCAVLC8T is a hardware-efficient version of the Huffman algorithm that employs eight static tables to avoid the cost of the on-the-fly Huffman statistical analysis. The best table to encode a block is selected at run time using a context evaluation, resulting in a context-adaptive configuration. The use of RFCAVLC8T reaches an average compression rate higher than 35% for the evaluated video sequences, with computational cost of a single VLC. Dieison Silveira, Guilherme Povala, Lívia Amaral, Bruno Zatt, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 6 |
| 2014 | Sample adaptive offset filter hardware design for HEVC encoderabstractThis work presents a hardware design for the Sample Adaptive Offset filter, which is an innovation brought by the new video coding standard HEVC. The architectures focus on the encoder side and include both classification methods used in SAO, the Band Offset and Edge Offset, and also the statistical calculations for the offset generation. The proposed architectures feature two sample buffers, classification units for both SAO types and the statistical collection unit. The architectures were described in VHDL and synthesized to an Altera Stratix V FPGA. The synthesis results show that the proposed architectures achieve 364MHz and are capable to process 44 QFHD (3840×2160) frames per second using 8,040 ALUTs of the target device hardware resources. Fabiane Rediess, Ruhan A. Conceição, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
VCIP | 4 |
| 2014 | A complexity reduction algorithm for depth maps intra prediction on the 3D-HEVCabstractThis paper proposes a complexity reduction algorithm for the depth maps intra prediction of the emerging 3D High Efficiency Video Coding standard (3D-HEVC). The 3D-HEVC introduces a new set of specific tools for the depth map coding that includes four Depth Modeling Modes (DMM) and these new features have inserted extra effort on the intra prediction. This extra effort is undesired and contributes to increasing the power consumption, which is a huge problem especially for embedded-systems. For this reason, this paper proposes a complexity reduction algorithm for the DMM 1, called Gradient-Based Mode One Filter (GMOF). This algorithm applies a filter to the borders of the encoded block and determines the best positions to evaluate the DMM 1, reducing the computational effort of DMM 1 process. Experimental analysis showed that GMOF is capable to achieve, in average, a complexity reduction of 9.8% on depth maps prediction, when evaluating under Common Test Conditions (CTC), with minor impacts on the quality of the synthesized views. Gustavo Sanchez, Mário Saldanha, Gabriel Balota, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
VCIP | 5 |
| 2013 | A hardware friedly motion estimation algorithm for the emergent HEVC standard and its low power hardware designabstractIn this paper the hardware friendly Multi Point Diamond Search (MPDS) motion estimation (ME) algorithm is evaluated in the High Efficiency Video Coding Standard (HEVC). The MPDS algorithm was implemented in HEVC reference software and its efficiency is compared with the standard HEVC fast algorithm, the Enhanced Predictive Zonal Search (EPZS). The evaluation result, in average, shows loses of only 1.7% in the compression rate and 0.05% in PSNR. However, the main advantage of the MPDS algorithm is its hardware friendly aspect and this paper also presents its hardware design focused on real time processing HD 1080p videos. The designed architecture is capable to process HD 1080p videos in real time when synthesized for TSCM 90nm technology. The MPDS algorithm has obtained a good tradeoff among quality, compression rate and costs for its hardware implementation. Gustavo Sanchez, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 2 |
| 2013 | A lossless approach for external memory bandwidth reduction in video coding systems and its VLSI architectureabstractThis paper presents the Reference Frame Context Adaptive Variable-Length Coder (RFCAVLC), which is a lossless solution to external memory bandwidth reduction in current video coding systems. The proposed approach is based on an adaptation of the traditional Huffman algorithm, and it uses eight static tables to avoid the cost of the on-the-fly statistical analysis. The best table to encode a block is defined using a context evaluation, resulting in a context-adaptive configuration. The use of RFCAVLC reached an average compression rate higher than 31% for the evaluated video sequences. The architectures that implement the RFCAVLC encoder and decoder were designed and synthesized to an FPGA device. The RFCAVLC design is able to reach real-time encoding for WQSXGA (3200 × 2048 pixels) at 30 fps. The synthesis results show that this solution can be easily coupled to a complete video encoder system with negligible hardware overhead and without compromising the throughput for real-time high-definition multimedia applications. Dieison Silveira, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICME | 2 |
| 2013 | Iterative random search: a new local minima resistant algorithm for motion estimation in high-definition videos
Marcelo Schiavon Porto, Cassio Cristani, Pargles Dall'Oglio, Mateus Grellert, Júlio C. B. de Mattos, Sergio Bampi, Luciano Volcan Agostini |
Multim. Tools Appl. | 1 |
| 2012 | High performance hardware architectures for the inverse Rotational Transform of the emerging HEVC standardabstractThis paper presents a dedicated hardware architecture for the Rotational Transform (ROT), which is one of the novel tools proposed for the emergent HEVC video coding standard. The main goal of this coding tool is to achieve higher energy compaction of the main transform coefficient matrix, minimizing the quantization error and improving the efficiency of the entropy encoding. Five versions of this architecture were implemented, using either a fully combinational structure or a pipeline with nine stages. The designed architectures were described in VHDL and synthesized for an Altera Stratix III FPGA. The synthesis results show that all versions can process very high resolution videos, such as QFHD, in real time. The version with the highest processing rate achieved a maximum operation frequency of 260.15 MHz. This architecture reaches a processing rate of 2.08 billion samples per second, allowing it to process UHDTV videos in real time. Henrique Avila Vianna, Gustavo Sanchez, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 3 |
| 2012 | Spread and Iterative Search: A High Quality Motion Estimation Algorithm for High Definition Videos and Its VLSI DesignabstractThis paper presents the Spread and Iterative Search (S&IS) motion estimation algorithm, which uses a random spread evaluation together with a central iterative evaluation to avoid local minima falls and to increase the image quality for high definition videos. Considering Full HD videos, S&IS reached an average PSNR gain of 1.41dB when compared to Diamond Search (DS), with an increase of about four times in the number of evaluated blocks. When compared to Full Search (FS), the S&IS achieved an average PSNR loss of 1.56 dB, evaluating 73 times less blocks than FS. An efficient architecture for the S&IS algorithm is also presented in this paper. The architecture was designed targeting in real time processing (30 frames per seconds) for QFHD videos (3840×2160 pixels). The architecture was described in VHDL and synthesized for and Altera Stratix 4 FPGA and for ST90nm standard cells technology. Booth syntheses show that the architecture is able to process QFHD frames in real time. The standard cells version is able to reach also a good trade-off among area, memory and power consumption, processing QFHD videos with 62.2 mW. Gustavo Sanchez, Luciano Volcan Agostini, Felipe Sampaio, Marcelo Schiavon Porto, Sergio Bampi |
ICME | 4 |
| 2010 | Gop structure adaptive to the video content for efficient H.264/AVC encodingabstractThis paper presents a new method for high efficiency video coding using an adaptive GOP structure based on video content for the H.264/AVC standard. The available H.264/AVC encoders typically use static GOP sizes that define how the frames I (Intra), P (Predictive) and B (Bi-predictive) are positioned during de coding process. However, by analyzing the video content it is possible to identify the optimum position for each type of frame inside the GOP. The proposed method analyses the video content and finds the best position for inserting I frames in the video sequence. Thus the GOP structure can assume different sizes, depending on the video content. The results for test sequences and real videos show that the proposed method can significantly reduce the required bit rate, comparing to the static GOP sizes, with reduced PSNR losses. The proposed adaptive GOP presents a gain, in terms of bit rate reduction for real movies, of 8.6%, 15%, 24.7% and 40.8% in comparison with static GOP sizes 32, 16, 8 and 4, respectively. Bruno Zatt, Marcelo Schiavon Porto, Jacob Scharcanski, Sergio Bampi |
ICIP | 2 |
| 2008 | A high throughput and low cost diamond search architecture for HDTV motion estimationabstractThis paper presents a high throughput and low cost architecture for motion estimation using a sub-sampled diamond search algorithm (SDS). The quality of SDS was compared with full search through software implementations and the results are presented. The designed hardware considered a search area of 100times100 samples, with blocks of 16times16 pixels. The architecture was described in VHDL and mapped to a Xilinx Virtex-4 FPGA. Synthesis results indicate that SDS is able to run at 185.7 MHz, using only 3541 LUTs. This architecture can reach real time for HDTV (1920times1080 pixels) in the worst case, and it can process 120 HDTV frames per second in the average case. Marcelo Schiavon Porto, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin |
ICME | 1 |
| 2007 | High Throughput Hardware Architecture for Motion Estimation with 4: 1 Pel Subsampling Targeting Digital Television Applications
Marcelo Schiavon Porto, Luciano Volcan Agostini, Leandro Rosa, Altamiro Amadeu Susin, Sergio Bampi |
PSIVT | 1 |