Cláudio Machado Diniz

dblp:27/1742 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2023 Multiversion Low-Power Hardware Accelerator for the AV1 Interpolation Filters
abstract
One of the new tools included in the AV1 video codec is the adaptive filtering scheme used in the sample interpolation process. This scheme includes three different filter families called Regular, Sharp and Smooth, offering high flexibility for motion estimation (ME) and motion compensation (MC). However, the high number of interpolation filters also leads to greater complexity and energy consumption, since the generation of samples at sub-pixel position is a costly process. This paper proposes a low-power and high-throughput hardware accelerator focused on the AV1 interpolation filters called Multiversion Interpolation Processor (MVIP). The accelerator includes the three AV1 interpolation filter families, with versions that employ operand isolation for power reduction in unused filters. The accelerator also includes a precise MVIP assuming the MC scenario, besides two approximate versions to reduce the cost on the ME scenario. The proposed design is able to process 8K video at 50fps in MC and 2,656.14 Msamples/sec in ME, with a power dissipation of 41.30mW.
Daiane Freitas, Mateus Grellert, Cláudio Machado Diniz, Guilherme Corrêa 0001
ISCAS3
2022 Improving Content-Aware Video Streaming in Congested Networks with In-Network Computing
abstract
Network congestion and packet loss pose an ever-increasing challenge to video streaming. Despite the research efforts toward making video encoding schemes resilient to lossy network conditions, forwarding devices have not considered monitoring packet content to prioritize packets and minimize the impact of packet loss on video transmission. In this work, we advocate in favor of in-network computing employing a packet drop algorithm and an in-network hardware module to devise a solution for improving content-aware video streaming in congested network. Results show that our approach can reduce intra-predicted packet loss by over 80% at negligible resource usage and performance costs.
Leonardo Gobatto, Mateus Saquetti, Cláudio Machado Diniz, Bruno Zatt, Weverton Luis da Costa Cordeiro, José Rodrigo Azambuja
ISCAS3
2022 Multiple Transform Selection Hardware Design for 4K@60fps Real-Time Versatile Video Coding
abstract
One of the main innovations introduced in the Versatile Video Coding (VVC) standard is the possibility to employ and combine different types of transforms for residual coding through a tool named as Multiple Transform Selection (MTS). This improved flexibility leads to a high computational cost, requiring efficient hardware designs for the transform module to achieve real-time processing. This work presents a dedicated hardware design for the MTS module. The architecture is capable of processing several block sizes and it implements all the allowed transform combinations of the MTS tool. The obtained results show that the architecture is capable of processing up to 4K@60fps videos in real time with a frequency of 279 MHz and a power dissipation of 583 mW. Also, when compared with related works, the proposed solution shows competitive results.
Bianca Silveira, Luiz Neto, Daniel Palomino 0001, Cláudio Machado Diniz, Guilherme Corrêa 0001
ISCAS4
2015 A deblocking filter hardware architecture for the high efficiency video coding standard
Cláudio Machado Diniz, Muhammad Shafique 0001, Felipe Vogel Dalcin, Sergio Bampi, Jörg Henkel
DATE1
2015 A Reconfigurable Hardware Architecture for Fractional Pixel Interpolation in High Efficiency Video Coding
abstract
We present a novel reconfigurable hardware architecture for interpolation filtering in high efficient video coding that adapts to run-time changes of the number of interpolation filter calls and thereby provides a high potential of energy efficiency. It employs a picture-based prediction scheme to estimate the number of interpolation filter calls at run-time by monitoring the group of pictures history based on video coding structure knowledge. Reconfigurable acceleration engines are developed that can adapt to different filter types. Dynamic composition of different instances of these engines enables different implementation versions with area versus throughput tradeoff. A run-time selection scheme determines the best implementation version for each picture based on the throughput requirements. Compared to state-of-the-art, our architecture reduces resource usage by 57% while supporting various throughputs and video resolutions.
Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 Run-time accelerator binding for tile-based mixed-grained reconfigurable architectures
abstract
Run-time mixed-grained reconfigurable architectures emerged as an efficient solution to deal with the heterogeneous and at-design-time unpredictable nature of advanced applications. Due to interconnection limitations, the reconfigurable elements are grouped into tiles communicating through an on-chip network. State-of-the-art run-time accelerator binding schemes, i.e., mapping the accelerators to elements in the physical reconfigurable array, do not deal with such tile-based architectures. We propose a new scheme for run-time accelerator binding into our tile-based mixed-grained reconfigurable architecture. By means of an advanced video encoding application, we illustrate that our scheme reduces the inter-tile communication overhead by up to 44% (avg. 23%).
Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
FPL1
2013 High-throughput interpolation hardware architecture with coarse-grained reconfigurable datapaths for HEVC
abstract
Fractional-pel interpolation for motion estimation and motion compensation is one of the key computational hotspots in the new High Efficient Video Coding (HEVC) standard. This work presents a high-throughput interpolation hardware architecture to improve performance of HEVC encoding and decoding. It employs two acceleration engines for luma and chroma filtering, each with 12-pel-parallel coarse-grained reconfigurable interpolation datapaths. An adaptive scheduling scheme manages the operation of these interpolation datapaths in different ways depending upon the prediction unit (PU) size and the execution scenario (i.e. motion estimation or motion compensation). We have implemented our hardware architecture in 150 nm technology. Compared to state-of-the-art techniques [12], our architecture required 49% less hardware area, while processing QFHD (3840×2160) resolution @ 30 fps.
Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
ICIP1
2012 Real-time block matching motion estimation onto GPGPU
abstract
This work presents an efficient method to map Motion Estimation (ME) algorithms onto General Purpose Graphic Processing Unit (GPGPU) architectures using CUDA programming model. Our method jointly exploits the massive parallelism available in current GPGPU devices and the parallelization potential of ME algorithms: Full Search (FS) and Diamond Search (DS). Our main goal is to evaluate the feasibility of achieving real-time high-definition video encoding performance running on GPUs. For comparison reasons, multi-core parallel and distributed versions of these algorithms were developed using OpenMP and MPI (Message Passing Interface) libraries, respectively. The CUDA-based solutions achieve the highest speed-up in comparison with OpenMP and MPI versions for both algorithms and, when compared to the state-of-the-art, our FS and DS solutions reach up to 18x and 11x speed-up, respectively.
Eduarda Monteiro, Marilena Maule, Felipe Sampaio, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi
ICIP4
2011 SHBS: A heuristic for fast inter mode decision of H.264/AVC standard targeting VLSI design
abstract
In the Rate-Distortion Optimization technique for H.264/AVC, the process of choosing the best mode is performed through exhaustive executions of the whole encoding process, which increases significantly the encoder complexity, sometimes even forbidding its use in real time video coding applications. In order to reduce the number of calculations necessary to determine the best inter-frame mode, this work proposes the SHBS (Stationarity, Heterogeneity and Border Strength) heuristic. The use of SHBS causes a reduction of 168 times in the encoding iterations, with a better PSNR, at the cost of a relatively small bit-rate increase. The SHBS heuristic was designed in hardware targeting FPGAs and this architecture achieved an operation frequency of 118 MHz, being able to process up to 438 HD 1080p frames per second.
Guilherme Corrêa 0001, Daniel Palomino 0001, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi
ICME3
2011 A high throughput H.264/AVC intra-frame encoding loop architecture for HD1080p
abstract
In this work we present a high throughput hardware architecture for the H.264/AVC intra-frame encoder exploiting the parallelism of intra prediction, forward and inverse transforms and quantization. Since there is a strong data dependency between the intra prediction and the image reconstruction loop, the latency of this path is a key design issue in order to provide high performance coding. Considering that 77% of the total intra-encoding computation is spent in these modules, our architecture handles a 4-pixel wide intra prediction module and a 16-pixel wide reconstruction loop. Compared to the state-of-the-art our approach reduces by 47% the number of cycles to process a macroblock. Running at 150 MHz our architecture guarantees encoding of 61 HD1080p frames per second. The developed architecture requires 73.4 MHz to real-time encode HD1080p, which is a 46% reduction of the frequency requirement compared to the state-of-the-art.
Cláudio Machado Diniz, Bruno Zatt, Cristiano Thiele, Altamiro Amadeu Susin, Sergio Bampi, Felipe Sampaio, Daniel Palomino 0001, Luciano Volcan Agostini
ISCAS1
2011 Applying CUDA Architecture to Accelerate Full Search Block Matching Algorithm for High Performance Motion Estimation in Video Encoding
abstract
This work presents a parallel GPU-based solution for the Motion Estimation (ME) process in a video encoding system. We propose a way to partition the steps of Full Search block matching algorithm in the CUDA architecture. A comparison among the performance achieved by this solution with a theoretical model and two other implementations (sequential and parallel using OpenMP library) is made as well. We obtained a O(n^2/log^2n) speed-up which fits the proposed theoretical model considering different search areas. It represents up to 600x gain compared to the serial implementation, and 66x compared to the parallel OpenMP implementation.
Eduarda Monteiro, Bruno Boessio Vizzotto, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi
SBAC-PAD3
2010 Timing and interface communication analysis of H.264/AVC encoder using SystemC model
abstract
This work presents a detailed timing and communication analysis for an H.264/AVC video encoder architecture using a SystemC model. The model was described using different abstraction levels in order to evaluate specific characteristics of each component module. The target encoder is defined to be able for H.264/AVC real-time encoding for 1080p video sequences at 30 fps and was modeled as a two-stage macro-pipeline system composed by eight component modules: Macroblock buffer, Intra- and Inter-Frame Predictors, Mode Decision, Forward and Inverse Transforms and Quantization, Reference Memory Write and Entropy Encoder (CAVLC). The bandwidth of each internal connection and of external memory interface was evaluated. The timing behavior and the data dependencies were characterized and summarized in a timing diagram in order to define design constraints and provide an accurate system specification when compared to a H.264/AVC encoder in the literature.
Bruno Zatt, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi
VLSI-SoC2
2009 A real time H.264/AVC intra frame prediction hardware architecture for HDTV 1080P video
abstract
This work presents an intra frame prediction hardware architecture for H.264/AVC baseline/main profile encoder which performs real time processing of HDTV 1080p videos. It is achieved by exploring the parallelism of intra prediction and by reducing the latency for Intra 4times4 processing, which is the intra encoding bottleneck. Synthesis results on Xilinx Virtex-II Pro FPGA and TSMC 0.18 mum standard-cells indicate that this architecture is able to real time encode HDTV 1080p video operating at 110 MHz. Our architecture can encode HD1080p, 720p and SD video in real time at a frequency 25% lower when compared to similar works.
Cláudio Machado Diniz, Bruno Zatt, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi
ICME1
2007 A Pipelined 8x8 2-D Forward DCT Hardware Architecture for H.264/AVC High Profile Encoder
Thaísa Leal da Silva, Cláudio Machado Diniz, João Alberto Vortmann, Luciano Volcan Agostini, Altamiro Amadeu Susin, Sergio Bampi
PSIVT2