VLDB 2026 Research / reviewers in the wild / expert
Daniel Palomino 0001
dblp:49/7554 · also Daniel Munari Palomino
· DBLP profile ↗
36ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-0409-8335ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Approximate SRAM-Based Memory System for Intra-Frame Prediction in VVC EncodersabstractThis paper introduces a memory system for the Versatile Video Coding (VVC) standard, aimed at optimizing energy consumption in intra-frame prediction through approximate storage techniques. We propose a system that employs data approximation due to dynamic voltage scaling Static Random Access Memory (SRAM) in two distinct buffers: the Neighbor Samples Buffer (NeighSB) and the Original Samples Buffer (OrigSB). Our findings indicate the feasibility of approximate storage exploitation of our memory system, with minor coding efficiency drops at controlled error rates. As NeighSB and OrigSB presented distinct resilience behaviors, we evaluate different approximation strengths to obtain the most suitable coding efficiency, enabling still effective energy savings. Finally, we investigate the energy savings achieved through different SRAM operation levels, yielding reductions of 42%-62%. Yasmin Camargo, Matheus Isquierdo, Daniel Palomino 0001, Bruno Zatt, Felipe Sampaio |
ISCAS | 3 |
| 2025 | Exploiting Approximate SRAM for Energy-Efficient Integer Motion Estimation on VVC EncodersabstractVideo coding is a critical technology for enabling many modern applications. However, the high complexity of state-of-the-art encoders leads to energy consumption challenges, particularly in memory systems. This paper proposes an integer motion estimation (IME) system using approximate SRAM memories to enhance the energy efficiency of VVC encoders. The approximate SRAM memories are employed for both the current block and the search area buffers by reducing the supply voltage. The proposed IME system is evaluated using a customized tool that simulates the approximation effects on both buffers based on real energy measurements from a 28nm SRAM. This tool is integrated with VVenC, a fast VVC implementation, to assess the impact on coding efficiency and energy consumption. Experimental results show that the IME system using approximate SRAM can reduce energy consumption in reading operations by up to 55%, with a low impact on coding efficiency. Matheus Isquierdo, Felipe Sampaio, Bruno Zatt, Nikil Dutt, Daniel Palomino 0001 |
ISCAS | 5 |
| 2025 | Analysis of Real-Time Hardware-based HEVC Encoders on GPU and Mobile PlatformsabstractRecent hardware encoders in GPUs and mobile SoCs enable real-time high-resolution video processing by constraining the supported encoding tools available in video coding standards. These constraints are also used to meet strict power, area, and memory limits of mobile platforms. This paper presents an analysis of the hardware-based High Efficiency Video Coding (HEVC) present in the high-performance NVIDIA NVENC within the RTX 4070Ti GPU, and the power-efficient encoder present in the Snapdragon 8 Gen 2 chip within the Samsung Galaxy S23+ smartphone. The analysis is performed in two perspectives: (1) the tool set constraints employed by each implementation are identified by a bitstream analysis on UHD encoded videos, and (2) a compression-efficiency evaluation of both encoders through a rate vs. distortion and Bjontegaard-Delta Rate (BD-Rate) analysis against the HEVC Test Model reference software. The results reveal the different design trade-offs between the platforms, offering valuable insights for hardware designers by highlighting the implementation choices of major industry players like NVIDIA and Qualcomm. Allan Schuch, Daniel Palomino 0001, Marcelo Schiavon Porto |
VCIP | 3 |
| 2025 | Improving Coding Efficiency of Massive Parallel Intra Prediction Using Alternative ReferencesabstractExploring massive parallelism is a common strategy to mitigate the processing time of modern video encoding standards. Nonetheless, data dependencies challenge parallelism exploitation, especially during intra prediction, where the reconstructed adjacent blocks are used as references. Some works use the original frame samples as references to decouple adjacent blocks and allow parallelism. Still, the original samples are static and cannot model the nuances of different bitrates. In this context, this work seeks to improve the coding efficiency of parallel intra prediction implementations by using alternative reference samples based on low-pass filters that better represent the nuances of different bitrates for any partitioning structure. Variations in multiple aspects of the filters are considered, such as their dimension and also the precision and distribution of their coefficients. Experimental evaluations assessed the similarity of such alternative samples when compared to the regular ones, in addition to their impacts on coding efficiency and the processing overhead required to obtain such samples. The results from such experiments demonstrate that the alternative references improve coding efficiency when compared to the original samples, especially at lower bitrates. Furthermore, the additional filtering stage poses negligible timing overhead in most computing systems. Iago Storch, Nuno Roma, Daniel Palomino 0001, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Fast Intra Mode Decision Using Machine Learning for the Versatile Video Coding StandardabstractThis paper presents a fast intra mode decision solution for the VVC standard using machine learning. The idea is to reorder the evaluation of modes performed by the Rate-Distortion Optimization (RDO) process according to the modes occurrence rate. Based on the new evaluation order, three Decision Tree models were trained to skip the modes less likely to be chosen. The results show that the proposed solution achieves time savings of up to 15.57% with coding efficiency degradation of only 0.41% on average. When compared with related works, the proposed solution shows competitive results. Adson Duarte, Bruno Zatt, Guilherme Corrêa 0001, Daniel Palomino 0001 |
ISCAS | 4 |
| 2022 | GM-RF: An AV1 Intra-Frame Fast Decision Based on Random ForestabstractThis paper presents the Grouping of Modes based on Random Forest (GM-RF), a fast decision algorithm for the AOMedia Video 1 (AV1) intra-frame prediction applying machine learning (ML). AV1 implements a wide variety of intra-frame prediction tools, significantly increasing the required computational effort. The GM-RF uses trained Random Forest (RF) models to reduce the number of intra-frame prediction modes evaluated for each encoded block. Experimental results show that the GM-RF achieves an average time savings of 50.19%, with a BD-BR of 7.41%. Compared with related works, GM-RF reached time savings from 5.6 to 10 times higher at a cost of a higher BDBR. To the best of the authors’ knowledge, this is the first solution in the literature using ML to reduce the AV1 intra-frame prediction computational effort. Pablo Rosa, Daniel Palomino 0001, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 2 |
| 2022 | Mode-Adaptive Subsampling of SAD/SSE Operations for Intra Prediction Cost ReductionabstractModern video encoders, such as the recently proposed AV1 and VVC, offer significant encoding gains at the cost of a corresponding increase of the computational effort. This is the case of the adopted intra prediction techniques, comprehending an increased number of prediction modes and range. To mitigate this computational cost, the presented work proposes a new mode-adaptive algorithm that significantly reduces the number of SAD/SSE operations during intra prediction, by generating an optimized subsampling pattern adaptive to each prediction mode. The method can be applied to any video codec and, when applied to AV1, it led to an encoding time reduction and BD-BR impact of 15.36% and 0.6%, respectively, or 7.97% and –0.02%, depending on the selected subsampling parameters. When implemented in hardware, the proposed technique provides an effective reduction as high as 75% of both the area and power on the modified distortion calculation module. Marcel Moscarelli Corrêa, Nuno Roma, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 3 |
| 2022 | Fast Affine Motion Estimation for VVC using Machine-Learning-Based Early Search TerminationabstractThe Affine Motion Estimation (AME) was introduced in the Versatile Video Coding (VVC) standard to allow for the detection of non-translational transformations during inter-frame prediction. Although providing important coding efficiency gains, this new tool represents 43% of the motion estimation (ME) complexity. However, an analysis over the AME step shows that the Affine motion vectors are often generated without resulting in the best ME prediction. This paper proposes a AME early search termination based on supervised machine learning. Six Random Forest models were trained with features obtained during the encoding process to accurately predict whether the AME step should be executed, partially executed or skipped, avoiding unnecessary calculations. As result, the proposed solution achieves an average time saving of 46.94% in the AME step with a coding efficiency loss of only 0.18%. Adson Duarte, Luciano Volcan Agostini, Bruno Zatt, Guilherme Corrêa 0001, Marcelo Schiavon Porto, Daniel Palomino 0001 |
ISCAS | 7 |
| 2022 | Multiple Transform Selection Hardware Design for 4K@60fps Real-Time Versatile Video CodingabstractOne of the main innovations introduced in the Versatile Video Coding (VVC) standard is the possibility to employ and combine different types of transforms for residual coding through a tool named as Multiple Transform Selection (MTS). This improved flexibility leads to a high computational cost, requiring efficient hardware designs for the transform module to achieve real-time processing. This work presents a dedicated hardware design for the MTS module. The architecture is capable of processing several block sizes and it implements all the allowed transform combinations of the MTS tool. The obtained results show that the architecture is capable of processing up to 4K@60fps videos in real time with a frequency of 279 MHz and a power dissipation of 583 mW. Also, when compared with related works, the proposed solution shows competitive results. Bianca Silveira, Luiz Neto, Daniel Palomino 0001, Cláudio Machado Diniz, Guilherme Corrêa 0001 |
ISCAS | 3 |
| 2022 | GPU-Acceleration of Affine Prediction in the Versatile Video CodingabstractThe Versatile Video Coding standard introduced a series of novel tools to improve the coding efficiency. However, these tools caused a massive increase in encoder computational complexity, and the affine prediction comprehends a significant share of this complexity. In this context, this work proposes an affine prediction modeling aiming at GPU implementation to accelerate the affine prediction by drawing the most parallelism out of such platforms. This modeling explores parallelism in two levels: conducting the prediction of multiple coding tree units simultaneously and breaking down the prediction into multiple highly-parallel stages. Since classical parallelization approaches for translational motion estimation are not very efficient for affine prediction, the proposed work explores the novel properties and parallelization possibilities introduced by affine prediction. Experimental results show that when applied to blocks 128x128, the proposed work can speed up the affine prediction by 57.21 times when compared to a fully sequential encoder, with a small coding efficiency penalty of 0.16% BD-BR. Iago Storch, Daniel Palomino 0001, Sergio Bampi |
ISCAS | 2 |
| 2022 | FastInter360: A Fast Inter Mode Decision for HEVC 360 Video CodingabstractThis paper presents FastInter360, a fast inter mode decision algorithm for accelerating the encoding of ERP 360 videos. The development of FastInter360 involves an in-depth and comprehensive set of evaluations performed to understand the differences in the encoder’s behavior when encoding 360 and conventional videos. These evaluations showed that due to the texture distortions resulting from projection, the encoder presents a specific behavior when encoding 360 videos, making it more likely to use a recurrent set of encoding modes when processing 360 videos. Besides, the coding efficiency is less sensible to approximations in some encoding steps depending on the frame region. FastInter360 is then proposed to reduce the encoding complexity by exploiting these differences. FastInter360 comprises three algorithms that accelerate the encoding by performing early decision by SKIP mode, reducing integer motion estimation search range, and adjusting fractional motion estimation precision. Furthermore, each of these algorithms behaves according to distortion intensity, performing greater complexity reduction in more distorted regions. When employed altogether, these algorithms compose FastInter360, which is able to achieve an average complexity reduction of 22.84% with a coding efficiency loss of 0.652% BD-BR, on average, making FastInter360 competitive with literature works. Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | RDE-MOGA: Automatic Selection of Rate-Distortion-Energy Control Points for Video Encoders Using Muti-Objetive Genetic AlgorithmabstractControlling energy consumption of video encoders is a complex multi-objective optimization problem of great importance. In this work we propose the RDE-MOGA, an multi-objective genetic algorithm capable of finding energetically efficient configurations for the HEVC encoder and replacing the current sensitivity analysis methodologies in the development of energy controllers. The utilization of our algorithm improved its efficiency in 60% whereas increasing the range of achievable reductions of the controller in at least 50%. Furthermore, the algorithm proved capable of sustaining 30% energy reduction at a cost of 3.45 BD-BR loss. Italo Machado, Marilton S. de Aguiar, Marcelo Schiavon Porto, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ICASSP | 5 |
| 2020 | Spatially Adaptive Intra Mode Pre-Selection for ERP 360 Video CodingabstractIn this work, we propose a spatially adaptive HEVC intra mode pre-selection for equirectangular (ERP) 360 video coding. The proposed technique exploits the spatial characteristics of 360 video in the ERP projection to reduce the complexity of intra prediction mode selection. The number of intra modes evaluated in Rate-Distortion Optimization is reduced based on a score technique that is adaptive to the frame region being encoded. Results show that the proposed technique achieves a complexity reduction of 16.5% with low coding efficiency penalties. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Guilherme Corrêa 0001, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
ICASSP | 6 |
| 2020 | ASIC Solution for the Directional Intra Prediction of the AV1 Encoder Targeting UHD 4K VideosabstractAOMedia Video 1 (AV1), developed by the Alliance for Open Media consortium and released in 2018, is an open-source and royalty-free video format. It was designed to deliver substantial compression gains over its predecessor VP9 whilst keeping hardware feasibility and a practical decoding complexity. When compared to state-of-the-art formats, AV1 has more complex encoder tools, including the intra prediction which is the focus of this work. This paper presents a highly parallelized ASIC solution for the directional intra prediction module. It supports all 56 directional modes defined in AV1 and is able to process all combinations of block partitions. When synthesized to the TSMC 40nm technology with a target frequency of 1,296MHz, the proposed design used an area of 455.8K gates and showed a power dissipation and energy consumption per predicted sample of 40.92mW and 0.055pJ/sample, respectively. The reached throughput supports the processing of 60 frames per second for UHD 4K videos (3840×2160 pixels). No other work was found in the literature with a hardware design supporting the AV1 intra prediction directional modes. Marcel Moscarelli Corrêa, Luiz Neto, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 3 |
| 2020 | Low-Power and Memory-Aware Approximate Hardware Architecture for Fractional Motion Estimation Interpolation on HEVCabstractNowadays, current video coding standards like the High Efficiency Video Coding (HEVC) implement several complex coding tools, like the Fractional Motion Estimation (FME). An alternative to improve performance and save power is allying the hardware acceleration with approximate computing solutions, focusing on such complex tools. In this work, we present a low-power and memory-aware hardware architecture for the HEVC FME interpolator, proposing the development of two novel hardware designs for the interpolation filters, called Approximate Unified FME Filters (AUFF). These solutions exploit the usage of approximate computing at both algorithmic and data levels, leading to a reduction in dissipated power and memory bandwidth. The proposed design is capable of real-time interpolation of UHD (Ultra High Definition) 4K and 8K videos when synthesized using a 40 nm standard-cell library, with a power dissipation ranging from 22.04 to 62.06 mW. Wagner Penny, Guilherme Corrêa 0001, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Gabriel L. Nazar, Bruno Zatt |
ISCAS | 4 |
| 2020 | Efficient Hardware Design for the AV1 CDEF Filter Targeting 4K UHD VideosabstractDeveloped by the AOMedia industry consortium, the AOM Video 1 (AV1) is an open-source and royalty-free video encoder released in June 2018. The Constrained Directional Enhancement Filter (CDEF) is one of the three AV1 in-loop filters and it is the focus of this work. The CDEF has the goal to reduce ringing artifacts generated with the encoding process, acting as a directional deringing filter. This paper presents a hardware design for the AV1 CDEF targeting real-time processing of 4K Ultra High Definition (UHD) videos. The architecture was synthesized to ASIC using the 40nm TSMC library, requiring 185 kgates and with a power dissipation of 43 mW when running at 93 MHz, reaching the frame rate of 60 frames per second (fps). To the best of the author's knowledge, there is no other work in the literature with dedicated hardware design for the AV1 CDEF. Eduardo Zummach, Roberta Palau, Jones Goebel, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 4 |
| 2020 | ERP-Based CTU Splitting Early Termination for Intra Prediction of 360 videosabstractThis work presents an Equirectangular projection (ERP) based Coding Tree Unit (CTU) splitting early termination algorithm for the High Efficiency Video Coding (HEVC) intra prediction of 360-degree videos. The proposed algorithm adaptively employs early termination in the HEVC CTU splitting based on distortion properties of the ERP projection, that generate homogeneous regions at the top and bottom portion of a video frame. Experimental results show an average of 24% time saving with 0.11% coding efficiency loss, significantly reducing the encoding complexity with minor impacts in the encoding efficiency. Besides, solution presents the best results considering the relation between time saving and coding efficiency when compared with all related works. Bernardo Beling, Iago Storch, Luciano Volcan Agostini, Bruno Zatt, Sergio Bampi, Daniel Palomino 0001 |
VCIP | 6 |
| 2020 | 4D-DCT Hardware Architecture for JPEG Pleno Light Field CodingabstractThis paper presents a 4D-DCT hardware architecture for Light Field Coding according to the JPEG Pleno standard. It is composed of two instances of 2D-DCT engines and a novel 4D Transposition Memory organization. Experimentally-defined fixed-point representation and LSB pruning techniques are employed do reduce hardware area and power dissipation. The proposed architecture operates over 4D-hypercubes of up to 8x8x8x8 samples and reaches performance to process 30 Lytro-like light fields per second dissipating 145.32mW at 825.75MHz. This is the first known 4D-DCT hardware architecture for light field coding and demonstrates the feasibility of such solutions on real-world systems. Matheus Jahnke, Jones Goebel, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 3 |
| 2020 | Power/QoS-Adaptive HEVC FME Hardware using Machine Learning-Based Approximation ControlabstractThis paper presents a machine learning-based adaptive approximate hardware design targeting the fractional motion estimation (FME) of HEVC encoder. Hardware designs targeting multiple levels of approximation are proposed, by changing FME filters coefficients and/or discarding taps. The level of approximation is defined by a decision tree, generated taking into account the behavior of several parameters of the encoding in order to predict homogeneous blocks, more suitable for more aggressive approximation without significant losses on quality of service (QoS). Instead of applying a specific level of approximation over the full video, different approximate FME accelerators are dynamically selected. Such a strategy is able to provide up to 50.54% of power reduction while keeping the QoS losses at 1.18% BD-BR. Wagner Penny, Daniel Palomino 0001, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 2 |
| 2019 | Online Machine Learning for Fast Coding Unit Decisions in HEVCabstractThe High Efficiency Video Coding standard introduced a flexible frame partitioning process that increased significantly compression rates in comparison to previous standards at the cost of a high computational cost. To accelerate frame partitioning decisions, this paper proposes a method that replaces the usual Rate-Distortion Optimization employed in Coding Unit size decision by a set of simpler decision tree models, which are built during encoding time by the C5 machine learning algorithm. The algorithm and the set of attributes employed in the model training process were chosen based on an extensive analysis that compared several options in terms of decision accuracy and training complexity. Experimental results show that the proposed method is capable of building accurate models for each video sequence, decreasing the HEVC encoding complexity in 34.4% on average with a compression efficiency loss of only 0.2% in comparison to the original HEVC reference encoder. Guilherme Corrêa 0001, Pargles Dall'Oglio, Daniel Palomino 0001, Luciano Volcan Agostini |
DCC | 3 |
| 2019 | FastIntra360: A Fast Intra-Prediction Technique for 360-Degrees Video Codingabstract360-degrees videos represent a whole sphere and enable the user to feel as if he is inside the scene. These videos demand more data than conventional videos to be represented, therefore they also must be compressed to be handled properly. However, current video coding standards only process rectangular videos, thus 360 videos must be represented in a flat fashion to be encoded. There are several projections to perform this and the currently most used one is the equirectangular projection (ERP), which transforms each parallel from the sphere into a row of the rectangle, resulting in a faithful representation of the equatorial area, and a stretched representation of the polar regions. This stretching in the polar regions tends to impact the behavior of intra-frame prediction, which is used to exploit the spatial redundancies in each frame. Therefore, this paper proposes FastIntra360 to accelerate the encoding of 360 videos. FastIntra360 is implemented in HEVC video coding standard [1], which is a recently established standard and poses high computational demand. During the development of FastIntra360, a set of videos were encoded and the behavior of the intra-prediction throughout the frame was extracted. Then, a statistical analysis was conducted over such data and it concluded that when encoding the polar regions of the frame, the prediction modes which exploit horizontal directions are selected more frequently than the remaining modes, whereas in the center of the frame all prediction modes present similar occurrence rates. FastIntra360 exploits this behavior to reduce the number of prediction modes evaluated in different regions of the frame to accelerate the encoding. FastIntra360 is developed in two variants: one considering three bands and other considering five bands, where each band is a horizontal stripe of the frame. Each band divides the frame samples into three or five stripes and performs the statistical analysis over these stripes individually. Both implementations were evaluated and compared against the HEVC Test Model version 16.16 (HM-16.16) according to time reduction and coding efficiency (considering BD-BR), where BD-BR represents the bitrate increase of the proposed technique. Experimental results showed that both implementations present good performance, reaching up to 16.5% complexity reduction with negligible BD-BR, that is, they present considerable complexity reduction whereas posing no harm to the video quality. Iago Storch, Bruno Zatt, Luciano Volcan Agostini, Luís Alberto da Silva Cruz, Daniel Palomino 0001 |
DCC | 5 |
| 2019 | Encoding Efficiency and Computational Cost Assessment of State-Of-The-Art Point Cloud CodecsabstractPoint clouds have recently emerged as a suitable solution to generate and display 3D digital models due to their capacity of representing high resolution images and videos through multiple viewpoints. However, as they are usually made up of thousands up to billions of points, advanced techniques of data compression are essential to store and transmit this type of data. This paper compares the two state-of-the-art solutions for point cloud compression, the Point Cloud Codec (PCC) and the Test Model Category 2 (TMC2), in terms of compression efficiency and encoding time. Experimental results show that the compression efficiency for geometry information is highly dependent upon the available bitrate for both TMC2 and PCC. However, for texture compression TMC2 almost always achieves the best results. The experiments have also shown that TMC2 presents a computational cost from 22.2 to 26 times larger the observed in PCC. Mateus M. Gonçalves, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Guilherme Corrêa 0001 |
ICIP | 3 |
| 2019 | High Throughput Hardware Design for AV1 Paeth and Smooth Intra ModesabstractDeveloped by AOMedia industry consortium and released in June 2018, AV1 is an open-source and royalty-free video coding format. The main goal of AV1 is to deliver substantial compression gains over state-of-the-art codecs such as VP9 and HEVC, while keeping a practical decoding complexity, hardware feasibility and its open and free status. This paper presents a high throughput hardware architecture for four important AV1 intra prediction coding modes: Paeth, Smooth, Smooth Vertical and Smooth Horizontal. The proposed architecture was designed to support all 19 block sizes specified by AV1 and to process every single combination of these blocks according to the 10-way partition tree, with a throughput of UHD 4K (3840×2160 pixels) videos at up to 30 frames per second. When synthesized to the TSMC 40nm cell library targeting a frequency of 648MHz, the proposed design used 109.57K gates and showed a power dissipation and an energy efficiency of 16.1mW and 1.23pJ/sample respectively. No other works were found in the literature describing hardware designs for AV1 intra prediction. Marcel Moscarelli Corrêa, Bianca Waskow, Bruno Zatt, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini |
ISCAS | 4 |
| 2018 | High-Throughput and Low-Power Integrated Direct/Inverse HEVC Quantization Hardware DesignabstractThis paper presents a high-throughput and low-power integrated HEVC direct/inverse quantization hardware design. The main focus of this design is to allow the evaluation of multiple coding modes during the residual encoding process of the HEVC for real-time Ultra-High Definition (UHD) video processing. The ASIC synthesis results, for a Nangate 45nm standard cell library, presented a maximum operational frequency of 1679.51MHz and a processing rate of 53.74 Gsps (giga samples per second). This throughput allows processing of real-time up to 72 coding modes for UHD 4K@60fps or up to nine coding modes for the UHD 8K@120fps while dissipating 369.37mW. Luciano Almeida Braatz, Bruno Zatt, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 3 |
| 2018 | OTED: Encoding Optimization Technique Targeting Energy-Efficient HEVC DecodingabstractThis work exploits the encoding for decoding concept by proposing an encoding optimization technique targeting energy-efficient HEVC decoding (called OTED). OTED changes the Rate-Distortion Optimization (RDO) calculation at the encoder side by adding the decoding energy estimation as a new variable to be considered. Along with this estimation, we use the Running Average Power Limit (RAPL) energy measurement tool to present real energy results and prove the efficiency of OTED. Experimental results show that the proposed algorithm can reach an energy reduction of up to 17.7% at the HEVC decoder with a small cost in coding efficiency. Besides being compliant with the standard HEVC decoder, the algorithm presents high levels of energy reduction for different encoding configurations. Furthermore, when compared with state-of-the-art works OTED presents the best relation between energy reduction consumption at the encoder side and coding efficiency. Douglas Corrêa, Guilherme Corrêa 0001, Daniel Palomino 0001, Bruno Zatt |
ISCAS | 3 |
| 2018 | Configurable Cache Memory Architecture for Low-Energy Motion EstimationabstractThe popularization of mobile devices and the increased demand for video applications from these devices necessitates the design of efficient video encoders such as HEVC. Since the Motion Estimation (ME) is the most processing and memory intensive unit in a video encoder, our focus is in the communication between external memory and the ME unit. The TZS algorithm is widely used in video encoders and has an unpredictable behavior, which leads to an unknown pattern of memory accesses, making SPMs ineffective solutions, for example. Therefore, this work proposes a configurable cache memory architecture for fast ME algorithms. This cache has settings that suit different video encoding scenarios. Six optimal cache configurations were defined based on our evaluation considering 23 video sequences, 4 QPs, and 32 different cache settings. External memory bandwidth savings of up to 96.84% were reached, representing a reduction from 25.48GB/s to 548.53MB/s in the best case. When compared to Level-C SPM and to a static 16KB 8-way associative cache, the proposed configurable cache achieves energy savings of up to 86.91% and 78.09%, respectively. Anderson Martins, Wagner Penny, Matheus Weber, Luciano Volcan Agostini, Marcelo Schiavon Porto, Daniel Palomino 0001, Júlio C. B. de Mattos, Bruno Zatt |
ISCAS | 6 |
| 2016 | Thermal optimization using adaptive approximate computing for video coding
Daniel Palomino 0001, Muhammad Shafique 0001, Altamiro Amadeu Susin, Jörg Henkel |
DATE | 1 |
| 2016 | Speedup-aware history-based tiling algorithm for the HEVC standardabstractThis paper proposes a history-based tiling algorithm aiming at the increase of speedup when using Tiles. The algorithm is composed of two independent steps that use workload history information to define the vertical and horizontal boundaries of the Tiles. The workload distribution of previous frames are used as reference to perform the tiling of the current frame exploiting the temporal similarity between neighboring frames. Experimental results show that the proposed algorithm outperforms the speedup when compared to uniform tiling by 6.85% on average for tested sequences, besides, there is no significant complexity increase and similar coding efficiency results. When compared to related works the proposed solution also presents better speedup results. Iago Storch, Daniel Palomino 0001, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 2 |
| 2014 | hevcDTM: Application-driven Dynamic Thermal Management for High Efficiency Video CodingabstractThis paper presents an application-driven algorithm for Dynamic Thermal Management (DTM) for the High Efficiency Video Coding (HEVC). For efficient design of such a DTM policy, we perform an offline thermal analysis of an HEVC encoder and demonstrate the impact of different video sequences and different coding configurations on the processor temperature. Our thermal analysis is leveraged to develop an efficient application-driven DTM policy that performs temperature-aware coding along with an application-driven control of DTM knobs (e.g., frequency scaling) in order to meet the temperature constraints while still providing high video quality (i.e. PSNR loss <; 0.01dB). For accurate thermal analysis and evaluation, we deploy an infrared camera-based thermal measurement setup that, on the contrary to state-of-the-art setups, does not require adding any extra layer on top of the measured chip, thus allowing the camera to accurately capture the infrared emissions from the die. Daniel Palomino 0001, Muhammad Shafique 0001, Hussam Amrouch, Altamiro Amadeu Susin, Jörg Henkel |
DATE | 1 |
| 2014 | TONE: adaptive temperature optimization for the next generation video encodersabstractThis paper presents an adaptive temperature optimization technique for the next generation video encoders. It exploits both application-specific knowledge (i.e. video encoding configurations) and video content properties in order to efficiently manage the temperature of advanced video coding systems at the software layer. For designing an efficient technique, we perform an extensive offline analysis to understand the impact of different video properties and configurations on the CPU thermal profiles when processing the next generation video encoder. Our temperature optimization technique performs an application-level prediction of the temperature trend followed by an application-level thermal management policy. The policy dynamically manages the temperature by performing an adaptive encoder configuration selection while providing minimum penalties in terms of bit rate and video quality. The experimental results show that our policy meets temperature constraints with negligible encoding performance loss. Moreover, when compared to state-of-the-art techniques, our policy provides a relatively reduced video quality loss while still meeting the temperature constraints. Daniel Palomino 0001, Muhammad Shafique 0001, Altamiro Amadeu Susin, Jörg Henkel |
ISLPED | 1 |
| 2013 | Adaptive content-based Tile partitioning algorithm for the HEVC standardabstractThis paper proposes a content-based Tile partitioning algorithm designed to exploit video properties in the tiling process, aiming at the reduction of the coding losses generated by the use of Tiles. In the proposed algorithm two steps are performed to define the vertical and horizontal Tile boundaries willing to group the high correlated samples into the same Tile partition. The definition of the Tiles' boundaries is performed by analyzing the variance map extracted from the raw picture information. Experimental results have shown that the proposed algorithm is able to reduce the inherent coding efficiency losses of using Tiles when compared to the conventional uniform spaced Tile partitions. Cauane Blumenberg, Daniel Palomino 0001, Sergio Bampi, Bruno Zatt |
PCS | 2 |
| 2013 | Fast HEVC intra mode decision algorithm based on new evaluation order in the Coding Tree BlockabstractThis paper presents a fast mode decision algorithm for the HEVC intra prediction. A new evaluation order in the Coding Tree Block (CTB) allows the use of modes from low level PUs to be used as reference to the current PU decision. In this paper we use this idea to develop a fast intra mode decision algorithm that can be configured to run in two different complexity modes, relaxed and aggressive. Experimental results have shown that our algorithm achieved encoding time savings of almost 60% with negligible loss in the compression efficiency when compared to the full RDO based decision. Besides, our mode decision algorithm presented the best result in terms of time saving per compression efficiency when compared with all related works. Daniel Palomino 0001, Eduardo Cavichioli, Altamiro Amadeu Susin, Luciano Volcan Agostini, Muhammad Shafique 0001, Jörg Henkel |
PCS | 1 |
| 2012 | A memory aware and multiplierless VLSI architecture for the complete Intra Prediction of the HEVC emerging standardabstractThis work proposes a hardware architecture for the Intra Frame Prediction of the emerging High Efficiency Video Coding (HEVC) standard. The architecture was designed considering all innovative features of the Intra Prediction included in the HEVC, i.e. all modes and all Prediction Units (PU) sizes. Performance and memory accesses are a problem in the HEVC intra prediction and hardware architecture designs are good alternative to solve these issues, especially when energy-efficient solutions are targeted. Buffers and internal memories were used in the designed architecture to decrease the number of external memory accesses. Two independent data paths processing eight samples in parallel and a deep and multiplierless pipeline were designed to increase the throughput. The architecture was synthesized using an IBM 65nm CMOS technology. The results have shown that the architecture is able to process 30 HD720p frames per second and 13 HD1080p frames per second when running at 500 MHz, reducing in 95% the accesses to the external memory. Daniel Palomino 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin |
ICIP | 1 |
| 2011 | SHBS: A heuristic for fast inter mode decision of H.264/AVC standard targeting VLSI designabstractIn the Rate-Distortion Optimization technique for H.264/AVC, the process of choosing the best mode is performed through exhaustive executions of the whole encoding process, which increases significantly the encoder complexity, sometimes even forbidding its use in real time video coding applications. In order to reduce the number of calculations necessary to determine the best inter-frame mode, this work proposes the SHBS (Stationarity, Heterogeneity and Border Strength) heuristic. The use of SHBS causes a reduction of 168 times in the encoding iterations, with a better PSNR, at the cost of a relatively small bit-rate increase. The SHBS heuristic was designed in hardware targeting FPGAs and this architecture achieved an operation frequency of 118 MHz, being able to process up to 438 HD 1080p frames per second. Guilherme Corrêa 0001, Daniel Palomino 0001, Cláudio Machado Diniz, Luciano Volcan Agostini, Sergio Bampi |
ICME | 2 |
| 2011 | A high throughput H.264/AVC intra-frame encoding loop architecture for HD1080pabstractIn this work we present a high throughput hardware architecture for the H.264/AVC intra-frame encoder exploiting the parallelism of intra prediction, forward and inverse transforms and quantization. Since there is a strong data dependency between the intra prediction and the image reconstruction loop, the latency of this path is a key design issue in order to provide high performance coding. Considering that 77% of the total intra-encoding computation is spent in these modules, our architecture handles a 4-pixel wide intra prediction module and a 16-pixel wide reconstruction loop. Compared to the state-of-the-art our approach reduces by 47% the number of cycles to process a macroblock. Running at 150 MHz our architecture guarantees encoding of 61 HD1080p frames per second. The developed architecture requires 73.4 MHz to real-time encode HD1080p, which is a 46% reduction of the frequency requirement compared to the state-of-the-art. Cláudio Machado Diniz, Bruno Zatt, Cristiano Thiele, Altamiro Amadeu Susin, Sergio Bampi, Felipe Sampaio, Daniel Palomino 0001, Luciano Volcan Agostini |
ISCAS | 7 |
| 2009 | Low latency and high throughput dedicated loop of transforms and quantization focusing in the H.264/AVC Intra PredictionabstractThis paper presents an efficient architectural design for a dedicated transforms and quantization loop. This design targeted the Intra Prediction of the H.264/AVC standard. The architecture was designed intending to achieve the best possible relation between throughput, latency and hardware resources consumption. The latency and throughput of this loop are extremely important to define the intra prediction performance. The use of hardware was reduced through the reuse of the same datapath for different calculations. The architecture was synthesized to Altera Stratix III FPGA and to the TSMC 0.18 ¿m standard-cells technology. The architecture, when mapped to standard-cells, reaches a processing rate of 114 HDTV frames per second, attending the intra prediction restrictions. Daniel Palomino 0001, Felipe Sampaio, Robson Dornelles, Luciano Volcan Agostini |
ICIP | 1 |