EDBT 2026 Demo / reviewers in the wild / expert
Mateus Grellert
dblp:29/9850 · also Mateus Grellert da Silva
· DBLP profile ↗
33ranked-venue papers
6as first author
20since 2021 · last 2025
0000-0003-0600-7054ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Interaction Detection in Images of Therapy Sessions with Children with Autism Spectrum DisorderabstractAutism spectrum disorder (ASD) affects many children and limits their social interaction, communication, and behavioral skills. Regular follow-ups with qualified professionals assist in the patients' development through sessions and progress evaluations. This progress is manually recorded by professionals, which can lead to errors in analysis. To assist with these annotations, an automation process is proposed using pose estimation and object detection techniques in computer vision to detect interaction between participants in therapy sessions with children with ASD. For this purpose, the Yolov8 object detector and Yolov8-Pose estimator are applied. Subsequently, interaction detection is performed using the results obtained from the previous predictions. Heuristics were introduced to compare these techniques. To enhance performance, solutions were explored to address the limitations encountered in object detection and pose estimation predictions. The results demonstrated a 15.8% improvement in precision for the pose estimation heuristic compared to the bounding box heuristic, achieving satisfactory performance for the proposed approach. Brenda Caroline Santos Mendes, Jônata Tyska Carvalho, Mateus Grellert |
CBMS | 3 |
| 2025 | Cross-Layer Approximate Hardware Design of Interpolation Filters for Fractional Motion Estimation in Versatile Video CodingabstractThe fact that we can stream video on multiple devices in our homes, on the go using mobile devices, or even while video chatting across the globe, even with low bandwidth, is owed to video coding. Versatile Video Coding (VVC) is the latest video coding standard, which introduces a range of innovative tools. One example is adopting an alternative fractional motion estimation (FME) filter that is part of the Advanced Motion Vector Resolution (AMVR) extension. This work introduces a lowpower hardware architecture accelerator specifically designed for Fractional Motion Estimation (FME) with support for the AMVR extension of VVC. Rafael da Silva, Ricardo Augusto da Luz Reis, Mateus Grellert |
VLSI-SoC | 3 |
| 2025 | An LSTM approach to predict emergency events using spatial features
Felipe Vieira 0001, Antônio Augusto Fröhlich, Mateus Grellert |
Appl. Intell. | 3 |
| 2025 | Cross-Layer Approximate Design of Low-Power Fractional Motion Estimation Accelerators for VVCabstractThe versatile video coding (VVC) standard introduces several innovative tools designed to enhance coding efficiency compared with its predecessors. One example is the adoption of an alternative filter for fractional motion estimation (FME) that is part of the adaptive motion vector resolution (AMVR) extension. While this allows a more precise motion representation, it also incurs more complexity for hardware implementations that aim at supporting most of VVC features. This work introduces a low-power hardware architecture accelerator specifically designed for FME with support for the AMVR extension of VVC. The proposed solution enables a systematic exploration of the design space of cross-layer approximate computing by combining approximations at both the operator and algorithm levels. This is achieved through the design of two novel architectures (2TAxA/4Tand2TAxA/2TAxA) alongside a newly proposed approximate filter applicable to both regular and alternative interpolation modes. Furthermore, we evaluate eight different approximate adder (AA) topologies to optimize power–quality tradeoffs. Experimental results demonstrate that for a complete FME multifilter interpolation unit (MIU) and maintaining an image quality threshold of$\text {SSIM} \geq 0.88$, our method achieves up to 72% power savings and 59.64% area savings. Rafael da Silva, Pedro Tauã Lopes Pereira, Mateus Grellert, Ricardo Augusto da Luz Reis |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Stimming Behavior Dataset - Unifying Stereotype Behavior Dataset in the WildabstractNon-intrusive vision-assisted methods that can recognize complex human behaviors can help the diagnosis and treatment of neurological disorders where stimming behaviors are prominent, such as Autism Spectrum Disorder (ASD). Machine learning methods, especially those related to computer vision are a promising direction for this kind of application, however, the effectiveness of these methods depends on the existence of good-quality datasets. Building a dataset in this domain is challenging due to privacy matters and also to the need to conduct experiments mostly with children. Therefore, we propose a consolidated, annotated, and standardized Stimming Behavior Dataset. This unified dataset can enable researchers to have access to a larger pool of data, classified and annotated for the stimming behaviors, enabling a faster and more streamlined approach when working with this kind of video collection. The Autism Stimming Behavior Dataset - ASBD11https://github.com/OckerGui/Stimming-Behavior-Dataset can be used for the development of machine learning algorithms that can detect behavioral patterns associated with autism spectrum disorder and other applications where this kind of analysis is important. The dataset is comprised of 165 annotated short duration clips, based on excerpts from 155 distinct publicly accessible Youtube videos., with 154 unique subjects of ages spanning from toddlers to young adults and divided into 48 female and 106 male individuals. The average duration of the clips is 10.6 seconds. The annotated clips show not only the URL and class of the video but the moment where the relevant behavior starts and its duration. In addition to providing a unified dataset, this paper also includes predictions of each category based on a classic I3D model trained in the Kinetics 400 dataset, which serves as a baseline for the action-recognition methods and analysis on the previous source datasets and models, providing insights for future action-recognition-based systems that can be used to assist in decisions regarding ASD diagnosis and treatment. Guilherme Ocker Ribeiro, Mateus Grellert, Jônata Tyska Carvalho |
CBMS | 2 |
| 2023 | Hypersaline Tidal Flats Detection Using Deep Learning Over 37 Years of Landsat DataabstractHypersaline tidal flats are plane areas usually related to mangrove forests, acting as guard and buffer against rising sea levels, and as maintainer of regional biodiversity. Such areas are primarily impacted by anthropogenic and natural activities, such as sea-salt extraction and pollution, so identifying and monitoring them is an important and challenging task. The present work uses a U-shaped Convolutional Neural Network architecture to systematically classify such formations over Landsat imagery. A large dataset containing data from 1985 to 2021 of the Brazilian Coastal Zone is used to train and evaluate our model. Experimental results show that the total area increased by 58.6 km2from 1985 to 2001, and decreased by approximately 92 km2from 2001 to 2021, representing a total reduction of ≈ 33.34 km2for the entire period. We also show that our model outperforms a related solution trained with the same dataset, achieving 70% and 86% for 1985 and 2020 respectively, against 69% and 82%. Maria Luize Pinheiro, Luiz Cortinhas, Cesar Guerreiro Diniz, Raian Vargas Maretto, Mateus Grellert |
IGARSS | 5 |
| 2023 | Multiversion Low-Power Hardware Accelerator for the AV1 Interpolation FiltersabstractOne of the new tools included in the AV1 video codec is the adaptive filtering scheme used in the sample interpolation process. This scheme includes three different filter families called Regular, Sharp and Smooth, offering high flexibility for motion estimation (ME) and motion compensation (MC). However, the high number of interpolation filters also leads to greater complexity and energy consumption, since the generation of samples at sub-pixel position is a costly process. This paper proposes a low-power and high-throughput hardware accelerator focused on the AV1 interpolation filters called Multiversion Interpolation Processor (MVIP). The accelerator includes the three AV1 interpolation filter families, with versions that employ operand isolation for power reduction in unused filters. The accelerator also includes a precise MVIP assuming the MC scenario, besides two approximate versions to reduce the cost on the ME scenario. The proposed design is able to process 8K video at 50fps in MC and 2,656.14 Msamples/sec in ME, with a power dissipation of 41.30mW. Daiane Freitas, Mateus Grellert, Cláudio Machado Diniz, Guilherme Corrêa 0001 |
ISCAS | 2 |
| 2022 | Low-Complexity Multi-Type Tree Partitioning for Versatile Video Coding Based on Machine LearningabstractThe Versatile Video Coding (VVC) standard introduces new types of frame partitioning structures, such as QuadTrees (QT) and Multi-Type Trees (MTT). To achieve the best compression efficiency, for each block of pixels the encoder performs a recursive search over the partitioning possibilities, which also impacts significantly on the encoding complexity and processing time. This work proposes a machine learning-based solution for quick block partitioning decisions. A set of fourteen Random Forests were trained using data gathered during the encoding process and the models were employed to decide whether vertical and horizontal partitions are required for each candidate block. The proposed solution leads to an average encoding time reduction of 34.76% at the cost of a compression efficiency loss of 1.03%. Matheus Lindino, Bruno Zatt, Mateus Grellert, Guilherme Corrêa 0001 |
ICIP | 3 |
| 2022 | Exploring Approximate Comparator Circuits on Power Efficient Design of Decision TreesabstractIn recent years, Approximate Computing has been gaining space as a technique for tackling energy requirements in error-resilient applications, while the usage of Machine Learning systems has steadily increased. This work explores two different approaches for approximation in comparator circuits and their impact on Decision Tree applications, observing power and accuracy metrics. Two gate-level architectures are proposed for dedicated comparators, approximating 25% or 50% of least significant bits using different techniques. The circuits were described in 7 nm FinFET technology. The approximate comparators were then evaluated in a Decision Tree classification model using five continuous and mixed attribute datasets. The 25% LSB approximate comparator proposed improves the energy efficiency in Decision Tree applications, reducing from 12% up to 84% the power per inference while presenting minor deviations in accuracy compared to the exact baseline. Pedro Aquino Silva, Mateus Grellert, Cristina Meinhardt |
VLSI-SoC | 2 |
| 2022 | Approximation Workflow for Energy-Efficient Comparators in Decision Tree ApplicationsabstractThe increasing use of Machine Learning applications has caused a high demand for design techniques targeting the trade-off between energy consumption and accuracy in inference models. In this scenario, Decision Trees are models widely used in embedded systems and Internet of Things applications. They are less resource intensive than Neural Networks, and maintain acceptable accuracy results for various problems. This project proposes a framework for evaluating the improvement of the power-accuracy trade-off in decision tree models using approximate circuits. The results of the electrical evaluation and accuracy provided in each step are obtained by applying different approximation and quantization techniques. This information can be used together with power-accuracy metrics to find the best comparator circuits for different applications of Decision Trees and dedicated hardware requirements. Pedro Aquino Silva, Mateus Grellert, Cristina Meinhardt |
VLSI-SoC | 2 |
| 2022 | A framework for designing power-efficient inference accelerators in tree-based learning applications
Brunno Abreu, Mateus Grellert, Sergio Bampi |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | C2PAx: Complexity-Aware Constant Parameter Approximation for Energy-Efficient Tree-Based Machine Learning AcceleratorsabstractTree-based machine learning models, like random forests and decision trees, are low-complexity solutions that provide an efficient prediction for a wide range of applications. These models are particularly interesting for energy-constrained platforms since they can be implemented with simple logical operations. Tree-based accelerators are also intrinsically resilient to errors, and this can be leveraged to boost energy efficiency with approximate computing techniques. The key operations in these models are comparisons to constants, making comparators excellent candidates for approximation. This paper presents a technique to approximate comparisons to constants called C2PAx, which is capable of reducing the area and energy of tree-based accelerators. The method consists in finding alternative constants that reduce circuit area while keeping an efficient prediction performance. It is also shown that the selection of the constant parameters directly influences both hardware complexity and model performance, demanding cross-layer optimization. For that, we extend an existing framework that generates VLSI tree-based accelerators, inserting our approximation proposal that allows selecting the constant parameters that maximize energy efficiency at the cost of minor accuracy drops. Simulation results demonstrate that C2PAx outperforms the Don’t Care logic approximation technique when accuracy and energy are jointly considered. C2PAx trades accuracy for significant reductions in the VLSI area, power, delay, and energy consumption compared to precise models. Brunno Abreu, Guilherme Paim, Mateus Grellert, Sergio Bampi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Logic Synthesis Meets Machine Learning: Trading Exactness for GeneralizationabstractLogic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning. Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee |
DATE | 26 |
| 2021 | Relying on a Rate Constraint to Reduce Motion Estimation ComplexityabstractThis paper proposes a rate-based candidate elimination strategy for Motion Estimation, which is considered one of the main sources of encoder complexity. We build from findings of previous works that show that selected motion vectors are generally near the predictor to propose a solution that uses the motion vector bitrate to constrain the candidate search to a subset of the original search window, resulting in less distortion computations. The proposed method is not tied to a particular search pattern, which makes it applicable to several ME strategies. The technique was tested in the VVC reference software implementation and showed complexity reductions of over 80% at the cost of an average 0.74% increase in BD-Rate with respect to the original TZ Search algorithm in the LDP configuration. Gabriel B. Sant'Anna, Luiz Henrique Cancellier, Ismael Seidel, Mateus Grellert, José Luís Güntzel |
ICASSP | 4 |
| 2021 | Quality and Complexity Assessment of Learning-Based Image Compression SolutionsabstractThis work presents an analysis of state-of-the-art learning-based image compression techniques. We compare 8 models available in the Tensorflow Compression package in terms of visual quality metrics and processing time, using the KODAK data set. The results are compared with the Better Portable Graphics (BPG) and the JPEG2000 codecs. Results show that JPEG2000 has the lowest execution times compared with the fastest learning-based model, with a speedup of $1.46 \times$ in compression and $30 \times$ in decompression. However, the learning-based models achieved improvements over JPEG2000 in terms of quality, specially for lower bitrates. Our findings also show that BPG is more efficient in terms of PSNR, but the learning models are better for other quality metrics, and sometimes even faster. The results indicate that learning-based techniques are promising solutions towards a future mainstream compression method. João Dick, Brunno Abreu, Mateus Grellert, Sergio Bampi |
ICIP | 3 |
| 2021 | Fast Logic Optimization Using Decision TreesabstractThis work evaluates the use of Decision Trees (DTs) methods for a fast logic minimization of Boolean functions. The proposed DT approach is compared to traditional Espresso logic minimizer and the minimization algorithms available in the ABC tool. The methods are compared with respect to the execution time, number of nodes and number of logic levels. The DT methods proved to be a faster alternative, reducing time by an average of 52% and 5.5% when compared to Espresso and ABC respectively, while keeping competitive results in terms of AIG depth and number of nodes. Additionally, in order to obtain smaller circuits at the cost of approximate results we tested DTs with limited tree depth. The trade-offs between synthesis time, circuit area and accuracy are also discussed. Compared to ABC, limiting the maximum tree depth leads to time savings of up to 52%, up to 86% less number of nodes, and up to 48% lower AIG depth, while maintaining acceptable accuracy results. Brunno Abreu, Augusto Andre Souza Berndt, Isac de Souza Campos, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi |
ISCAS | 6 |
| 2021 | Design of Energy-Efficient Gaussian Filters by Combining Refactoring and Approximate AddersabstractThe Gaussian image filter is a compute-intensive approach to reduce undesirable artifacts and generally serves as a pre-processing technique for emerging applications related to visual computing systems. This work evaluates alternatives for the design of power-efficient Gaussian Filters. The proposed optimization strategy combines: 1) a refactored function to minimize the arithmetic operations, and 2) a design space exploration investigating different approximation scenarios applied to the full adders. The exact version of our refactored Gaussian Filter architecture reduces the total power consumption and the circuit area by 18% and 12%, respectively compared with the baseline Gaussian Filter architecture. Moreover, the combination of different approximation levels with the refactored architecture provides design options with power reductions from 21% to 59% compared with the baseline Gaussian Filter architecture. Marcio Monteiro, Pedro Aquino Silva, Ismael Seidel, Mateus Grellert, Leonardo Bandeira Soares, José Luís Güntzel, Cristina Meinhardt |
ISCAS | 4 |
| 2021 | Complexity and Coding Efficiency Assessment of the Versatile Video Coding StandardabstractThe Versatile Video Coding standard was finalized by the Joint Video Exploration Team in July 2020 and is currently considered the state-of-the-art video compression technology. VVC significantly improves coding efficiency compared to HEVC thanks to several new features and tools that incur a large increase in computational cost. This paper presents a complexity and coding efficiency assessment of VVC divided into three analyses, focusing on: (1) the impact of using SIMD optimizations in the VVC Test Model software, (2) the impact of limiting partitioning structures when encoding, and (3) the computational cost associated to each encoding tool in VVC. Experimental results show that SIMD optimizations accelerate the encoding time by 40%, on average, and that limiting the available partitioning structures can decrease encoding time between 40% and 73%. The software profiling revealed that inter-frame prediction is responsible for almost half of the total encoding time. Finally, the paper also presents an analytic discussion on tools and partitioning possibilities that are rarely chosen in the mode decision process despite their high impact in coding complexity. Ícaro Siqueira, Guilherme Corrêa 0001, Mateus Grellert |
ISCAS | 3 |
| 2021 | Hardware-Friendly Search Patterns for the Versatile Video Coding Fractional Motion EstimationabstractThe recently finalized Versatile Video Coding (VVC) standard brings a number of new tools to substantially improve coding efficiency, which resulted in a significant increase in complexity. Therefore, the use of such new standard in mobile devices requires not only the use of dedicated VLSI hardware architectures, but also the adoption of techniques that could reduce its complexity to meet the real-time and energy efficiency requirements. In this context, the most intensive encoding tools, such as the Fractional Motion Estimation (FME), are the first ones to be considered for optimization. Thereby, this work presents three different search patterns for the FME that are able to reduce the complexity of VVC by creating fixed search windows around the most relevant candidates. Our experiments show that these patterns result in acceptable decreases in coding efficiency, with average BD-Rates in the range of 0.56% - 0.33%, for the LD-P configuration, and 0.47% - 0.21% for the RA configuration, while allowing for significant hardware optimization. Comparing to a state-of-the-art FME hardware architecture that searches over 48 candidates the three proposed patterns can operate in higher frequencies to achieve the same throughput and require less area, leading to less power consumption and higher energy efficiency. Particularly, one of the three patterns may lead to a 60% area reduction and 53% less dynamic power consumption. Vanio Rodrigues Filho, Marcio Monteiro, Ismael Seidel, Mateus Grellert, José Luís Güntzel |
MMSP | 4 |
| 2021 | SAD or SATD? How the Distortion Metric Impacts a Fractional Motion Estimation VLSI ArchitectureabstractVideo coding systems have to deal with a number of tradeoffs. The decision of adopting a specific distortion metric in the Fractional Motion Estimation (FME) step, for instance, presents a designer with a tradeoff between energy and coding efficiency. This paper analyzes such a tradeoff considering two of the most known and used distortion metrics, the Sum of Absolute Differences (SAD) and the Sum of Absolute Transformed Differences (SATD), within a High Efficiency Video Coding (HEVC)-compatible FME hardware architecture. We show that the SATD-based FME architecture is 1.94 times larger than the SAD-based one and consumes 2.07 times more energy. Ismael Seidel, Vanio Rodrigues Filho, Mateus Grellert, Luciano Volcan Agostini, José Luís Güntzel |
MMSP | 3 |
| 2020 | VLSI Design of Tree-Based Inference for Low-Power Learning ApplicationsabstractThe use of Machine Learning techniques in battery-powered devices has increased in recent years. Therefore, evaluating power-accuracy trade-offs of inference models has become important. This work explores decision trees architectures, analyzing the effects of model complexity and approximation in power and accuracy. By quantizing the inputs to limited widths, we increased the accuracy of the models in up to 8.7% compared to precise cases. In terms of power, the variation in model complexity was more significant than in input width, as we obtained reductions of up to 97% by decreasing the tree depth, compared to 88% when decreasing the width. Brunno Abreu, Mateus Grellert, Sergio Bampi |
ISCAS | 2 |
| 2019 | Coding Tree Early Termination for Fast HEVC Transrating Based on Random ForestsabstractVideo transrating has become an essential task in streaming service providers that need to transmit and deliver different versions of the same content for a multitude of users that operate under different network conditions. As the transrating operation is comprised of a decoding and an encoding step in sequence, a huge computational cost is required in such large-scale services, especially when considering the use of complex state-of-the-art codecs, such as the High Efficiency Video Coding (HEVC). This work proposes an early-termination method for complexity reduction of the HEVC transrating based on Random Forests, which use features obtained from the HEVC decoding process to accelerate the coding tree decisions during the re-encoding process. Experimental results show that the proposed method achieves an average transrating time reduction of 47.09% at the cost of a negligible bitrate increase of 0.292%. Thiago Luiz Alves Bubolz, Mateus Grellert, Bruno Zatt, Guilherme Corrêa 0001 |
ICASSP | 2 |
| 2019 | Fast Coding Unit Partition Decision for HEVC Using Support Vector MachinesabstractDespite the several speedup methods proposed in the literature, the computational complexity of High Efficiency Video Coding (HEVC) video encoding is still a problem. This paper proposes a fast coding unit (CU) partition decision for use in HEVC encoders based on support vector machine (SVM)-trained offline. The SVM classifiers, features, and training procedures are described in detail, and a justification for the use of SVMs is provided. The trained classifiers are incorporated into a modified reference encoder in the form of a fast CU partition decision algorithm, which decides if the exhaustive search for the best partition is continued or terminated prematurely. Using the proposed method, an average complexity reduction of 48% is achieved with a 0.48% Bjontegaard-Delta bitrate (BD-BR) loss using the random access coding configuration, 44% reduction with a 0.62% BD-BR loss for the Low Delay B, and a 41% reduction with a 0.6% BD-BR loss for the Low Delay P configuration. We also tested our approach under constant bitrate conditions, achieving a 47% reduction in encoding time with a 1.11% loss in the BD-BR. In addition, a decision threshold adaptation is also proposed to allow adjusting the rate-distortion/complexity trade-off of our solution. With this approach, the computational complexity reduction can be varied from 34.9% (with a 0.13% loss in the BD-BR) up to 52.4% (with a 1.11% BD-BR loss) using the random access configuration. Compared with the state-of-the-art solutions, our decision scheme outperforms the related works in terms of combined rate distortion and complexity. Mateus Grellert, Bruno Zatt, Sergio Bampi, Luís Alberto da Silva Cruz |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Learning-Based Complexity Reduction and Scaling for HEVC EncodersabstractThis article proposes a fast Coding Unit (CU) partition decision for use in HEVC encoders based on Decision Tree classifiers. The trees are employed in a modified low-complexity encoder that implements a fast CU partition decision algorithm. Using the proposed method, an average complexity reduction of 47.8% is achieved with a Bjontegaard Delta bitrate (BD-BR) loss of 0.24% in the Random Access coding configuration, and a 42.8% complexity reduction with a 0.19% BD-BR loss in the Low Delay B configuration. A decision threshold analysis is also presented to assess the rate-distortion-complexity trade-off of the proposed method at different complexity points, varying the complexity reduction from 28% (with a 0.04% loss in BD-BR) up to 60% (with a 3.6% BD-BR loss) using the Random Access configuration. A comparison with related works shows that the proposed method outperforms competing solutions in terms of both rate-distortion efficiency and complexity reduction. Mateus Grellert, Sergio Bampi, Guilherme Corrêa 0001, Bruno Zatt, Luís Alberto da Silva Cruz |
ICASSP | 1 |
| 2018 | Fast HEVC Transrating using Random ForestsabstractThis article describes a fast transrating solution for HEVC based on classification and machine learning techniques. Two classifiers are trained to predict the range of CTU quadtree depths that will be searched to find the best CTU partitioning. Three approaches are proposed for reducing the number of features used by the classifiers, two based on feature selection, and one based on feature transformation using autoencoders. A full transrating framework based on x265 is built for model training and evaluation. Experimental results using the x265 encoder show that an average 41.81% computational complexity reduction can be achieved at the cost of a tolerable 0.29% Bjontegaard-Delta bitrate, outperforming competing methods. Mateus Grellert, Tiago J. L. Oliveira, Carlos Rafael Duarte, Luís Alberto da Silva Cruz |
VCIP | 1 |
| 2016 | Energy-aware cache assessment of HEVC decodingabstractThis paper presents a thorough analysis of energy consumption of a software HEVC decoder. The evaluation utilizes a framework developed herein specifically to estimate the energy consumption in all levels of cache hierarchies. Our framework is based on analytical models combined with memory profiling; tools. Energy analyses of several cache hierarchies executing HEVC decoding with different input bit streams were carried out. Our results point to the most suited cache parameters for each video resolution. The energy was estimated for a 32nm CMOS technology. Our study includes different tradeoffs between energy efficiency and capacity, associativity, and main memory bandwidth. Our detailed analysis shows that the higher are the cache features, the more efficient is the energy consumption. The main memory bandwidth evaluation shows that the energy consumption increases with the main memory bandwidth requirement. Full HD video resolutions require up to 90 times higher bandwidth and 57 times more energy than class D resolutions. Eduarda Monteiro, Mateus Grellert, Sergio Bampi, Bruno Zatt |
ISCAS | 2 |
| 2016 | Complexity-scalable HEVC encodingabstractHEVC encoders impose several challenges in resource-constrained embedded applications, especially under real-time and battery constraints. This paper proposes a complexity-scalable encoder that is able to achieve considerable time savings while maintaining an efficient rate-distortion-complexity tradeoff. To design the system, a complexity analysis of HEVC-supported parameters as well as new ones introduced in this work is presented. To build the configurations for each target savings, a Complexity Target Satisfaction algorithm was designed. This algorithm was able to reduce the optimization space by approximately 600 times, producing efficient configurations with up to 90% time savings. The complete system was implemented and tested against a state-of-the-art complexity management implementation. The results proved the efficiency of our solution, as it achieves more time savings and better compression for savings of 60% and higher. Mateus Grellert, Sergio Bampi, Bruno Zatt |
PCS | 1 |
| 2015 | Rate-distortion and energy performance of HEVC and H.264/AVC encoders: A comparative analysisabstractA quantitative, systematic, and detailed analysis of the energy impacts of the tools that comprise two of the most recent video coding standards: the High Efficiency Video Coding (HEVC) and the H.264/AVC is presented. Our comparative study measures the energy consumption effects of important video-coding parameters, like Search Range (SR), Quantization Parameter (QP), and video resolution on both encoders. The obtained results for HEVC showed, for the Random Access (RA) prediction structure, gains of 25% in BD-Rate over H.264/AVC at the expense of 17% higher energy consumption. A new metric we defined herein, called BD-Energy, was used in the SR analysis, and the results from this investigation showed HEVC achieved an energy consumption up to 37.6% higher for a BD-Rate gain of 32.2%. The QP analysis demonstrated that the energy consumption gap between both encoders varies greatly as QP increases, resulting in a 15.08% difference from QP 22 to QP 37, on average. The major finding from our work is that the HEVC encoder presents better results in the energy/compression trade-off, but this efficiency is reduced as encoding becomes more complex, as our results discovered that the HEVC energy consumption scales faster. Eduarda Monteiro, Mateus Grellert, Sergio Bampi, Bruno Zatt |
ISCAS | 2 |
| 2013 | Hardware-software collaborative complexity reduction scheme for the emerging HEVC intra encoderabstractHigh Efficiency Video Coding (HEVC/H.265) is an emerging standard for video compression that provides almost double compression efficiency at the cost of major computational complexity increase as compared to current industry-standard Advanced Video Coding (AVC/H.264). This work proposes a collaborative hardware and software scheme for complexity reduction in an HEVC Intra encoding system, with run-time adaptivity. Our scheme leverages video content properties which drive the complexity management layer (software) to generate a highly probable coding configuration. The intra prediction size and direction are estimated for the prediction unit which provides reduced computational-complexity. At the hardware layer, specialized coprocessors with enhanced reusability are employed as accelerators. Additionally, depending upon the video properties, the software layer administers the energy management of the hardware coprocessors. Experimental results show that a complexity reduction of up to 60 % and the energy reduction up to 42 % are achieved. Muhammad Usman Karim Khan, Muhammad Shafique 0001, Mateus Grellert, Jörg Henkel |
DATE | 3 |
| 2013 | An adaptive workload management scheme for HEVC encodingabstractManaging the complexity of the emerging HEVC standard is a matter of academic and industrial research since its earlier versions. The sophisticated and computation-intensive tools involved in the encoding process must be leveraged if real-time applications are considered. In this paper, we propose a workload management scheme for dynamically controlling the computational complexity of HEVC, under user-defined operation frequency and target FPS. Our scheme receives these two parameters as input and aims to meet the target FPS by adjusting different encoding parameters during execution time. Experiments demonstrate that our scheme successfully meets the target FPS while introducing negligible rate-distortion losses. A comparison with state-of-the-art shows that our scheme is capable of achieving a time reduction of up to 43% for Full HD sequences, with a maximum loss of 0.03 dB in Y-PSNR and a 3.5% increase in bitrate. Mateus Grellert, Muhammad Shafique 0001, Muhammad Usman Karim Khan, Luciano Volcan Agostini, Júlio C. B. de Mattos, Jörg Henkel |
ICIP | 1 |
| 2013 | Iterative random search: a new local minima resistant algorithm for motion estimation in high-definition videos
Marcelo Schiavon Porto, Cassio Cristani, Pargles Dall'Oglio, Mateus Grellert, Júlio C. B. de Mattos, Sergio Bampi, Luciano Volcan Agostini |
Multim. Tools Appl. | 4 |
| 2012 | Motion Vectors Merging: Low Complexity Prediction Unit Decision Heuristic for the Inter-prediction of HEVC EncodersabstractThis paper presents the Motion Vectors Merging (MVM) heuristic, which is a method to reduce the HEVC inter-prediction complexity targeting the PU partition size decision. In the HM test model of the emerging HEVC standard, computational complexity is mostly concentrated in the inter-frame prediction step (up to 96% of the total encoder execution time, considering common test conditions). The goal of this work is to avoid several Motion Estimation (ME) calls during the PU inter-prediction decision in order to reduce the execution time in the overall encoding process. The MVM algorithm is based on merging NxN PU partitions in order to compose larger ones. After the best PU partition is decided, ME is called to produce the best possible rate-distortion results for the selected partitions. The proposed method was implemented in the HM test model version 3.4 and provides an execution time reduction of up to 34% with insignificant rate-distortion losses (0.08 dB drop and 1.9% bitrate increase in the worst case). Besides, there is no related work in the literature that proposes PU-level decision optimizations. When compared with works that target CU-level fast decision methods, the MVM shows itself competitive, achieving results as good as those works. Felipe Sampaio, Sergio Bampi, Mateus Grellert, Luciano Volcan Agostini, Júlio C. B. de Mattos |
ICME | 3 |
| 2011 | A multilevel data reuse scheme for Motion Estimation and its VLSI designabstractMotion Estimation (ME) in video coding is a vital component that excels not only in computational complexity, but off-chip memory bandwidth as well. These two issues are considered critical constraints in terms of High Definition (HD) video coding, since a large volume of data must be processed. The multilevel data reuse scheme proposed in this paper is able to reduce the off-chip memory bandwidth, with direct impact in throughput and energy consumption. This scheme explores the concept of overlapped Search Windows (SW) in more than one level and poses no harm to video quality. Comparisons with related works show that this solution provides the best tradeoff between the use of on-chip memory and reduction of the off-chip memory bandwidth. The data reuse scheme was applied in a ME architecture and the synthesis results show that this solution presented the lowest use of hardware resources and the highest operation frequency among related works. The proposed architecture is able to process 1080p videos at 25 fps, and the reduction ratio of off-chip memory access achieved by the architecture is greater than 95% when compared to the traditional method. Mateus Grellert, Felipe Sampaio, Júlio C. B. de Mattos, Luciano Volcan Agostini |
ISCAS | 1 |