Felipe Sampaio

dblp:38/7554 · also Felipe Martin Sampaio · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
3since 2021 · last 2025
0000-0002-5700-2221ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Approximate SRAM-Based Memory System for Intra-Frame Prediction in VVC Encoders
abstract
This paper introduces a memory system for the Versatile Video Coding (VVC) standard, aimed at optimizing energy consumption in intra-frame prediction through approximate storage techniques. We propose a system that employs data approximation due to dynamic voltage scaling Static Random Access Memory (SRAM) in two distinct buffers: the Neighbor Samples Buffer (NeighSB) and the Original Samples Buffer (OrigSB). Our findings indicate the feasibility of approximate storage exploitation of our memory system, with minor coding efficiency drops at controlled error rates. As NeighSB and OrigSB presented distinct resilience behaviors, we evaluate different approximation strengths to obtain the most suitable coding efficiency, enabling still effective energy savings. Finally, we investigate the energy savings achieved through different SRAM operation levels, yielding reductions of 42%-62%.
Yasmin Camargo, Matheus Isquierdo, Daniel Palomino 0001, Bruno Zatt, Felipe Sampaio
ISCAS5
2025 Exploiting Approximate SRAM for Energy-Efficient Integer Motion Estimation on VVC Encoders
abstract
Video coding is a critical technology for enabling many modern applications. However, the high complexity of state-of-the-art encoders leads to energy consumption challenges, particularly in memory systems. This paper proposes an integer motion estimation (IME) system using approximate SRAM memories to enhance the energy efficiency of VVC encoders. The approximate SRAM memories are employed for both the current block and the search area buffers by reducing the supply voltage. The proposed IME system is evaluated using a customized tool that simulates the approximation effects on both buffers based on real energy measurements from a 28nm SRAM. This tool is integrated with VVenC, a fast VVC implementation, to assess the impact on coding efficiency and energy consumption. Experimental results show that the IME system using approximate SRAM can reduce energy consumption in reading operations by up to 55%, with a low impact on coding efficiency.
Matheus Isquierdo, Felipe Sampaio, Bruno Zatt, Nikil Dutt, Daniel Palomino 0001
ISCAS2
2022 Video Decoder Improvements with Near-Data Speculative Motion Compensation Processing
abstract
Video decoder implementations are still evolving as they directly affect a large fraction of embedded systems nowadays. In this context, Versatile Video Coding (VVC) brings increased compression efficiency, which comes with extra over-head in terms of computational effort and energy consumption. At the same time, emerging Near-Data Processing (NDP) architectures promise drastic time and energy cuts for applications with data streaming behavior. In this paper, a speculative Motion Compensation (MC) is proposed to enable video decoders improvements through the exploitation of NDP. We adopted a large-vector SIMD-based NDP system (called VIMA) that provides high-performance operations over 2 K vectors. The proposed strategy leverages the correlation between the prediction modes and the motion data between spatially neighboring blocks within a frame to speculatively perform the MC for an entire region of 2Kx128 samples. MC interpolation kernels were implemented using VIMA and x86 AVX-256 SIMD libraries. Our NDP-based kernel implementation allows speedup of $1.9\times$ to $22\times$ compared to the x86 baseline solutions. Stepping forward, based on a coalescence estimation, our strategy can properly handle interpolation misses, achieving MC performance improvements from 7% to 64%.
Garrenlus de Souza, José Rodrigo Azambuja, Bruno Zatt, Marco A. Z. Alves, Sergio Bampi, Felipe Sampaio
ISCAS6
2020 Memory Assessment Of Versatile Video Coding
abstract
This paper presents a memory assessment of the next-generation Versatile Video Coding (VVC). The memory analyses are performed adopting as a baseline the state-of-the-art High-Efficiency Video Coding (HEVC). The goal is to offer insights and observations of how critical the memory requirements of VVC are aggravated, compared to HEVC. The adopted methodology consists of two sets of experiments: (1) an overall memory profiling and (2) an inter-prediction specific memory analysis. The results obtained in the memory profiling show that VVC access up to 13.4x more memory than HEVC. Moreover, the inter-prediction module remains (as in HEVC) the most resource-intensive operation in the encoder: 60%-90% of the memory requirements. The inter-prediction specific analysis demonstrates that VVC requires up to 5.3x more memory accesses than HEVC. Furthermore, our analysis indicates that up to 23% of such growth is due to VVC novel-CU sizes (larger than 64x64).
Arthur Cerveira, Luciano Volcan Agostini, Bruno Zatt, Felipe Sampaio
ICIP4
2019 Hybrid Scratchpad Video Memory Architecture for Energy-Efficient Parallel HEVC
abstract
A hybrid scratchpad video memory (Hy-SVM) for energy-efficient tiles-parallelized high-efficiency video coding (HEVC) is presented here. The key ideas behind the Hy-SVM include: application-specific design and management; combined multiple levels of private and shared memories that jointly exploit intra-tile and inter-tiles data reuse; scratchpad memories (SPMs) as on-chip data storage; SRAM; and STT-RAM hybrid design. We propose a design methodology for the Hy-SVM that leverages application-specific properties to properly define the SPMs parameters. The inter-tiles data reuse potential of parallel HEVC is exploited by our run-time overlap prediction scheme, which identifies the redundant memory access behavior by analyzing monitored past frames encoding. Based on the predicted overlap characteristics, the Hy-SVM integrates memory access management units to control the access dynamics to the private/shared SPM levels. Furthermore, adaptive access management units (APMUs) can strongly reduce on-chip energy consumption due to the predicted overlap formation. The experimental results demonstrate the Hy-SVM overall energy savings of 11%-64% (4-tile) and 8%-46% (8-tile) when compared with related works. From the external memory perspective, the Hy-SVM can improve data reuse, resulting in 14%-59% of off-chip energy consumption (compared with no inter-tiles data reuse scenarios). In addition, our APMU contributes by reducing on-chip energy consumption of the Hy-SVM by 58%, on average. Thus, compared with related works, the Hy-SVM presents the lowest on-chip energy consumption. Moreover, the overhead of implementing our management units insignificantly affects the performance- and energy-efficiency of the Hy-SVM.
Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Jörg Henkel, Sergio Bampi
IEEE Trans. Circuits Syst. Video Technol.1
2015 Approximation-aware Multi-Level Cells STT-RAM cache architecture
abstract
Current manycore processors exhibit large on-chip last-level caches that may reach sizes of 32MB - 128MB and incur high power/energy consumption. The emerging Multi-Level Cells (MLC) STT-RAM memory technology improves the capacity and energy efficiency issues of large-sized memory banks. However, MLC STT-RAM incurs non-negligible protection overhead to ensure reliable operations when compared to the Single-Level Cells (SLC) STT-RAM. In this paper, we propose an approximation-aware MLC STT-RAM cache architecture, which is partially-protected to restrict the reliability overhead and in turn leverages variable resilience characteristics of different applications for adaptively curtailing the protection overhead under a given error tolerance level. It thereby improves the energy-efficiency of the cache while meeting the reliability requirements. Our cache architecture is equipped with a latency-aware hardware module for double-error correction. To achieve high energy efficiency, approximation-aware read and write policies are proposed that perform approximate storage management while tolerating some errors bounded within the user-provided tolerance level. The architecture also facilitates runtime control on the quality of applications' results. We perform a case study on the next-generation advanced video encoding that exhibit memory-intensive functional blocks with varying resilience properties and support for parallelism. Experimental results demonstrate that our approximation-aware MLC STT-RAM based cache architecture can improve the energy efficiency compared to state-of-the-art fully-protected caches (7%-19%, on average), while incurring minimal quality penalties in the output (-0.219% to -0.426%, on average). Furthermore, our architecture supports complete error protection coverage for all cache data when processing non-resilient application. The hardware overhead to implement our approximation-aware management negligibly affects the energy efficiency (0.15%-1.3% of overhead) and the access latency (only 0.02%-1.56% of overhead).
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
CASES1
2014 dSVM: Energy-efficient distributed Scratchpad Video Memory Architecture for the next-generation High Efficiency Video Coding
abstract
An energy-efficient distributed Scratchpad Video Memory Architecture (dSVM) for the next-generation parallel High Efficiency Video Coding is presented. Our dSVM combines private and overlapping (shared) Scratchpad Memories (SPMs) to support data reuse within and across different cores concurrently executing multiple parallel HEVC threads. We developed a statistical method to size and design the organization of the SPMs along with a supporting memory reading policy for energy efficiency. The key is to leverage the HEVC and video content knowledge. Furthermore, we integrate an adaptive power management policy for SPMs to manage the power states of different memory parts at run time depending upon the varying video content properties. Our experimental results illustrate that our dSVM architecture reduces the overall memory energy consumption by up to 51%-61% compared to parallelized state-of-the-art solutions [11]. The dSVM external memory energy savings increase with an increasing number of parallel HEVC threads and size of search window. Moreover, our SPM power management reacts to the current video properties and achieves up to 54% on-chip leakage energy savings.
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
DATE1
2014 Energy-efficient architecture for advanced video memory
abstract
An energy-efficient hybrid on-chip video memory architecture (enHyV) is presented that combines private and shared memories using a hybrid design (i.e., SRAM and emerging STT-RAM). The key is to leverage the application-specific properties to efficiently design and manage the enHyV. To increase STT-RAM lifetime, we propose a data management technique that alleviates the bit-toggling write occurrences. An adaptive power management is also proposed for static-energy savings. Experimental results illustrate that enHyV reduces on-chip static memory energy compared to SRAM-only version of enHyV and to state-of-art AMBER hybrid video memory [9] by 66%-75% and 55%-76%, respectively. Furthermore, negligible external memory energy consumption is required for reference frames communication (98% lower than state-of-the-art Level C+ technique [18]). Our data management significantly improves the enHyV STT-RAM lifetime, achieving 0.83 of normalized lifetime (near to the optimal case). Our hybrid memory design and management incur low overhead in terms of latency and dynamic energy.
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
ICCAD1
2014 Content-driven memory pressure balancing and video memory power management for parallel high efficiency video coding
abstract
We present a novel content-driven memory pressure balancing and video memory power management scheme for parallel High Efficiency Video Coding (HEVC). The key is to leverage the application-specific knowledge to balance the (instant) access pressure on Scratchpad-based Video Memories (SVMs) for parallelized video processing. Our scheme accurately predicts the memory requirements of each processing core based on monitored memory usage and leverages this knowledge to perform a categorization of different video regions. Afterwards, it employs an adaptive policy for memory pressure balancing by rescheduling encoding of different video blocks based on their categories. This balancing also facilitates our scheme to perform efficient power-gating of unused parts of SVMs. Experimental results show that our scheme reduces the variations in the memory pressure by 37%-83% when compared to the traditional raster scan processing for 4- and 16-core parallelized HEVC encoder. Our content-driven power management saves 56% (on average) of SVM leakage energy.
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
ISLPED1
2013 Energy-efficient memory hierarchy for motion and disparity estimation in multiview video coding
abstract
This work presents an energy-efficient memory hierarchy for Motion and Disparity Estimation on Multiview Video Coding employing a Reference Frames-Centered Data Reuse (RCDR) scheme. In RCDR the reference search window becomes the center of the motion/disparity estimation processing flow and calls for processing all blocks requesting its data. By doing so, RCDR avoids multiple search window retransmissions leading to reduced number of external memory accesses, thus memory energy reduction. To deal with out-of-order processing and further reduce external memory traffic, a statistics-based partial results compressor is developed. The on-chip video memory energy is reduced by employing a statistical power gating scheme and candidate blocks reordering. Experimental results show that our reference-centered memory hierarchy outperforms the state-of-the-art [7][13] by providing reduction of up to 71% for external memory energy, 88% on-chip memory static energy, and 65% on-chip memory dynamic energy.
Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel
DATE1
2013 Content-adaptive reference frame compression based on intra-frame prediction for multiview video coding
abstract
This paper presents a content-adaptive reference frame compression scheme to alleviate the large overhead of external memory communication during the Motion and Disparity Estimation process in Multiview Video Coding (MVC). Our scheme is based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples. The intra-prediction residue is compressed by a path composed of non-linear quantization and Huffman-based entropy encoder. Four different quantization strengths and Huffman tables were statistically defined. They are dynamically selected according to a content adaptation strategy, which classifies the original blocks based on their spatial homogeneity. Experimental results show that the proposed content-adaptive compression scheme is able to reduce the external memory accesses by up to 63% along with negligible losses in the MVC encoder rate-distortion performance. Compared to the best available related work [12] our content-adaptive reference frame compression achieves 39% reduced external memory accesses, while still providing a BD-PSNR increase of 0.03dB.
Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Jörg Henkel, Sergio Bampi
ICIP1
2012 Real-time block matching motion estimation onto GPGPU
abstract
This work presents an efficient method to map Motion Estimation (ME) algorithms onto General Purpose Graphic Processing Unit (GPGPU) architectures using CUDA programming model. Our method jointly exploits the massive parallelism available in current GPGPU devices and the parallelization potential of ME algorithms: Full Search (FS) and Diamond Search (DS). Our main goal is to evaluate the feasibility of achieving real-time high-definition video encoding performance running on GPUs. For comparison reasons, multi-core parallel and distributed versions of these algorithms were developed using OpenMP and MPI (Message Passing Interface) libraries, respectively. The CUDA-based solutions achieve the highest speed-up in comparison with OpenMP and MPI versions for both algorithms and, when compared to the state-of-the-art, our FS and DS solutions reach up to 18x and 11x speed-up, respectively.
Eduarda Monteiro, Marilena Maule, Felipe Sampaio, Cláudio Machado Diniz, Bruno Zatt, Sergio Bampi
ICIP3
2012 A memory aware and multiplierless VLSI architecture for the complete Intra Prediction of the HEVC emerging standard
abstract
This work proposes a hardware architecture for the Intra Frame Prediction of the emerging High Efficiency Video Coding (HEVC) standard. The architecture was designed considering all innovative features of the Intra Prediction included in the HEVC, i.e. all modes and all Prediction Units (PU) sizes. Performance and memory accesses are a problem in the HEVC intra prediction and hardware architecture designs are good alternative to solve these issues, especially when energy-efficient solutions are targeted. Buffers and internal memories were used in the designed architecture to decrease the number of external memory accesses. Two independent data paths processing eight samples in parallel and a deep and multiplierless pipeline were designed to increase the throughput. The architecture was synthesized using an IBM 65nm CMOS technology. The results have shown that the architecture is able to process 30 HD720p frames per second and 13 HD1080p frames per second when running at 500 MHz, reducing in 95% the accesses to the external memory.
Daniel Palomino 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Altamiro Amadeu Susin
ICIP2
2012 Motion Vectors Merging: Low Complexity Prediction Unit Decision Heuristic for the Inter-prediction of HEVC Encoders
abstract
This paper presents the Motion Vectors Merging (MVM) heuristic, which is a method to reduce the HEVC inter-prediction complexity targeting the PU partition size decision. In the HM test model of the emerging HEVC standard, computational complexity is mostly concentrated in the inter-frame prediction step (up to 96% of the total encoder execution time, considering common test conditions). The goal of this work is to avoid several Motion Estimation (ME) calls during the PU inter-prediction decision in order to reduce the execution time in the overall encoding process. The MVM algorithm is based on merging NxN PU partitions in order to compose larger ones. After the best PU partition is decided, ME is called to produce the best possible rate-distortion results for the selected partitions. The proposed method was implemented in the HM test model version 3.4 and provides an execution time reduction of up to 34% with insignificant rate-distortion losses (0.08 dB drop and 1.9% bitrate increase in the worst case). Besides, there is no related work in the literature that proposes PU-level decision optimizations. When compared with works that target CU-level fast decision methods, the MVM shows itself competitive, achieving results as good as those works.
Felipe Sampaio, Sergio Bampi, Mateus Grellert, Luciano Volcan Agostini, Júlio C. B. de Mattos
ICME1
2012 Spread and Iterative Search: A High Quality Motion Estimation Algorithm for High Definition Videos and Its VLSI Design
abstract
This paper presents the Spread and Iterative Search (S&IS) motion estimation algorithm, which uses a random spread evaluation together with a central iterative evaluation to avoid local minima falls and to increase the image quality for high definition videos. Considering Full HD videos, S&IS reached an average PSNR gain of 1.41dB when compared to Diamond Search (DS), with an increase of about four times in the number of evaluated blocks. When compared to Full Search (FS), the S&IS achieved an average PSNR loss of 1.56 dB, evaluating 73 times less blocks than FS. An efficient architecture for the S&IS algorithm is also presented in this paper. The architecture was designed targeting in real time processing (30 frames per seconds) for QFHD videos (3840×2160 pixels). The architecture was described in VHDL and synthesized for and Altera Stratix 4 FPGA and for ST90nm standard cells technology. Booth syntheses show that the architecture is able to process QFHD frames in real time. The standard cells version is able to reach also a good trade-off among area, memory and power consumption, processing QFHD videos with 62.2 mW.
Gustavo Sanchez, Luciano Volcan Agostini, Felipe Sampaio, Marcelo Schiavon Porto, Sergio Bampi
ICME3
2011 Run-time adaptive energy-aware motion and disparity estimation in multiview video coding
abstract
This paper presents a novel run-time adaptive energy-aware Motion and Disparity Estimation (ME, DE) architecture for Multiview Video Coding (MVC). It incorporates efficient memory access and data prefetching techniques for jointly reducing the on/off-chip memory energy consumption. A dynamically expanding search window is constructed at run time to reduce the off-chip memory accesses. Considering the multi-stage processing nature of advanced fast ME/DE schemes, a reduced-sized multi-bank on-chip memory is employed which can be power-gated depending upon the video properties. As a result, when tested for various video sequence, our approach provides a dynamic energy reduction of 82--96% for the off-chip memory and a leakage energy reduction of 57--75% for the on-chip memory compared to the Level-C and Level-C+ [7] prefetching techniques (which are the prominent data reuse and prefetching techniques in ME for video coding). The proposed ME/DE architecture is synthesized using a 65nm IBM low power technology. Compared to state-of-the-art MVC ME/DE hardware [14], our architecture provides 66% and 72% reduction in the area and power consumption, respectively. Moreover, our scheme achieves 30fps ME/DE 4-view HD1080p encoding with a power consumption of 74mW.
Bruno Zatt, Muhammad Shafique 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel
DAC3
2011 A high throughput H.264/AVC intra-frame encoding loop architecture for HD1080p
abstract
In this work we present a high throughput hardware architecture for the H.264/AVC intra-frame encoder exploiting the parallelism of intra prediction, forward and inverse transforms and quantization. Since there is a strong data dependency between the intra prediction and the image reconstruction loop, the latency of this path is a key design issue in order to provide high performance coding. Considering that 77% of the total intra-encoding computation is spent in these modules, our architecture handles a 4-pixel wide intra prediction module and a 16-pixel wide reconstruction loop. Compared to the state-of-the-art our approach reduces by 47% the number of cycles to process a macroblock. Running at 150 MHz our architecture guarantees encoding of 61 HD1080p frames per second. The developed architecture requires 73.4 MHz to real-time encode HD1080p, which is a 46% reduction of the frequency requirement compared to the state-of-the-art.
Cláudio Machado Diniz, Bruno Zatt, Cristiano Thiele, Altamiro Amadeu Susin, Sergio Bampi, Felipe Sampaio, Daniel Palomino 0001, Luciano Volcan Agostini
ISCAS6
2011 A multilevel data reuse scheme for Motion Estimation and its VLSI design
abstract
Motion Estimation (ME) in video coding is a vital component that excels not only in computational complexity, but off-chip memory bandwidth as well. These two issues are considered critical constraints in terms of High Definition (HD) video coding, since a large volume of data must be processed. The multilevel data reuse scheme proposed in this paper is able to reduce the off-chip memory bandwidth, with direct impact in throughput and energy consumption. This scheme explores the concept of overlapped Search Windows (SW) in more than one level and poses no harm to video quality. Comparisons with related works show that this solution provides the best tradeoff between the use of on-chip memory and reduction of the off-chip memory bandwidth. The data reuse scheme was applied in a ME architecture and the synthesis results show that this solution presented the lowest use of hardware resources and the highest operation frequency among related works. The proposed architecture is able to process 1080p videos at 25 fps, and the reduction ratio of off-chip memory access achieved by the architecture is greater than 95% when compared to the traditional method.
Mateus Grellert, Felipe Sampaio, Júlio C. B. de Mattos, Luciano Volcan Agostini
ISCAS2
2009 Low latency and high throughput dedicated loop of transforms and quantization focusing in the H.264/AVC Intra Prediction
abstract
This paper presents an efficient architectural design for a dedicated transforms and quantization loop. This design targeted the Intra Prediction of the H.264/AVC standard. The architecture was designed intending to achieve the best possible relation between throughput, latency and hardware resources consumption. The latency and throughput of this loop are extremely important to define the intra prediction performance. The use of hardware was reduced through the reuse of the same datapath for different calculations. The architecture was synthesized to Altera Stratix III FPGA and to the TSMC 0.18 ¿m standard-cells technology. The architecture, when mapped to standard-cells, reaches a processing rate of 114 HDTV frames per second, attending the intra prediction restrictions.
Daniel Palomino 0001, Felipe Sampaio, Robson Dornelles, Luciano Volcan Agostini
ICIP2