EDBT 2026 Demo / reviewers in the wild / expert
Jones Goebel
dblp:184/4479 · also Jones W. Goebel, Jones William Goebel
· DBLP profile ↗
8ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0002-5937-9794ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | High-Throughput Design for a Multi-Size DCT-II Targeting the AV1 EncoderabstractThis paper presents a dedicated multi-size hardware design for the Discrete Cosine Transform type II (DCT-II) of AV1 encoder. The DCT-II is one of four transform kernels supported by AV1; however, DCT-II is used in all configurations defined by AV1. Moreover, the 1D DCT-II can be applied for five different sizes ranging from 4-point up to 64-point. The 1D multi-size DCT-II was designed to process multiple transform sizes in parallel, always processing 64 samples in parallel for any size. The presented solution can process UHD 8K videos at 60 frames per second when running at 46.6 MHz, with a power dissipation of 44.48 mW and an area of 261.28 Kgates. To the best of authors' knowledge, this is the first work in the literature presenting a hardware design for the AV1 DCT-II transform. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 1 |
| 2023 | A High-Throughput Hardware Design for the AV1 Decoder IntrapredictionabstractThe Alliance for Open Media (AOMedia) (AV1) was released in 2018 as a royalty-free and open-source video codec. AV1 was developed by the AOMedia that is composed of many leading tech companies. AV1 has the goal to process ultrahigh definition (UHD) 8K (7680$\times4320$pixels) and 4K videos (3840$\times2160$pixels) and to achieve high coding efficiency, which leads to increased complexity when compared to other codecs in the market, such as VP9, HEVC, and H.264. This article presents the AV1 intraprediction decoder (AVID), a dedicated high-throughput hardware design for the AV1 decoder intraprediction supporting the AV1 68 prediction modes and 19 block sizes. The proposed architecture can decode UHD 4K videos at 120 frames/s in the worst case, requiring an operation frequency of 279.93 MHz and demanding a total area of 234.45 kgates with a power dissipation of 27.74 mW. The comparison with related works showed that AVID reached the smallest area and very competitive power results. To the best of the authors’ knowledge, this is the first article detailing the hardware design of a complete decoder for intraprediction targeting the AV1 codec. Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | A High-Throughput Design for the H.266/VVC Low-Frequency Non-Separable TransformabstractThis paper presents a high throughput hardware design for the Low-Frequency Non-Separable Transform (LFNST) of the Versatile Video Coding (H.266/VVC) standard. The LFNST is a secondary transform used to transform the coefficients already transformed by the DCT-II as primary transform over the residues from the directional intra prediction. The LFNST architecture was designed to process Ultra-High Definition (UHD) videos with $4098 \times 2160$ pixels (4K) at 60 frames per second. Our solution presents an area utilization of 99.13 kgates and a power dissipation of 38.50 mW, when running at 186.62 MHz and considering the worst-case operation (processing the LFNST $4\times 4$ through TU size of $4\times 4$). Jones Goebel, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 1 |
| 2020 | Efficient Hardware Design for the AV1 CDEF Filter Targeting 4K UHD VideosabstractDeveloped by the AOMedia industry consortium, the AOM Video 1 (AV1) is an open-source and royalty-free video encoder released in June 2018. The Constrained Directional Enhancement Filter (CDEF) is one of the three AV1 in-loop filters and it is the focus of this work. The CDEF has the goal to reduce ringing artifacts generated with the encoding process, acting as a directional deringing filter. This paper presents a hardware design for the AV1 CDEF targeting real-time processing of 4K Ultra High Definition (UHD) videos. The architecture was synthesized to ASIC using the 40nm TSMC library, requiring 185 kgates and with a power dissipation of 43 mW when running at 93 MHz, reaching the frame rate of 60 frames per second (fps). To the best of the author's knowledge, there is no other work in the literature with dedicated hardware design for the AV1 CDEF. Eduardo Zummach, Roberta Palau, Jones Goebel, Daniel Palomino 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto |
ISCAS | 3 |
| 2020 | 4D-DCT Hardware Architecture for JPEG Pleno Light Field CodingabstractThis paper presents a 4D-DCT hardware architecture for Light Field Coding according to the JPEG Pleno standard. It is composed of two instances of 2D-DCT engines and a novel 4D Transposition Memory organization. Experimentally-defined fixed-point representation and LSB pruning techniques are employed do reduce hardware area and power dissipation. The proposed architecture operates over 4D-hypercubes of up to 8x8x8x8 samples and reaches performance to process 30 Lytro-like light fields per second dissipating 145.32mW at 825.75MHz. This is the first known 4D-DCT hardware architecture for light field coding and demonstrates the feasibility of such solutions on real-world systems. Matheus Jahnke, Jones Goebel, Daniel Palomino 0001, Guilherme Corrêa 0001, Luciano Volcan Agostini, Marcelo Schiavon Porto, Bruno Zatt |
VCIP | 2 |
| 2016 | High-throughput and memory-aware hardware of a sub-pixel interpolator for multiple video coding standardsabstractReal-time operation and low-power dissipation in video coding systems have become important research challenges, especially in mobile devices with limited battery and computational resources. There are many video coding standards coexisting in the market nowadays, so it is important for current devices to support different video coding standards. This paper presents a multi-standard luminance sub-samples interpolator hardware design for the Motion Compensation (MC) and Fractional Motion Estimation (FME), with support to MPEG-2/4, H.264/AVC, HEVC, and AVS/2 video coding standards. Our design is able to save hardware resources through an optimized filter organization, totally compliant with the focused standards and capable to interpolate samples for UHD 4320p@60fps at real time. The 45nm standard-cell library implementation dissipates 10mW, when processing according MPEG-2 standard, up to 46.4mW when processing AVS2. Guilherme Paim, Jones Goebel, Wagner Penny, Bruno Zatt, Marcelo Schiavon Porto, Luciano Volcan Agostini |
ICIP | 2 |
| 2016 | An efficient sub-sample interpolator hardware for VP9-10 standardsabstractThis paper presents a hardware design for the sub-sample interpolator used in FME (Fractional Motion Estimation) and MC (Motion Compensation) stages according to the VP9 and VP10 video-coding standards. The proposed architecture is able to save hardware resources through an optimized-filter organization whereas reaching high-throughput and low-power dissipation. The hardware design was described in Verilog and synthesized for ASIC technology. The synthesis results were generated for 45nm Nangate standard cells and demonstrate that the developed architecture is able to process 2160p@60fps videos with a power dissipation of 2.34mW focusing on a VP9-10 decoder. Guilherme Paim, Wagner Penny, Jones Goebel, Vladimir Afonso, Altamiro Amadeu Susin, Marcelo Schiavon Porto, Bruno Zatt, Luciano Volcan Agostini |
ICIP | 3 |
| 2016 | An HEVC multi-size DCT hardware with constant throughput and supporting heterogeneous CUsabstractThis paper presents an efficient hardware design for the Discrete Cosine Transform (DCT) of High Efficiency Video Coding standard (HEVC). This hardware supports all HEVC transform sizes: 4×4, 8×8, 16×16, and 32×32 including any combination of the Transform Unit (TU) sizes. The proposed DCT architecture has a constant throughput of 32 coefficients per cycle, independently of the transform sizes combination. The architecture was synthesized for a Nangate 45nm standard-cell library and the power analysis was made considering real input vectors. The synthesis results show a very good tradeoff between area, power dissipation and processing rates. The architecture is able to process 1.6G coeff/s when running at 50MHz dissipating 24.2 mW. These results allow a processing rate of 30 HD 1080p frames per second when evaluating 17 HEVC prediction modes. Jones Goebel, Guilherme Paim, Luciano Volcan Agostini, Bruno Zatt, Marcelo Schiavon Porto |
ISCAS | 1 |