Roberto Azevedo

dblp:234/2391 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0001-5473-506XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image Compression
abstract
Generative neural image compression supports data representation at extremely low bitrate, synthesizing details at the client and consistently producing highly realistic images. By leveraging the similarities between quantization error and additive noise, diffusion-based generative image compression codecs can be built using a latent diffusion model to "denoise" the artifacts introduced by quantization. However, we identify three critical gaps in previous approaches following this paradigm (namely, the noise level, noise type, and discretization gaps) that result in the quantized data falling out of the data distribution known by the diffusion model. In this work, we propose a novel quantization-based forward diffusion process with theoretical foundations that tackles all three aforementioned gaps. We achieve this through universal quantization with a carefully tailored quantization schedule and a diffusion model trained with uniform noise. Compared to previous work, our proposal produces consistently realistic and detailed reconstructions, even at very low bitrates. In such a regime, we achieve the best rate-distortion-realism performance, outperforming previous related works.
Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers
CVPR2
2025 Spatiotemporal Diffusion Priors for Extreme Video Compression
abstract
Diffusion models have recently demonstrated impressive results in image compression, where the strong spatial prior enables the synthesis of fine details rather than allocating bits to transmit them. In this work, we propose to extend this paradigm to video compression by utilizing a generative spatiotemporal prior and present the first codec based on a video diffusion model. Our method operates by performing longcontext interpolation guided by sparse inter-frame predictions, thus requiring minimal motion information. To this end, we develop a sparse, bidirectional optical flow which serves as a bitrate-efficient motion conditioning in the diffusion decoding process. The resulting codec can compress videos to extremely low rates (as low as 0.01 bits per pixel) while maintaining realistic textures and motion, and outperforms both neural and traditional baselines on several benchmark datasets. Our method shows state-of-the art performance in perceptually-oriented distortion metrics, and, when considering rate-realism, we achieve an improvement in FID score of up to 73.3 at the same bitrate compared to the leading traditional video codec, VTM. Overall, we present an important first work examining spatiotemporal diffusion priors for video compression.
Lucas Relic, André Emmenegger, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Christopher Schroers
PCS3
2024 Combining Frame and GOP Embeddings for Neural Video Representation
abstract
Implicit neural representations (INRs) were recently proposed as a new video compression paradigm, with existing approaches performing on par with HEVC. However, such methods only perform well in limited settings, e.g., specific model sizes, fixed aspect ratios, and low-motion videos. We address this issue by proposing T-NeRV, a hybrid video INR that combines framespecific embeddings with GOP-specific features, providing a lever for content-specific fine-tuning. We employ entropy-constrained training to jointly optimize our model for rate and distortion and demonstrate that T-NeRV can thereby automatically adjust this lever during training, effectively fine-tuning itself to the target content. We evaluate T-NeRVon the UVG dataset, where it achieves state-of-the-art results on the video representation task, outperforming previous works by up to 3dB PSNR on challenging high-motion sequences. Further, our method improves on the compression performance of pre vious methods and is the first video INR to outperform HEVC on all UVG sequences.
Jens Eirik Saethre, Roberto Azevedo, Christopher Schroers
CVPR2
2024 Lossy Image Compression with Foundation Diffusion Models
abstract
Abstract Incorporating diffusion models in the image compression domain has the potential to produce realistic and detailed reconstructions, especially at extremely low bitrates. Previous methods focus on using diffusion models as expressive decoders robust to quantization errors in the conditioning signals. However, achieving competitive results in this manner requires costly training of the diffusion model and long inference times due to the iterative generative process. In this work we formulate the removal of quantization error as a denoising task, using diffusion to recover lost information in the transmitted image latent. Our approach allows us to perform less than 10% of the full diffusion generative process and requires no architectural changes to the diffusion model, enabling the use of foundation models as a strong prior without additional fine tuning of the backbone. Our proposed codec outperforms previous methods in quantitative realism metrics, and we verify that our reconstructions are qualitatively preferred by end users, even when other methods use twice the bitrate.
Lucas Relic, Roberto Azevedo, Markus Gross 0001, Christopher Schroers
ECCV (61)2
2024 Efficient Video Encoder Autotuning via Offline Bayesian Optimization and Supervised Learning
abstract
Modern video encoders are complex software containing dozens of parameters, which allows them to be configured to different scenarios, requirements, or specific titles or scenes. Besides the number of parameters, the inter-dependency between them adds to the complexity of finding a per-title optimized combination of encoding parameters. Even though good practices in the industry have emerged, with the definition of presets per content type (e.g., film vs. cartoon), such practices are suboptimal for specific titles or scenes. Indeed, finding the best encoding parameters for a piece of content is currently a mix of best practices and trial-and-error artwork. We propose an efficient video encoder autotuner based on offline Bayesian optimization and supervised machine learning. Our proposal uses Bayesian optimization to search for a per-title best encoding parameter set offline to generate a dataset. Then, we use the generated dataset to train machine learning models that can map features extracted from the content to the best encoding parameters. Our experiments show that our generated dataset can find a combination of parameters that improves up to approximately −14.49% BD-Rate (0.77 BD-PSNR) and −11.59% BD-Rate (2.12 BD-VMAF) when optimizing for PSNR and VMAF, respectively. In comparison, our prediction models can recover ~80% of such performance while requiring only one fast encoding (compared to hundreds of encodes of a search optimization).
Roberto Azevedo, Yuanyi Xue, Xuewei Meng, Scott Labrozzi, Christopher Schroers
MMSP1
2023 Video Compression with Entropy-Constrained Neural Representations
abstract
Encoding videos as neural networks is a recently proposed approach that allows new forms of video processing. However, traditional techniques still outperform such neural video representation (NVR) methods for the task of video compression. This performance gap can be explained by the fact that current NVR methods: i) use architectures that do not efficiently obtain a compact representation of temporal and spatial information; and ii) minimize rate and distortion disjointly (first overfitting a network on a video and then using heuristic techniques such as post-training quantization or weight pruning to compress the model). We propose a novel convolutional architecture for video representation that better represents spatio-temporal information and a training strategy capable of jointly optimizing rate and distortion. All network and quantization parameters are jointly learned end-to-end, and the post-training operations used in previous works are unnecessary. We evaluate our method on the UVG dataset, achieving new state-of-the-art results for video compression with NVRs. Moreover, we deliver the first NVR-based video compression method that improves over the typically adopted HEVC benchmark (x265, disabled b-frames, “medium” preset), closing the gap to autoencoder-based video compression techniques.
Carlos Gomes, Roberto Azevedo, Christopher Schroers
CVPR2
2023 Neural Video Compression with Spatio-Temporal Cross-Covariance Transformers
abstract
Although existing neural video compression~(NVC) methods have achieved significant success, most of them focus on improving either temporal or spatial information separately. They generally use simple operations such as concatenation or subtraction to utilize this information, while such operations only partially exploit spatio-temporal redundancies. This work aims to effectively and jointly leverage robust temporal and spatial information by proposing a new 3D-based transformer module: Spatio-Temporal Cross-Covariance Transformer (ST-XCT). The ST-XCT module combines two individual extracted features into a joint spatio-temporal feature, followed by 3D convolutional operations and a novel spatio-temporal-aware cross-covariance attention mechanism. Unlike conventional transformers, the cross-covariance attention mechanism is applied across the feature channels without breaking down the spatio-temporal features into local tokens. Such design allows for modeling global cross-channel correlations of the spatio-temporal context while lowering the computational requirement. Based on ST-XCT, we introduce a novel transformer-based end-to-end optimized NVC framework. ST-XCT-based modules are integrated into various key coding components of NVC, such as feature extraction, frame reconstruction, and entropy modeling, demonstrating its generalizability. Extensive experiments show that our ST-XCT-based NVC proposal achieves state-of-the-art compression performances on various standard video benchmark datasets.
Lucas Relic, Roberto Azevedo, Yang Zhang 0003, Markus Gross 0001, Dong Xu 0001, Luping Zhou, Christopher Schroers
ACM Multimedia3
2023 Large-Scale Multi-Site Subjective Assessment on Image Banding Artifacts
abstract
Banding largely occurs due to the finite bit depth representation of digital media and displays. It is a unique type of artifact that is both challenging to accurately measure and difficult to properly mitigate. These artifacts can also be observed across a wide spectrum of streaming services, no matter if professional studio productions or user-generated content. Inspired by the need to monitor and measure these banding artifacts, we present a large-scale multi-site subjective study on image banding artifacts. We designed a novel two-question-based subjective evaluation protocol that goes beyond giving ratings at the frame level, but also collects subjective opinions at the sub-frame level. The high-quality subjective data are collected from a cohort of 56 participants across 3 physical sites with calibrated environments, covering diverse backgrounds and experiences. Finally, we demonstrate that none of the common banding metrics are highly correlated with the subjective data. The paper calls for a need of continued efforts in modeling the perceptual effect of banding artifacts.
Yuanyi Xue, Roberto Azevedo, Xuchang Huangfu, Yang Zhang 0003, Christopher Schroers, Scott Labrozzi
QoMEX2
2021 Microdosing: Knowledge Distillation for GAN Based Compression
Leonhard Helminger, Roberto Azevedo, Abdelaziz Djelouah, Markus Gross 0001, Christopher Schroers
BMVC2