María Santamaría 0001

dblp:147/0758 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2023
0000-0003-1946-0712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2023 Overfitting NN loop-filters in video coding
abstract
Overfitting is usually regarded as a negative condition since it impairs the generalisation power of a model. Nevertheless, overfitting a Neural Network (NN) on test data may be advantageous to improve the compression efficiency of image/video coding tools and systems. Previous research has demonstrated the benefits of NN overfitting for post-processing operations, i.e. post-filters, but not yet for actual decoding tools. Generally, the NN is overfitted on test data at the encoder end, and the weight update is coded and sent to the decoder end along the image/video bitstream. The proposed approach follows this strategy. In particular, the overfitting of the Low Operation Point (LOP) loop-filter in NN-based Video Coding (NNVC) software is studied. The overall approach yields Bjøntegaard Delta rate (BD-rate) of -7.74%, -13.73% and -12.49%, for the Y, U and V components, respectively. Out of these coding gains, 1.21%, 6.43% and 5.52%, for the Y, U and V components, are attributed to the overfitting. The boost in the coding gains comes with only 1.5% more complexity, due to the multiplier parameters introduced during the overfitting.
Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela, Tapio Elomaa
VCIP2
2022 Content-Adaptive Neural Network Post-Processing Filter with NNR-Coded Weight-Updates
abstract
Neural Network (NN) filters improve the perceptual quality of reconstructed videos by reducing compression artefacts. For content adaptation, a few NN-filters use over-fitting. As the adaptation signal is a weight-update, compression is required to minimise significant bitrate overheads. Most approaches, however, use generic data compression algorithms, which are inadequate for coding NN weight-updates. This work introduces a content-adaptive NN post-processing filter with weight-updates coded using the Neural Network compression and Representation (NNR) standard. The bitrate overhead is further decreased by over-fitting only a subset of weights, selected via energy-based analysis. The proposed filter saved about 4.57% (Y), 10.33% (Cb), 6.53% (Cr) Bjøntegaard Delta rate (BD-rate) on top of the Versatile Video Coding (VVC) Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0, in Random Access (RA) configuration. Compared to the non-over-fitted NN, the performance was doubled; and compared to 7z, NNR reduced the bitrate of the weight-update by ∼64%.
María Santamaría 0001, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela
ICIP1
2022 Low-precision post-filtering in video coding
abstract
Neural Networks (NNs) have demonstrated their effectiveness in tackling challenges involving multimedia content. In the video coding field, NNs are actively exploited as novel tools that complement conventional signal processing tools, as well as end-to-end coding solutions. Since NNs use commonly floating-point arithmetic, different results may be generated in different computing environments, leading to discrepancies and even corrupted reconstructions. Accordingly, this issue is solved by employing fixed-point arithmetic instead. This paper studies the quantisation of a 32-bit floating-point (float32) NN post-filter to 32-bit fixed-point (int32) and 16-bit fixed-point (int16). On top of the VVC Test Model (VTM) 11.0 NN-based Video Coding (NNVC) 1.0, the coding gains of the float32 post-filter are 5.01% (Y), 18.95% (Cb) and 17.33% Cr. Compared to the float32 inference, the quantised models produce coding losses: 0.01% (Y, Cb and Cr) for the int32 inference and 0.45% (Y), 1.97% (Cb) and 1.19% (Cr) for the int16 inference. Nevertheless, the fixed-point approaches achieve bit exact matches in different computing environments. Moreover, the decoding time with int16 is about half the decoding time of int32.
Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela
ISM2
2021 Content-adaptive convolutional neural network post-processing filter
abstract
Neural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration.
María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj
ISM1
2021 Coding of volumetric content with MIV using VVC subpictures
abstract
Storage and transport of six degrees of freedom (6DoF) dynamic volumetric visual content for immersive applications requires efficient compression. ISO/IEC MPEG has recently been working on a standard that aims to efficiently code and deliver 6DoF immersive visual experiences. This standard is called the MIV. MIV uses regular 2D video codecs to code the visual data. MPEG jointly with ITU-T VCEG, has also specified the VVC standard. VVC introduced recently the concept of subpicture. This tool was specifically designed to provide independent accessibility and decodability of sub-bitstreams for omnidirectional applications. This paper shows the benefit of using subpictures in the MIV use-case. While different ways in which subpictures could be used in MIV are discussed, a particular case study is selected. Namely, subpictures are used for parallel encoding and to reduce the number of decoder instances. Experimental results show that the cost of using subpictures in terms of bitrate overhead is negligible (0.1% to 0.4%), when compared to the overall bitrate. The number of decoder instances on the other hand decreases by a factor of two.
María Santamaría 0001, Vinod Kumar Malamal Vadakital, Lukasz Kondrad, Antti Hallapuro, Miska M. Hannuksela
MMSP1
2021 Learn to overfit better: finding the important parameters for learned image compression
abstract
For most machine learning systems, overfitting is an undesired behavior. However, overfitting a model to a test image or a video at inference time is a favorable and effective technique to improve the coding efficiency of learning-based image and video codecs. At the encoding stage, one or more neural networks that are part of the codec are finetuned using the input image or video to achieve a better coding performance. The encoder en-codes the input content into a content bitstream. If the finetuned neural network is part (also) of the decoder, the encoder signals the weight update of the finetuned model to the decoder along with the content bitstream. At the decoding stage, the decoder first updates its neural network model according to the received weight update, and then proceeds with decoding the content bitstream. Since a neural network contains a large number of parameters, compressing the weight update is critical to reducing bitrate overhead. In this paper, we propose learning-based methods to find the important parameters to be overfitted, in terms of rate-distortion performance. Based on simple distribution models for variables in the weight update, we derive two objective functions. By optimizing the proposed objective functions, the importance scores of the parameters can be calculated and the important parameters can be determined. Our experiments on lossless image compression codec show that the proposed method significantly outperforms a prior-art method where overfitted parameters were selected based on heuristics. Furthermore, our technique improved the compression performance of the state-of-the-art lossless image compression codec by 0.1 bit per pixel.
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, María Santamaría 0001, Yat-Hong Lam, Miska M. Hannuksela
VCIP4
2018 Estimation of Rate Control Parameters for Video Coding Using CNN
abstract
Rate-control is essential to ensure efficient video delivery. Typical rate-control algorithms rely on bit allocation strategies, to appropriately distribute bits among frames. As reference frames are essential for exploiting temporal redundancies, intra frames are usually assigned a larger portion of the available bits. In this paper, an accurate method to estimate number of bits and quality of intra frames is proposed, which can be used for bit allocation in a rate-control scheme. The algorithm is based on deep learning, where networks are trained using the original frames as inputs, while distortions and sizes of compressed frames after encoding are used as ground truths. Two approaches are proposed where either local or global distortions are predicted.
María Santamaría 0001, Ebroul Izquierdo, Saverio G. Blasi, Marta Mrak
VCIP1
2014 Edge-Based Coding Tree Unit Partitioning Strategy in Inter Prediction
María Santamaría 0001, María Trujillo
CIARP1