EDBT 2026 Demo / reviewers in the wild / expert
Marc Gorriz
dblp:210/2508 · also Marc Górriz, Marc Górriz Blanch
· DBLP profile ↗
16ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-7636-4981ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Camera Motion-Conditioned Motion Estimation for Neural Video Coding for Cloud GamingabstractThe rising popularity of cloud gaming highlights the importance of effective video compression for this domain. Despite the potential of neural video codecs to surpass traditional codecs, their application to cloud gaming videos generally achieves limited performance primarily due to two reasons: (1) most advanced neural video codecs are designed for natural videos with only a few optimized for cloud gaming content, and (2) the distinctive characteristics of cloud gaming videos including abrupt and large camera movements, coupled with repetitive textures, make the direct application of existing codecs suboptimal. To this end, this paper introduces an enhanced neural video codec for cloud gaming with camera motion-conditioned motion estimation. By leveraging the unique feature of cloud gaming, we effectively utilize the camera motion information to condition motion estimation with an attention mechanism, overcoming the pronounced challenges of motion estimation and facilitating the accurate learning of optical flow. Furthermore, we optimize the loss function of the motion estimation network by employing a comprehensive loss across multiple feature levels, ensuring the learned optical flow is multi-faceted and well-suited to subsequent motion compensation. Extensive experiments demonstrate the effectiveness of the proposed codec, providing improved motion estimation capabilities and superior rate-distortion performance. Moreover, the robustness of our codec to rapid camera movements is validated, which makes it highly suitable for cloud gaming scenarios. Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Multim. | 5 |
| 2025 | Towards Reliable Identification of Diffusion-based Image ManipulationsabstractChanging facial expressions, gestures, or background details may dramatically alter the meaning conveyed by an image.
Notably, recent advances in diffusion models greatly improve the quality of image manipulation while also opening the door to misuse.
Identifying changes made to authentic images, thus, becomes an important task, constantly challenged by new diffusion-based editing tools.
To this end, we propose a novel approach for ReliAble iDentification of inpainted AReas (RADAR).
RADAR builds on existing foundation models and combines features from different image modalities.
It also incorporates an auxiliary contrastive loss that helps to isolate manipulated image patches.
We demonstrate these techniques to significantly improve both the accuracy of our method and its generalisation to a large number of diffusion models.
To support realistic evaluation, we further introduce BBC-PAIR, a new comprehensive benchmark, with images tampered by 28 diffusion models.
Our experiments show that RADAR achieves excellent results, outperforming the state-of-the-art in detecting and localising image edits made by both seen and unseen diffusion models.
Further information about our code, data and models, including separate licensing terms, will be publicly available at https://alex-costanzino.github.io/radar/. Alex Costanzino, Woody Bayliss, Juil Sock, Marc Gorriz, Danijela Horak, Ivan Laptev, Philip Torr 0001, Fabio Pizzati |
NeurIPS | 4 |
| 2025 | Enhanced neural video compression for cloud gaming videos with aligned frame generationabstractThe burgeoning popularity of cloud gaming makes it critical for efficient video compression to relieve the growing bandwidth pressure. While existing neural video coding approaches have demonstrated strong compression potential on natural videos, there is an absence of efficient neural codecs dedicated to gaming videos. To bridge this gap, in this paper, we propose an end-to-end neural video compression method designed specifically for cloud gaming videos. By effectively utilizing the unique camera motion information inherent to cloud gaming, the previous reconstructed frame is maximally aligned to the current frame through a learningbased module with multiple losses, which then replaces the previous reconstructed frame for optical flow estimation. By significantly reducing the displacement between two consecutive frames caused by camera motion, the motion estimation accuracy is enhanced, effectively handling the large and abrupt motion scenarios frequently present in gaming videos. Furthermore, the aligned tensor obtained in the previous step is used to enhance the latent prior of the entropy model, providing a superior temporal prior for coding. Extensive experimental results demonstrate the superior performance of our proposed method compared to one of the previous state-of-the-art approaches, DCVC-HEM, providing significant progress in end-to-end neural compression in cloud gaming videos Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
Expert Syst. Appl. | 5 |
| 2025 | Curriculum learning-based slimmable cross-component prediction for video coding
Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Juil Sock, Fei Yang 0004, Luis Herranz |
Neurocomputing | 3 |
| 2025 | Lightweight Deep Exemplar Colorization via Semantic Attention-Guided Laplacian PyramidabstractExemplar-based colorization aims to generate plausible colors for a grayscale image with the guidance of a color reference image. The main challenging problem is finding the correct semantic correspondence between the target image and the reference image. However, the colors of the object and background are often confused in the existing methods. Besides, these methods usually use simple encoder-decoder architectures or pyramid structures to extract features and lack appropriate fusion mechanisms, which results in the loss of high-frequency information or high complexity. To address these problems, this article proposes a lightweight semantic attention-guided Laplacian pyramid network (SAGLP-Net) for deep exemplar-based colorization, exploiting the inherent multi-scale properties of color representations. They are exploited through a Laplacian pyramid, and semantic information is introduced as high-level guidance to align the object and background information. Specially, a semantic guided non-local attention fusion module is designed to exploit the long-range dependency and fuse the local and global features. Moreover, a Laplacian pyramid fusion module based on criss-cross attention is proposed to fuse high frequency components in the large-scale domain. An unsupervised multi-scale multi-loss training strategy is further introduced for network training, which combines pixel loss, color histogram loss, total variance regularisation, and adversarial loss. Experimental results demonstrate that our colorization method achieves better subjective and objective performance with lower complexity than the state-of-the-art methods. Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Marta Mrak, Juil Sock, Fei Yang 0004, Luis Herranz |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | A slimmable framework for practical neural video compressionabstractDeep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs. Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz |
Neurocomputing | 4 |
| 2024 | Task-Switchable Pre-Processor for Image Compression for Multiple Machine Vision TasksabstractVisual content is increasingly being processed by machines for various automated content analysis tasks instead of being consumed by humans. Despite the existence of several compression methods tailored for machine tasks, few consider real-world scenarios with multiple tasks. In this paper, we aim to address this gap by proposing a task-switchable pre-processor that optimizes input images specifically for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. The proposed task-switchable pre-processor adeptly maintains relevant semantic information based on the specific characteristics of different downstream tasks, while effectively suppressing irrelevant information to reduce bitrate. To enhance the processing of semantic information for diverse tasks, we leverage pre-extracted semantic features to modulate the pixel-to-pixel mapping within the pre-processor. By switching between different modulations, multiple tasks can be seamlessly incorporated into the system. Extensive experiments demonstrate the practicality and simplicity of our approach. It significantly reduces the number of parameters required for handling multiple tasks while still delivering impressive performance. Our method showcases the potential to achieve efficient and effective compression for machine vision tasks, supporting the evolving demands of real-world applications. Mingyi Yang, Fei Yang 0004, Luka Murn, Marc Gorriz, Juil Sock, Shuai Wan, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Chroma Intra Prediction With Lightweight Attention-Based Neural NetworksabstractNeural networks can be successfully used for cross-component prediction in video coding. In particular, attention-based architectures are suitable for chroma intra prediction using luma information because of their capability to model relations between difierent channels. However, the complexity of such methods is still very high and should be further reduced, especially for decoding. In this paper, a cost-effective attention-based neural network is designed for chroma intra prediction. Moreover, with the goal of further improving coding performance, a novel approach is introduced to utilize more boundary information effectively. In addition to improving prediction, a simplification methodology is also proposed to reduce inference complexity by simplifying convolutions. The proposed schemes are integrated into H.266/Versatile Video Coding (VVC) pipeline, and only one additional binary block-level syntax flag is introduced to indicate whether a given block makes use of the proposed method. Experimental results demonstrate that the proposed scheme achieves up to −0.46%/−2.29%/−2.17% BD-rate reduction on Y/Cb/Cr components, respectively, compared with H.266/VVC anchor. Reductions in the encoding and decoding complexity of up to 22% and 61%, respectively, are achieved by the proposed scheme with respect to the previous attention-based chroma intra prediction method while maintaining coding performance. Chengyi Zou, Shuai Wan, Tiannan Ji, Marc Gorriz, Marta Mrak, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Efficient Super-Resolution for Compression Of Gaming VideosabstractDue to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming videos at low bitrates. Most state-of-the-art super-resolution networks generalize over a variety of RGB inputs and use a unified network architecture for frames of different levels of degradation, leading to high complexity and redundancy. Since games usually consist of a limited number of fixed scenarios, we specialize one model for each scenario and assign appropriate network capacities for different QPs to perform super-resolution under the guidance of reconstructed high-quality luma components. Experimental results show that our framework achieves a superior quality-complexity trade-off compared to the ESRnet baseline, saving at most 93.59% parameters while maintaining comparable performance. The compression efficiency compared to HEVC is also improved by more than 17% BD-rate gain. Luka Murn, Luis Herranz, Fei Yang 0004, Marta Mrak, Wei Zhang 0072, Shuai Wan, Marc Gorriz |
ICASSP | 8 |
| 2023 | Semantic Preprocessor for Image Compression for MachinesabstractVisual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision. Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak |
ICASSP | 5 |
| 2022 | DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed VideoabstractIn this paper, we propose a deformable convolution-based generative adversarial network (DCNGAN) for perceptual quality enhancement of compressed videos. DCNGAN is also adaptive to the quantization parameters (QPs). Compared with optical flows, deformable convolutions are more effective and efficient to align frames. Deformable convolutions can operate on multiple frames, thus leveraging more temporal information, which is beneficial for enhancing the perceptual quality of compressed videos. Instead of aligning frames in a pairwise manner, the deformable convolution can process multiple frames simultaneously, which leads to lower computational complexity. Experimental results demonstrate that the proposed DCNGAN outperforms other state-of-the-art compressed video quality enhancement algorithms. Saiping Zhang, Luis Herranz, Marta Mrak, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
ICASSP | 4 |
| 2022 | Towards Lightweight Neural Network-based Chroma Intra Prediction for Video CodingabstractIn video compression the luma channel can be useful for predicting chroma channels (Cb, Cr), as has been demonstrated with the Cross-Component Linear Model (CCLM) used in Versatile Video Coding (VVC) standard. More recently, it has been shown that neural networks can even better capture the relationship among different channels. In this paper, a new attention-based neural network is proposed for cross-component intra prediction. With the goal to simplify neural network design, the new framework consists of four branches: boundary branch and luma branch for extracting features from reference samples, attention branch for fusing the first two branches, and prediction branch for computing the predicted chroma samples. The proposed scheme is integrated into VVC test model together with one additional binary block-level syntax flag which indicates whether a given block makes use of the proposed method. Experimental results demonstrate 0.31%/2.36%/2.00% BD-rate reductions on Y/Cb/Cr components, respectively, on top of the VVC Test Model (VTM) 7.0 which uses CCLM. Chengyi Zou, Shuai Wan, Marta Mrak, Marc Gorriz, Luis Herranz, Tiannan Ji |
ICIP | 4 |
| 2021 | Attention-based Stylisation for Exemplar Image ColourisationabstractExemplar-based colourisation aims to add plausible colours to a grayscale image using the guidance of a colour reference image. Most existing methods tackle the task as a style transfer problem, using a convolutional neural network (CNN) to obtain deep representations of the content of both inputs. Stylised outputs are then obtained by computing similarities between both feature representations in order to transfer the style of the reference to the content of the target input. However, in order to gain robustness towards dissimilar references, the stylised outputs need to be refined with a second colourisation network, which significantly increases the overall system complexity. This work reformulates the existing methodology introducing a novel end-to-end colourisation network that unifies the feature matching with the colourisation process. The proposed architecture integrates attention modules at different resolutions that learn how to perform the style transfer task in an unsupervised way towards decoding realistic colour predictions. Moreover, axial attention is proposed to simplify the attention operations and to obtain a fast but robust cost-effective architecture. Experimental validations demonstrate efficiency of the proposed methodology which generates high quality and visually appealing colourisation. Furthermore, the complexity of the proposed methodology is reduced compared to the state-of-the-art methods. Marc Gorriz, Issa Khalifeh, Noel E. O'Connor, Marta Mrak |
MMSP | 1 |
| 2021 | DVC-P: Deep Video Compression with Perceptual OptimizationsabstractRecent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual op-timizations (DVC-P), which aims at increasing perceptual quality of decoded videos. Our proposed DVC-P is based on Deep Video Compression (DVC) network, but improves it with perceptual optimizations. Specifically, a discriminator network and a mixed loss are employed to help our network trade off among distortion, perception and rate. Furthermore, nearest-neighbor interpolation is used to eliminate checkerboard artifacts which can appear in sequences encoded with DVC frameworks. Thanks to these two improvements, the perceptual quality of decoded sequences is improved. Experimental results demonstrate that, compared with the baseline DVC, our proposed method can generate videos with higher perceptual quality achieving 12.27% reduction in a perceptual BD- rate equivalent, on average. Saiping Zhang, Marta Mrak, Luis Herranz, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
VCIP | 4 |
| 2020 | Chroma Intra Prediction With Attention-Based CNN ArchitecturesabstractNeural networks can be used in video coding to improve chroma intra-prediction. In particular, usage of fully-connected networks has enabled better cross-component prediction with respect to traditional linear models. Nonetheless, state-of-the-art architectures tend to disregard the location of individual reference samples in the prediction process. This paper proposes a new neural network architecture for cross-component intra-prediction. The network uses a novel attention module to model spatial relations between reference and predicted samples. The proposed approach is integrated into the Versatile Video Coding (VVC) prediction pipeline. Experimental results demonstrate compression gains over the latest VVC anchor compared with state-of-the-art chroma intra-prediction methods based on neural networks. Marc Gorriz, Saverio G. Blasi, Alan F. Smeaton, Noel E. O'Connor, Marta Mrak |
ICIP | 1 |
| 2019 | End-to-End Conditional GAN-based Architectures for Image ColourisationabstractIn this work recent advances in conditional adversarial networks are investigated to develop an end-to-end architecture based on Convolutional Neural Networks (CNNs) to directly map realistic colours to an input greyscale image. Observing that existing colourisation methods sometimes exhibit a lack of colourfulness, this paper proposes a method to improve colourisation results. In particular, the method uses Generative Adversarial Neural Networks (GANs) and focuses on improvement of training stability to enable better generalisation in large multi-class image datasets. Additionally, the integration of instance and batch normalisation layers in both generator and discriminator is introduced to the popular U-Net architecture, boosting the network capabilities to generalise the style changes of the content. The method has been tested using the ILSVRC 2012 dataset, achieving improved automatic colourisation results compared to other methods based on GANs. Marc Gorriz, Marta Mrak, Alan F. Smeaton, Noel E. O'Connor |
MMSP | 1 |