EDBT 2026 Demo / reviewers in the wild / expert
Luka Murn
dblp:267/1405
· DBLP profile ↗
14ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0001-9041-647XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Camera Motion-Conditioned Motion Estimation for Neural Video Coding for Cloud GamingabstractThe rising popularity of cloud gaming highlights the importance of effective video compression for this domain. Despite the potential of neural video codecs to surpass traditional codecs, their application to cloud gaming videos generally achieves limited performance primarily due to two reasons: (1) most advanced neural video codecs are designed for natural videos with only a few optimized for cloud gaming content, and (2) the distinctive characteristics of cloud gaming videos including abrupt and large camera movements, coupled with repetitive textures, make the direct application of existing codecs suboptimal. To this end, this paper introduces an enhanced neural video codec for cloud gaming with camera motion-conditioned motion estimation. By leveraging the unique feature of cloud gaming, we effectively utilize the camera motion information to condition motion estimation with an attention mechanism, overcoming the pronounced challenges of motion estimation and facilitating the accurate learning of optical flow. Furthermore, we optimize the loss function of the motion estimation network by employing a comprehensive loss across multiple feature levels, ensuring the learned optical flow is multi-faceted and well-suited to subsequent motion compensation. Extensive experiments demonstrate the effectiveness of the proposed codec, providing improved motion estimation capabilities and superior rate-distortion performance. Moreover, the robustness of our codec to rapid camera movements is validated, which makes it highly suitable for cloud gaming scenarios. Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Multim. | 3 |
| 2025 | Enhanced neural video compression for cloud gaming videos with aligned frame generationabstractThe burgeoning popularity of cloud gaming makes it critical for efficient video compression to relieve the growing bandwidth pressure. While existing neural video coding approaches have demonstrated strong compression potential on natural videos, there is an absence of efficient neural codecs dedicated to gaming videos. To bridge this gap, in this paper, we propose an end-to-end neural video compression method designed specifically for cloud gaming videos. By effectively utilizing the unique camera motion information inherent to cloud gaming, the previous reconstructed frame is maximally aligned to the current frame through a learningbased module with multiple losses, which then replaces the previous reconstructed frame for optical flow estimation. By significantly reducing the displacement between two consecutive frames caused by camera motion, the motion estimation accuracy is enhanced, effectively handling the large and abrupt motion scenarios frequently present in gaming videos. Furthermore, the aligned tensor obtained in the previous step is used to enhance the latent prior of the entropy model, providing a superior temporal prior for coding. Extensive experimental results demonstrate the superior performance of our proposed method compared to one of the previous state-of-the-art approaches, DCVC-HEM, providing significant progress in end-to-end neural compression in cloud gaming videos Fei Yang 0004, Luka Murn, Juil Sock, Marc Gorriz, Shuai Wan, Wei Zhang 0072, Fuzheng Yang 0001, Luis Herranz |
Expert Syst. Appl. | 3 |
| 2025 | Curriculum learning-based slimmable cross-component prediction for video coding
Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Juil Sock, Fei Yang 0004, Luis Herranz |
Neurocomputing | 4 |
| 2025 | Lightweight Deep Exemplar Colorization via Semantic Attention-Guided Laplacian PyramidabstractExemplar-based colorization aims to generate plausible colors for a grayscale image with the guidance of a color reference image. The main challenging problem is finding the correct semantic correspondence between the target image and the reference image. However, the colors of the object and background are often confused in the existing methods. Besides, these methods usually use simple encoder-decoder architectures or pyramid structures to extract features and lack appropriate fusion mechanisms, which results in the loss of high-frequency information or high complexity. To address these problems, this article proposes a lightweight semantic attention-guided Laplacian pyramid network (SAGLP-Net) for deep exemplar-based colorization, exploiting the inherent multi-scale properties of color representations. They are exploited through a Laplacian pyramid, and semantic information is introduced as high-level guidance to align the object and background information. Specially, a semantic guided non-local attention fusion module is designed to exploit the long-range dependency and fuse the local and global features. Moreover, a Laplacian pyramid fusion module based on criss-cross attention is proposed to fuse high frequency components in the large-scale domain. An unsupervised multi-scale multi-loss training strategy is further introduced for network training, which combines pixel loss, color histogram loss, total variance regularisation, and adversarial loss. Experimental results demonstrate that our colorization method achieves better subjective and objective performance with lower complexity than the state-of-the-art methods. Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Marta Mrak, Juil Sock, Fei Yang 0004, Luis Herranz |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | A slimmable framework for practical neural video compressionabstractDeep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs. Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz |
Neurocomputing | 5 |
| 2024 | Task-Switchable Pre-Processor for Image Compression for Multiple Machine Vision TasksabstractVisual content is increasingly being processed by machines for various automated content analysis tasks instead of being consumed by humans. Despite the existence of several compression methods tailored for machine tasks, few consider real-world scenarios with multiple tasks. In this paper, we aim to address this gap by proposing a task-switchable pre-processor that optimizes input images specifically for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. The proposed task-switchable pre-processor adeptly maintains relevant semantic information based on the specific characteristics of different downstream tasks, while effectively suppressing irrelevant information to reduce bitrate. To enhance the processing of semantic information for diverse tasks, we leverage pre-extracted semantic features to modulate the pixel-to-pixel mapping within the pre-processor. By switching between different modulations, multiple tasks can be seamlessly incorporated into the system. Extensive experiments demonstrate the practicality and simplicity of our approach. It significantly reduces the number of parameters required for handling multiple tasks while still delivering impressive performance. Our method showcases the potential to achieve efficient and effective compression for machine vision tasks, supporting the evolving demands of real-world applications. Mingyi Yang, Fei Yang 0004, Luka Murn, Marc Gorriz, Juil Sock, Shuai Wan, Fuzheng Yang 0001, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Efficient Super-Resolution for Compression Of Gaming VideosabstractDue to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming videos at low bitrates. Most state-of-the-art super-resolution networks generalize over a variety of RGB inputs and use a unified network architecture for frames of different levels of degradation, leading to high complexity and redundancy. Since games usually consist of a limited number of fixed scenarios, we specialize one model for each scenario and assign appropriate network capacities for different QPs to perform super-resolution under the guidance of reconstructed high-quality luma components. Experimental results show that our framework achieves a superior quality-complexity trade-off compared to the ESRnet baseline, saving at most 93.59% parameters while maintaining comparable performance. The compression efficiency compared to HEVC is also improved by more than 17% BD-rate gain. Luka Murn, Luis Herranz, Fei Yang 0004, Marta Mrak, Wei Zhang 0072, Shuai Wan, Marc Gorriz |
ICASSP | 2 |
| 2023 | Semantic Preprocessor for Image Compression for MachinesabstractVisual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision. Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak |
ICASSP | 4 |
| 2023 | Query-Based Video Summarization with Pseudo Label SupervisionabstractExisting datasets for manually labelled query-based video summarization are costly and thus small, limiting the performance of supervised deep video summarization models. Self-supervision can address the data sparsity challenge by using a pretext task and defining a method to acquire extra data with pseudo labels to pre-train a supervised deep model. In this work, we introduce segment-level pseudo labels from input videos to properly model both the relationship between a pretext task and a target task, and the implicit relationship between the pseudo label and the human-defined label. The pseudo labels are generated based on existing human-defined frame-level labels. To create more accurate query-dependent video summaries, a semantics booster is proposed to generate context-aware query representations. Furthermore, we propose mutual attention to help capture the interactive information between visual and textual modalities. Three commonly-used video summarization benchmarks are used to thoroughly validate the proposed approach. Experimental results show that the proposed video summarization algorithm achieves state-of-the-art performance. Jia-Hong Huang, Luka Murn, Marta Mrak, Marcel Worring |
ICIP | 2 |
| 2023 | Efficient Convolution and Transformer-based Network for Video Frame InterpolationabstractVideo frame interpolation is an increasingly important research task with several key industrial applications in the video coding, broadcast and production sectors. Recently, transformers have been introduced to the field resulting in substantial performance gains. However, this comes at a cost of greatly increased memory usage, training and inference time. In this paper, a novel method integrating a transformer encoder and convolutional features is proposed. This network reduces the memory burden by close to 50% and runs up to four times faster during inference time compared to existing transformer-based interpolation methods. A dual-encoder architecture is introduced which combines the strength of convolutions in modelling local correlations with those of the transformer for long-range dependencies. Quantitative evaluations are conducted on various benchmarks with complex motion to showcase the robustness of the proposed method, achieving competitive performance compared to state-of-the-art interpolation networks. Issa Khalifeh, Luka Murn, Marta Mrak, Ebroul Izquierdo |
ICIP | 2 |
| 2022 | Complexity Reduction of Learned In-Loop Filtering in Video CodingabstractIn video coding, in-loop filters are applied on reconstructed video frames to enhance their perceptual quality, before storing the frames for output. Conventional in-loop filters are obtained by hand-crafted methods. Recently, learned filters based on convolutional neural networks that utilize attention mechanisms have been shown to improve upon traditional techniques. However, these solutions are typically significantly more computationally expensive, limiting their potential for practical applications. The proposed method uses a novel combination of sparsity and structured pruning for complexity reduction of learned in-loop filters. This is done through a three-step training process of magnitude-guided weight pruning, insignificant neuron identification and removal, and fine-tuning. Through initial tests we find that network parameters can be significantly reduced with a minimal impact on network performance. Woody Bayliss, Luka Murn, Ebroul Izquierdo, Qianni Zhang, Marta Mrak |
ISCAS | 2 |
| 2021 | GPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video SummarizationabstractTraditional video summarization methods generate fixed video representations regardless of user interest. Therefore such methods limit users' expectations in content search and exploration scenarios. Multi-modal video summarization is one of the methods utilized to address this problem. When multi-modal video summarization is used to help video exploration, a text-based query is considered as one of the main drivers of video summary generation, as it is user-defined. Thus, encoding both the text-based query and the video effectively is important for the task of multi-modal video summarization. In this work, a new method is proposed that uses a specialized attention network and contextualized word representations to tackle this task. The proposed model consists of a contextualized video summary controller, multi-modal attention mechanisms, an interactive attention network, and a video summary generator. Based on the evaluation of the existing multi-modal video summarization benchmark, experimental results show that the proposed model is effective with the increase of +5.88% in accuracy and +4.06% increase of F1-score, compared with the state-of-the-art method. https://github.com/Jhhuangkay/GPT2MVS-Generative-Pre-trained-Transformer-2-for-Multi-modal-Video-Summarization. Jia-Hong Huang, Luka Murn, Marta Mrak, Marcel Worring |
ICMR | 2 |
| 2021 | Interpreting Super-Resolution CNNs for Sub-Pixel Motion Compensation in Video CodingabstractMachine learning approaches for more efficient video compression have been developed thanks to breakthroughs in deep learning. However, they typically bring coding improvements at the cost of significant increases in computational complexity, making them largely unsuitable for practical applications. In this paper, we present open-source software for convolutional neural network-based solutions which improve the interpolation of reference samples needed for fractional precision motion compensation. Contrary to previous efforts, the networks are fully linear, allowing them to be interpreted, with a full interpolation filter set derived from trained models, making it simple to integrate in conventional video coding schemes. When implemented in the context of the state-of-the-art Versatile Video Coding (VVC) test model, the complexity of the learned interpolation schemes is significantly reduced compared to the interpolation with full neural networks, while achieving notable coding efficiency improvements on lower resolution video sequences. The open-source software package is available at https://github.com/bbc/cnn-fractional-motion-compensation under the 3-clause BSD license. Luka Murn, Alan F. Smeaton, Marta Mrak |
ACM Multimedia | 1 |
| 2020 | Interpreting CNN For Low Complexity Learned Sub-Pixel Motion Compensation In Video CodingabstractDeep learning has shown great potential in image and video compression tasks. However, it brings bit savings at the cost of significant increases in coding complexity, which limits its potential for implementation within practical applications. In this paper, a novel neural network-based tool is presented which improves the interpolation of reference samples needed for fractional precision motion compensation. Contrary to previous efforts, the proposed approach focuses on complexity reduction achieved by interpreting the interpolation filters learned by the networks. When the approach is implemented in the Versatile Video Coding (VVC) test model, up to 4.5% BD-rate saving for individual sequences is achieved compared with the baseline VVC, while the complexity of learned interpolation is significantly reduced compared to the application of full neural network. Luka Murn, Saverio G. Blasi, Alan F. Smeaton, Noel E. O'Connor, Marta Mrak |
ICIP | 1 |