EDBT 2026 Demo / reviewers in the wild / expert
Marta Mrak
dblp:79/4830
· DBLP profile ↗
70ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-7777-0252ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lightweight Deep Exemplar Colorization via Semantic Attention-Guided Laplacian PyramidabstractExemplar-based colorization aims to generate plausible colors for a grayscale image with the guidance of a color reference image. The main challenging problem is finding the correct semantic correspondence between the target image and the reference image. However, the colors of the object and background are often confused in the existing methods. Besides, these methods usually use simple encoder-decoder architectures or pyramid structures to extract features and lack appropriate fusion mechanisms, which results in the loss of high-frequency information or high complexity. To address these problems, this article proposes a lightweight semantic attention-guided Laplacian pyramid network (SAGLP-Net) for deep exemplar-based colorization, exploiting the inherent multi-scale properties of color representations. They are exploited through a Laplacian pyramid, and semantic information is introduced as high-level guidance to align the object and background information. Specially, a semantic guided non-local attention fusion module is designed to exploit the long-range dependency and fuse the local and global features. Moreover, a Laplacian pyramid fusion module based on criss-cross attention is proposed to fuse high frequency components in the large-scale domain. An unsupervised multi-scale multi-loss training strategy is further introduced for network training, which combines pixel loss, color histogram loss, total variance regularisation, and adversarial loss. Experimental results demonstrate that our colorization method achieves better subjective and objective performance with lower complexity than the state-of-the-art methods. Chengyi Zou, Shuai Wan, Marc Gorriz, Luka Murn, Marta Mrak, Juil Sock, Fei Yang 0004, Luis Herranz |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | A slimmable framework for practical neural video compressionabstractDeep learning is being increasingly applied to image and video compression in a new paradigm known as neural video compression. While achieving impressive rate–distortion (RD) performance, neural video codecs (NVC) require heavy neural networks, which in turn have large memory and computational costs and often lack important functionalities such as variable rate. These are significant limitations to their practical application. Addressing these problems, recent slimmable image codecs can dynamically adjust their model capacity to elegantly reduce the memory and computation requirements, without harming RD performance. However, the extension to video is not straightforward due to the non-trivial interplay with complex motion estimation and compensation modules in most NVC architectures. In this paper we propose the slimmable video codec framework (SlimVC) that integrates an slimmable autoencoder and a motion-free conditional entropy model. We show that the slimming mechanism is also applicable to the more complex case of video architectures, providing SlimVC with simultaneous control of the computational cost, memory and rate, which are all important requirements in practice. We further provide detailed experimental analysis, and describe application scenarios that can benefit from slimmable video codecs. Zhaocheng Liu, Fei Yang 0004, Defa Wang, Marc Gorriz, Luka Murn, Shuai Wan, Saiping Zhang, Marta Mrak, Luis Herranz |
Neurocomputing | 8 |
| 2024 | Chroma Intra Prediction With Lightweight Attention-Based Neural NetworksabstractNeural networks can be successfully used for cross-component prediction in video coding. In particular, attention-based architectures are suitable for chroma intra prediction using luma information because of their capability to model relations between difierent channels. However, the complexity of such methods is still very high and should be further reduced, especially for decoding. In this paper, a cost-effective attention-based neural network is designed for chroma intra prediction. Moreover, with the goal of further improving coding performance, a novel approach is introduced to utilize more boundary information effectively. In addition to improving prediction, a simplification methodology is also proposed to reduce inference complexity by simplifying convolutions. The proposed schemes are integrated into H.266/Versatile Video Coding (VVC) pipeline, and only one additional binary block-level syntax flag is introduced to indicate whether a given block makes use of the proposed method. Experimental results demonstrate that the proposed scheme achieves up to −0.46%/−2.29%/−2.17% BD-rate reduction on Y/Cb/Cr components, respectively, compared with H.266/VVC anchor. Reductions in the encoding and decoding complexity of up to 22% and 61%, respectively, are achieved by the proposed scheme with respect to the previous attention-based chroma intra prediction method while maintaining coding performance. Chengyi Zou, Shuai Wan, Tiannan Ji, Marc Gorriz, Marta Mrak, Luis Herranz |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Efficient Super-Resolution for Compression Of Gaming VideosabstractDue to the increasing demand for game-streaming services, efficient compression of computer-generated video is more critical than ever, especially when the available bandwidth is low. This paper proposes a super-resolution framework that improves the coding efficiency of computer-generated gaming videos at low bitrates. Most state-of-the-art super-resolution networks generalize over a variety of RGB inputs and use a unified network architecture for frames of different levels of degradation, leading to high complexity and redundancy. Since games usually consist of a limited number of fixed scenarios, we specialize one model for each scenario and assign appropriate network capacities for different QPs to perform super-resolution under the guidance of reconstructed high-quality luma components. Experimental results show that our framework achieves a superior quality-complexity trade-off compared to the ESRnet baseline, saving at most 93.59% parameters while maintaining comparable performance. The compression efficiency compared to HEVC is also improved by more than 17% BD-rate gain. Luka Murn, Luis Herranz, Fei Yang 0004, Marta Mrak, Wei Zhang 0072, Shuai Wan, Marc Gorriz |
ICASSP | 5 |
| 2023 | Semantic Preprocessor for Image Compression for MachinesabstractVisual content is being increasingly transmitted and consumed by machines rather than humans to perform automated content analysis tasks. In this paper, we propose an image preprocessor that optimizes the input image for machine consumption prior to encoding by an off-the-shelf codec designed for human consumption. To achieve a better trade-off between the accuracy of the machine analysis task and bitrate, we propose leveraging pre-extracted semantic information to improve the preprocessor’s ability to accurately identify and filter out task-irrelevant information. Furthermore, we propose a two-part loss function to optimize the preprocessor, consisted of a rate-task performance loss and a semantic distillation loss, which helps the reconstructed image obtain more information that contributes to the accuracy of the task. Experiments show that the proposed preprocessor can save up to 48.83% bitrate compared with the method without the preprocessor, and save up to 36.24% bitrate compared to existing preprocessors for machine vision. Mingyi Yang, Luis Herranz, Fei Yang 0004, Luka Murn, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001, Marta Mrak |
ICASSP | 8 |
| 2023 | Query-Based Video Summarization with Pseudo Label SupervisionabstractExisting datasets for manually labelled query-based video summarization are costly and thus small, limiting the performance of supervised deep video summarization models. Self-supervision can address the data sparsity challenge by using a pretext task and defining a method to acquire extra data with pseudo labels to pre-train a supervised deep model. In this work, we introduce segment-level pseudo labels from input videos to properly model both the relationship between a pretext task and a target task, and the implicit relationship between the pseudo label and the human-defined label. The pseudo labels are generated based on existing human-defined frame-level labels. To create more accurate query-dependent video summaries, a semantics booster is proposed to generate context-aware query representations. Furthermore, we propose mutual attention to help capture the interactive information between visual and textual modalities. Three commonly-used video summarization benchmarks are used to thoroughly validate the proposed approach. Experimental results show that the proposed video summarization algorithm achieves state-of-the-art performance. Jia-Hong Huang, Luka Murn, Marta Mrak, Marcel Worring |
ICIP | 3 |
| 2023 | Efficient Convolution and Transformer-based Network for Video Frame InterpolationabstractVideo frame interpolation is an increasingly important research task with several key industrial applications in the video coding, broadcast and production sectors. Recently, transformers have been introduced to the field resulting in substantial performance gains. However, this comes at a cost of greatly increased memory usage, training and inference time. In this paper, a novel method integrating a transformer encoder and convolutional features is proposed. This network reduces the memory burden by close to 50% and runs up to four times faster during inference time compared to existing transformer-based interpolation methods. A dual-encoder architecture is introduced which combines the strength of convolutions in modelling local correlations with those of the transformer for long-range dependencies. Quantitative evaluations are conducted on various benchmarks with complex motion to showcase the robustness of the proposed method, achieving competitive performance compared to state-of-the-art interpolation networks. Issa Khalifeh, Luka Murn, Marta Mrak, Ebroul Izquierdo |
ICIP | 3 |
| 2022 | DCNGAN: A Deformable Convolution-Based GAN with QP Adaptation for Perceptual Quality Enhancement of Compressed VideoabstractIn this paper, we propose a deformable convolution-based generative adversarial network (DCNGAN) for perceptual quality enhancement of compressed videos. DCNGAN is also adaptive to the quantization parameters (QPs). Compared with optical flows, deformable convolutions are more effective and efficient to align frames. Deformable convolutions can operate on multiple frames, thus leveraging more temporal information, which is beneficial for enhancing the perceptual quality of compressed videos. Instead of aligning frames in a pairwise manner, the deformable convolution can process multiple frames simultaneously, which leads to lower computational complexity. Experimental results demonstrate that the proposed DCNGAN outperforms other state-of-the-art compressed video quality enhancement algorithms. Saiping Zhang, Luis Herranz, Marta Mrak, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
ICASSP | 3 |
| 2022 | Towards Lightweight Neural Network-based Chroma Intra Prediction for Video CodingabstractIn video compression the luma channel can be useful for predicting chroma channels (Cb, Cr), as has been demonstrated with the Cross-Component Linear Model (CCLM) used in Versatile Video Coding (VVC) standard. More recently, it has been shown that neural networks can even better capture the relationship among different channels. In this paper, a new attention-based neural network is proposed for cross-component intra prediction. With the goal to simplify neural network design, the new framework consists of four branches: boundary branch and luma branch for extracting features from reference samples, attention branch for fusing the first two branches, and prediction branch for computing the predicted chroma samples. The proposed scheme is integrated into VVC test model together with one additional binary block-level syntax flag which indicates whether a given block makes use of the proposed method. Experimental results demonstrate 0.31%/2.36%/2.00% BD-rate reductions on Y/Cb/Cr components, respectively, on top of the VVC Test Model (VTM) 7.0 which uses CCLM. Chengyi Zou, Shuai Wan, Marta Mrak, Marc Gorriz, Luis Herranz, Tiannan Ji |
ICIP | 3 |
| 2022 | Complexity Reduction of Learned In-Loop Filtering in Video CodingabstractIn video coding, in-loop filters are applied on reconstructed video frames to enhance their perceptual quality, before storing the frames for output. Conventional in-loop filters are obtained by hand-crafted methods. Recently, learned filters based on convolutional neural networks that utilize attention mechanisms have been shown to improve upon traditional techniques. However, these solutions are typically significantly more computationally expensive, limiting their potential for practical applications. The proposed method uses a novel combination of sparsity and structured pruning for complexity reduction of learned in-loop filters. This is done through a three-step training process of magnitude-guided weight pruning, insignificant neuron identification and removal, and fine-tuning. Through initial tests we find that network parameters can be significantly reduced with a minimal impact on network performance. Woody Bayliss, Luka Murn, Ebroul Izquierdo, Qianni Zhang, Marta Mrak |
ISCAS | 5 |
| 2021 | GPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video SummarizationabstractTraditional video summarization methods generate fixed video representations regardless of user interest. Therefore such methods limit users' expectations in content search and exploration scenarios. Multi-modal video summarization is one of the methods utilized to address this problem. When multi-modal video summarization is used to help video exploration, a text-based query is considered as one of the main drivers of video summary generation, as it is user-defined. Thus, encoding both the text-based query and the video effectively is important for the task of multi-modal video summarization. In this work, a new method is proposed that uses a specialized attention network and contextualized word representations to tackle this task. The proposed model consists of a contextualized video summary controller, multi-modal attention mechanisms, an interactive attention network, and a video summary generator. Based on the evaluation of the existing multi-modal video summarization benchmark, experimental results show that the proposed model is effective with the increase of +5.88% in accuracy and +4.06% increase of F1-score, compared with the state-of-the-art method. https://github.com/Jhhuangkay/GPT2MVS-Generative-Pre-trained-Transformer-2-for-Multi-modal-Video-Summarization. Jia-Hong Huang, Luka Murn, Marta Mrak, Marcel Worring |
ICMR | 3 |
| 2021 | Assisting News Media Editors with Cohesive Visual StorylinesabstractCreating a cohesive, high-quality, relevant, media story is a challenge that news media editors face on a daily basis. This challenge is aggravated by the flood of highly-relevant information that is constantly pouring onto the newsroom. To assist news media editors in this daunting task, this paper proposes a framework to organize news content into cohesive, high-quality, relevant visual storylines. First, we formalize, in a nonsubjective manner, the concept of visual story transition. Leveraging it, we propose four graph based methods of storyline creation, aiming for global story cohesiveness. These where created and implemented to take full advantage of existing graph algorithms, ensuring their correctness and good computational performance. They leverage a strong ensemble-based estimator which was trained to predict story transition quality based on both the semantic and visual features present in the pair of images under scrutiny. A user study covered a total of 28 curated stories about sports and cultural events. Experiments showed that (i) visual transitions in storylines can be learned with a quality above 90%, and (ii) the proposed graph methods can produce cohesive storylines with a quality in the range of 88% to 96%. Gonçalo Marcelino, David Semedo, André Mourão, Saverio G. Blasi, João Magalhães, Marta Mrak |
ACM Multimedia | 6 |
| 2021 | Interpreting Super-Resolution CNNs for Sub-Pixel Motion Compensation in Video CodingabstractMachine learning approaches for more efficient video compression have been developed thanks to breakthroughs in deep learning. However, they typically bring coding improvements at the cost of significant increases in computational complexity, making them largely unsuitable for practical applications. In this paper, we present open-source software for convolutional neural network-based solutions which improve the interpolation of reference samples needed for fractional precision motion compensation. Contrary to previous efforts, the networks are fully linear, allowing them to be interpreted, with a full interpolation filter set derived from trained models, making it simple to integrate in conventional video coding schemes. When implemented in the context of the state-of-the-art Versatile Video Coding (VVC) test model, the complexity of the learned interpolation schemes is significantly reduced compared to the interpolation with full neural networks, while achieving notable coding efficiency improvements on lower resolution video sequences. The open-source software package is available at https://github.com/bbc/cnn-fractional-motion-compensation under the 3-clause BSD license. Luka Murn, Alan F. Smeaton, Marta Mrak |
ACM Multimedia | 3 |
| 2021 | Attention-based Stylisation for Exemplar Image ColourisationabstractExemplar-based colourisation aims to add plausible colours to a grayscale image using the guidance of a colour reference image. Most existing methods tackle the task as a style transfer problem, using a convolutional neural network (CNN) to obtain deep representations of the content of both inputs. Stylised outputs are then obtained by computing similarities between both feature representations in order to transfer the style of the reference to the content of the target input. However, in order to gain robustness towards dissimilar references, the stylised outputs need to be refined with a second colourisation network, which significantly increases the overall system complexity. This work reformulates the existing methodology introducing a novel end-to-end colourisation network that unifies the feature matching with the colourisation process. The proposed architecture integrates attention modules at different resolutions that learn how to perform the style transfer task in an unsupervised way towards decoding realistic colour predictions. Moreover, axial attention is proposed to simplify the attention operations and to obtain a fast but robust cost-effective architecture. Experimental validations demonstrate efficiency of the proposed methodology which generates high quality and visually appealing colourisation. Furthermore, the complexity of the proposed methodology is reduced compared to the state-of-the-art methods. Marc Gorriz, Issa Khalifeh, Noel E. O'Connor, Marta Mrak |
MMSP | 4 |
| 2021 | DVC-P: Deep Video Compression with Perceptual OptimizationsabstractRecent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual op-timizations (DVC-P), which aims at increasing perceptual quality of decoded videos. Our proposed DVC-P is based on Deep Video Compression (DVC) network, but improves it with perceptual optimizations. Specifically, a discriminator network and a mixed loss are employed to help our network trade off among distortion, perception and rate. Furthermore, nearest-neighbor interpolation is used to eliminate checkerboard artifacts which can appear in sequences encoded with DVC frameworks. Thanks to these two improvements, the perceptual quality of decoded sequences is improved. Experimental results demonstrate that, compared with the baseline DVC, our proposed method can generate videos with higher perceptual quality achieving 12.27% reduction in a perceptual BD- rate equivalent, on average. Saiping Zhang, Marta Mrak, Luis Herranz, Marc Gorriz, Shuai Wan, Fuzheng Yang 0001 |
VCIP | 2 |
| 2021 | User generated content for enhanced professional productions: a mobile application for content contributors and a study on the factors influencing their satisfaction and loyaltyabstractMotivated by a combination of social media, technological evolution, as well as new habits and preferences of TV content consumers, there is an increasing demand for enhancement of professional productions with user generated content. Studies have explored the potential and feasibility of this approach, indicating that footage from non-professionals can be effectively used to enrich the viewing experience. However, an important concern is whether such efforts are appealing to potential contributors, and what can actually impact their satisfaction and loyalty. Aiming to investigate these factors, this paper presents a mobile application for content contributors and a study involving 38 attendees of live events, using the application in the field. The events were hosted in two different countries, and transmitted by two well-known broadcasters. The results suggest that age, gender, technological expertise, and overall sharing attitude do not affect the satisfaction and loyalty of contributors. The differentiating factors, however, are the filming confidence and expertise of contributors, as well as the Wi-Fi/4G connectivity on-site. Implications of these findings are discussed and recommendations for similar endeavors are provided. Stavroula Ntoa, George Margetis, Fiona Rivera, Michael Evans, Ilia Adami, Georgios Mathioudakis, Rajitha Weerakkody, George Metaxakis, Ioannis Markopoulos, Marta Mrak, Constantine Stephanidis |
Multim. Tools Appl. | 10 |
| 2020 | Chroma Intra Prediction With Attention-Based CNN ArchitecturesabstractNeural networks can be used in video coding to improve chroma intra-prediction. In particular, usage of fully-connected networks has enabled better cross-component prediction with respect to traditional linear models. Nonetheless, state-of-the-art architectures tend to disregard the location of individual reference samples in the prediction process. This paper proposes a new neural network architecture for cross-component intra-prediction. The network uses a novel attention module to model spatial relations between reference and predicted samples. The proposed approach is integrated into the Versatile Video Coding (VVC) prediction pipeline. Experimental results demonstrate compression gains over the latest VVC anchor compared with state-of-the-art chroma intra-prediction methods based on neural networks. Marc Gorriz, Saverio G. Blasi, Alan F. Smeaton, Noel E. O'Connor, Marta Mrak |
ICIP | 5 |
| 2020 | Interpreting CNN For Low Complexity Learned Sub-Pixel Motion Compensation In Video CodingabstractDeep learning has shown great potential in image and video compression tasks. However, it brings bit savings at the cost of significant increases in coding complexity, which limits its potential for implementation within practical applications. In this paper, a novel neural network-based tool is presented which improves the interpolation of reference samples needed for fractional precision motion compensation. Contrary to previous efforts, the proposed approach focuses on complexity reduction achieved by interpreting the interpolation filters learned by the networks. When the approach is implemented in the Versatile Video Coding (VVC) test model, up to 4.5% BD-rate saving for individual sequences is achieved compared with the baseline VVC, while the complexity of learned interpolation is significantly reduced compared to the application of full neural network. Luka Murn, Saverio G. Blasi, Alan F. Smeaton, Noel E. O'Connor, Marta Mrak |
ICIP | 5 |
| 2019 | Fast Inter-prediction Based on Decision Trees for AV1 EncodingabstractThe AOMedia Video 1 (AV1) standard can achieve considerable compression efficiency thanks to the usage of many advanced tools and improvements, such as advanced inter-prediction modes. However, these come at the cost of high computational complexity of encoder, which may limit the benefits of the standard in practical applications. This paper shows that not all sequences benefit from using all such modes, which indicates that a number of encoder optimisations can be introduced to speed up AV1 encoding. A method based on decision trees is proposed to selectively decide whether to test all inter modes. Appropriate features are extracted and used to perform the decision for each block. Experimental results show that the proposed method can reduce the encoding time on average by 43.4% with limited impact on the coding efficiency. Jieon Kim, Saverio G. Blasi, André Seixas Dias, Marta Mrak, Ebroul Izquierdo |
ICASSP | 4 |
| 2019 | Decision Trees for Complexity Reduction in Video CompressionabstractThis paper proposes a method for complexity reduction in practical video encoders using multiple decision tree classifiers. The method is demonstrated for the fast implementation of the `High Efficiency Video Coding' (HEVC) standard, chosen because of its high bit rate reduction capability but large complexity overhead. Optimal partitioning of each video frame into coding units (CUs) is the main source of complexity as a vast number of combinations are tested. The decision tree models were trained to identify when the CU testing process, a time-consuming Lagrangian optimisation, can be skipped i.e a high probability that the CU can remain whole. A novel approach to finding the simplest and most effective decision tree model called `manual pruning' is described. Implementing the skip criteria reduced the average encoding time by 42.1% for a Bjøntegaard Delta rate detriment of 0.7%, for 17 standard test sequences in a range of resolutions and quantisation parameters. Natasha Westland, André Seixas Dias, Marta Mrak |
ICIP | 3 |
| 2019 | A Benchmark of Visual Storytelling in Social MediaabstractMedia editors in the newsroom are constantly pressed to provide a"like-being there" coverage of live events. Social media provides a disorganised collection of images and videos that media professionals need to grasp before publishing their latest news updated. Automated news visual storyline editing with social media content can be very challenging, as it not only entails the task of finding the right content but also making sure that news content evolves coherently over time. To tackle these issues, this paper proposes a benchmark for assessing social media visual storylines. The SocialStories benchmark, comprised by total of 40 curated stories covering sports and cultural events, provides the experimental setup and introduces novel quantitative metrics to perform a rigorous evaluation of visual storytelling with social media data. Gonçalo Marcelino, David Semedo, André Mourão, Saverio G. Blasi, Marta Mrak, João Magalhães |
ICMR | 5 |
| 2019 | End-to-End Conditional GAN-based Architectures for Image ColourisationabstractIn this work recent advances in conditional adversarial networks are investigated to develop an end-to-end architecture based on Convolutional Neural Networks (CNNs) to directly map realistic colours to an input greyscale image. Observing that existing colourisation methods sometimes exhibit a lack of colourfulness, this paper proposes a method to improve colourisation results. In particular, the method uses Generative Adversarial Neural Networks (GANs) and focuses on improvement of training stability to enable better generalisation in large multi-class image datasets. Additionally, the integration of instance and batch normalisation layers in both generator and discriminator is introduced to the popular U-Net architecture, boosting the network capabilities to generalise the style changes of the content. The method has been tested using the ILSVRC 2012 dataset, achieving improved automatic colourisation results compared to other methods based on GANs. Marc Gorriz, Marta Mrak, Alan F. Smeaton, Noel E. O'Connor |
MMSP | 2 |
| 2019 | Time-Constrained Video Delivery Using Adaptive Coding ParametersabstractIn some applications, video content needs to be encoded and uploaded to a remote destination within a pre-defined amount of time. In order to guarantee that the overall processing time does not exceed certain time constraints, a system that performs joint encoding and uploading time control is needed. Such a system requires the flexibility to control both the time spent in the encoding process and the bit-rate. This is because the latter significantly influences the transmission time, especially for transmissions under low-bandwidth constraints. This paper proposes a novel approach to address this challenge by adapting the quantization parameters (QP) of a video encoder in order to meet overall processing time requirements, including both encoding and uploading time. The proposed QP adaptation approach relies on mechanisms to accurately predict the encoding time and bit-rate during the encoding process for the incoming group of pictures. This in turn allows adequate QP selections that result in accurately meeting the overall time constraints. A comprehensive experimental evaluation shows that the proposed QP adaptation approach can accurately meet the overall time constraints by efficiently adapting the encoding process to different target times and bandwidth conditions. André Seixas Dias, Shenglan Huang, Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Complexity-Constrained Video Encoding and Delivery using Configuration Transfer MatrixabstractMany applications require video content to be encoded and uploaded under specific complexity constraints. While many speed-ups are available in practical video encoder implementations, it is difficult to predict the impact of such techniques on the actual content being encoded and therefore select the best configuration to meet the given constraints. A method is proposed in this paper to automatically select the encoder configuration in order to meet complexity constraints in terms of encoding and uploading time, using a pre-trained encoder configuration transfer matrix. The algorithm ensures that the content is processed within the specified targets, as presented in the experimental evaluation, where it is shown that the encoder can accurately meet specific constraints under a variety of conditions. Saverio G. Blasi, André Seixas Dias, Marta Mrak, Shenglan Huang, Ebroul Izquierdo |
PCS | 3 |
| 2018 | Estimation of Rate Control Parameters for Video Coding Using CNNabstractRate-control is essential to ensure efficient video delivery. Typical rate-control algorithms rely on bit allocation strategies, to appropriately distribute bits among frames. As reference frames are essential for exploiting temporal redundancies, intra frames are usually assigned a larger portion of the available bits. In this paper, an accurate method to estimate number of bits and quality of intra frames is proposed, which can be used for bit allocation in a rate-control scheme. The algorithm is based on deep learning, where networks are trained using the original frames as inputs, while distortions and sizes of compressed frames after encoding are used as ground truths. Two approaches are proposed where either local or global distortions are predicted. María Santamaría 0001, Ebroul Izquierdo, Saverio G. Blasi, Marta Mrak |
VCIP | 4 |
| 2017 | Performance evaluation of reverse tone mapping operators for dynamic range expansion of SDR video contentabstractWhen displaying Standard Dynamic Range (SDR) video on High Dynamic Range (HDR) displays a reverse tone mapping operation can be employed to expand the dynamic range of the SDR video to that offered by the display. This paper presents a subjective performance evaluation of existing reverse Tone Mapping Operators (rTMOs). The presented study evaluates the performance of rTMOs acting on both well exposed SDR video and video that exhibits exposure variation, as is often the case with user generated content (e.g. captured on mobile phones). The paper highlights flickering artefacts that arise in such cases and evaluates the effect on perceived quality of an adapted de-flickering method. Hu Hao, Yang Zhang 0003, Dimitris Agrafiotis, Matteo Naccari, Marta Mrak |
MMSP | 5 |
| 2017 | Complexity-driven rate-control for parallel HEVC codingabstractWhen using adaptive streaming, the content needs to be segmented so that clients can seamlessly switch to different rates depending on network conditions. On the video server each segment is stored in various bitrate representations, which are in practice provided by very fast encoders. Such encoders rely on parallelization strategies to limit the encoder complexity. Parallelization strongly affects the performance of rate-control (RC) algorithms, since different segments and parts of segments are encoded independently from each other. A new approach is proposed in this paper to tackle these issues, based on the optimization of the initial parameters of a state-of-the-art RC model for inter-predicted frames in an HEVC/H.265 codec. The model makes use of an estimate of the texture complexity of the first frame in the segment to efficiently tune the parameters depending on the target rate. The approach is consistently improving the accuracy of RC schemes as well as the visual quality, with negligible impact on the encoding efficiency. Juliette Carter, Saverio G. Blasi, Marta Mrak |
VCIP | 3 |
| 2016 | Optimised selection of structure of pictures for video codingabstractEncoders based on the High Efficiency Video Coding (HEVC) standard consider an input sequence as a succession of slices grouped in Structures of Pictures (SOP). The SOP used while encoding specifies many parameters, such as the coding order of frames, or the reference frames used during inter-prediction. Reference encoders typically make use of a fixed SOP structure of a given size, which is periodically repeated throughout the whole sequence. In this paper, the usage of unconventional SOP structures is first analysed, showing that most sequences benefit from usage of larger SOPs, and that the selection of the optimal SOP is highly content dependent. As a result, an algorithm is proposed to automatically select the optimal SOP size based on a low-complexity texture analysis of neighbouring frames. The algorithm is capable of adaptively changing SOP size during the encoding. Extensive evaluation shows that consistent bit-rate reductions are reported at the same objective quality as an effect of using the proposed algorithm. Vigneswaran Poobalasingam, Ebroul Izquierdo, Saverio G. Blasi, Marta Mrak |
MMSP | 4 |
| 2016 | Improved combined intra prediction for higher video compression efficiencyabstractIntra prediction is an important component of video compression. The High Efficiency Video Coding (HEVC) standard supports advanced Intra prediction tools to achieve remarkable Intra coding performance. However, higher compression efficiency performance may be achieved using Combined Intra Prediction (CIP). CIP consists in combining conventional Intra reference samples, with samples extracted from within the current block being encoded. In this paper, an improvement to the original CIP approach is proposed to achieve higher compression efficiency. Due to the fact the encoder does not have access to reconstruction samples while compressing a block, a controlled drift occurs at the block level between the Intra prediction blocks generated at the encoder and decoder when using CIP. The proposed approach aims at reducing such a drift, hence improving CIP performance. This Improved CIP is able to achieve up to 1.4 % BD-rate savings with respect to the HEVC reference software. The paper also proposes the combination of Improved CIP with an additional tool for enhancing the accuracy of Intra prediction. Up to 1.9 % BD-rate savings can be achieved using this approach. André Seixas Dias, Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
PCS | 3 |
| 2016 | An application of unified reference picture list for motion-compensated video compressionabstractModern video compression standards rely on motion-compensated prediction from multiple reference pictures. In past standardisation efforts following the introduction of bidirectional prediction, reference pictures were typically signalled in two independent lists, used for forward- and backward-prediction. Modern video coding standards, such as H.264/MPEG-4 Advanced Video Coding or High Efficiency Video Coding, extend this concept by removing the separation between preceding and succeeding reference pictures, and allow reference frames to be included in the two lists regardless of their temporal index. Therefore, the principle of using two separate lists can be considered obsolete. This paper evaluates an alternative concept, using a single, unified reference picture list. This alternative list, called “list unified” or LU, reduces reference picture signalling, while all functionalities for motion-compensated prediction using multiple references are preserved. Evaluation shows competitive video coding efficiency compared to the usage of two lists, with the advantage of simplified bitstream parsing and much improved reference picture flexibility. Sebastian Schwarz, Marta Mrak |
PCS | 2 |
| 2016 | Two-pass rate control for UHDTV delivery with HEVCabstractRate control has been regarded as an indispensable video coding tool for virtually any application involving video transmission. With the advent of many flexible tools introduced in the current state-of-the-art High Efficiency Video Coding (HEVC) standard, previous Rate-Distortion (RD) models used for rate control become insufficiently accurate. To overcome this issue, a new RD model has been recently proposed based on a robust correspondence between the rate and Lagrange multiplier λ. However, existing methods based on this model tend to perform sub-optimally after the scene change. In this paper, a two-pass rate control method is proposed, targeting Ultra High Definition Television (UHDTV) applications. In the first pass, a fast encoder with a reduced set of coding tools is used to obtain the data used for rate allocation and model parameter initialisation utilised during the second pass. Multiple encoding steps required to derive this information are avoided with the proposed variable quantization parameter framework. Experimental evaluation showed that the proposed two-pass rate control method achieves on average 4.4% BD-rate loss, compared with variable bit-rate encoding. That significantly outperforms the state-of-the-art HEVC rate control method with an average BD-rate loss of 8.8%. Ivan Zupancic, Ebroul Izquierdo, Matteo Naccari, Marta Mrak |
PCS | 4 |
| 2016 | Video Quality Evaluation Methodology and Verification Testing of HEVC Compression PerformanceabstractThe High Efficiency Video Coding (HEVC) standard (ITU-T H.265 and ISO/IEC 23008-2) has been developed with the main goal of providing significantly improved video compression compared with its predecessors. In order to evaluate this goal, verification tests were conducted by the Joint Collaborative Team on Video Coding of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29. This paper presents the subjective and objective results of a verification test in which the performance of the new standard is compared with its highly successful predecessor, the Advanced Video Coding (AVC) video compression standard (ITU-T H.264 and ISO/IEC 14496-10). The test used video sequences with resolutions ranging from 480p up to ultra-high definition, encoded at various quality levels using the HEVC Main profile and the AVC High profile. In order to provide a clear evaluation, this paper also discusses various aspects for the analysis of the test results. The tests showed that bit rate savings of 59% on average can be achieved by HEVC for the same perceived video quality, which is higher than a bit rate saving of 44% demonstrated with the PSNR objective quality metric. However, it has been shown that the bit rates required to achieve good quality of compressed content, as well as the bit rate savings relative to AVC, are highly dependent on the characteristics of the tested content. Thiow Keng Tan, Rajitha Weerakkody, Marta Mrak, Naeem Ramzan, Vittorio Baroncini, Jens-Rainer Ohm, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | High Dynamic Range Video Compression Exploiting Luminance MaskingabstractThe human visual system (HVS) exhibits nonlinear sensitivity to the distortions introduced by lossy image and video coding. This effect is due to the luminance masking, contrast masking, and spatial and temporal frequency masking characteristics of the HVS. This paper proposes a novel perception-based quantization to remove nonvisible information in high dynamic range (HDR) color pixels by exploiting luminance masking so that the performance of the High Efficiency Video Coding (HEVC) standard is improved for HDR content. A profile scaling based on a tone-mapping curve computed for each HDR frame is introduced. The quantization step is then perceptually tuned on a transform unit basis. The proposed method has been integrated into the HEVC reference model for the HEVC range extensions (HM-RExt), and its performance was assessed by measuring the bitrate reduction against the HM-RExt. The results indicate that the proposed method achieves significant bitrate savings, up to 42.2%, with an average of 12.8%, compared with HEVC at the same quality (based on HDR-visible difference predictor-2 and subjective evaluations). Yang Zhang 0003, Matteo Naccari, Dimitris Agrafiotis, Marta Mrak, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Adaptive quantisation in HEVC for contouring artefacts removal in UHD contentabstractContouring artefacts affect the visual experience of some particular types of compressed Ultra High Definition (UHD) sequences characterised by smoothly textured areas and gradual transitions in the value of the pixels. This paper proposes a technique to adjust the quantisation process at the encoder so that contouring artefacts are avoided. The devised method does not require any change at the decoder side and introduces a negligible coding rate increment (up to 3.4% for the same objective quality). This result compares favourably with the average 11.2% bit-rate penalty introduced by a method where the quantisation step is reduced in contour-prone areas. Nicolo Casali, Matteo Naccari, Marta Mrak, Riccardo Leonardi |
ICIP | 3 |
| 2015 | Region-adaptive quantisation for prevention of contouring in coded videoabstractFor some compressed UHD sequences, contouring artefacts appear in areas where smooth light variations occur, affecting the subjective quality of the decoded video sequences. This type of artefact appears even when light compression is applied. In this paper, an analysis on using fine quantisation in areas prone to contouring artefacts is presented, targeting a reduction in the visibility of contouring artefacts. A technique to further reduce the visibility of contouring artefacts by modifying the rate-distortion costs considered during the video encoding process is also proposed. Contouring artefacts can be prevented or significantly reduced with these techniques with a low impact in terms of rate-distortion performance for high target quality needed for UHD broadcasting. André Seixas Dias, Marta Mrak |
MMSP | 2 |
| 2015 | HEVC coding optimisation for Ultra High Definition television servicesabstractUltra High Definition TV (UHDTV) services are being trialled while UHD streaming services have already seen commercial débuts. The amount of data associated with these new services is very high thus extremely efficient video compression tools are required for delivery to the end user. The recently published High Efficiency Video Coding (HEVC) standard promises a new level of compression efficiency, up to 50% better than its predecessor, Advanced Video Coding (AVC). The greater efficiency in HEVC is obtained at much greater computational cost compared to AVC. A practical encoder must optimise the choice of coding tools and devise strategies to reduce the complexity without affecting the compression efficiency. This paper describes the results of a study aimed at optimising HEVC encoding for UHDTV content. The study first reviews the available HEVC coding tools to identify the best configuration before developing three new algorithms to further reduce the computational cost. The proposed optimisations can provide an additional 11.5% encoder speed-up for an average 3.1% bitrate increase on top of the best encoder configuration. Matteo Naccari, Andrea Gabriellini, Marta Mrak, Saverio G. Blasi, Ivan Zupancic, Ebroul Izquierdo |
PCS | 3 |
| 2015 | Frequency-Domain Intra Prediction Analysis and Processing for High-Quality Video CodingabstractMost of the advances in video coding technology focus on applications that require low bitrates, for example, for content distribution on a mass scale. For these applications, the performance of conventional coding methods is typically sufficient. Such schemes inevitably introduce large losses to the signal, which are unacceptable for numerous other professional applications such as capture, production, and archiving. To boost the performance of video codecs for high-quality content, better techniques are needed especially in the context of the prediction module. An analysis of conventional intra prediction methods used in the state-of-the-art High Efficiency Video Coding (HEVC) standard is reported in this paper, in terms of the prediction performance of such methods in the frequency domain. Appropriately modified encoder and decoder schemes are presented and used for this paper. The analysis shows that conventional intra prediction methods can be improved, especially for high frequency components of the signal which are typically difficult to predict. A novel approach to improve the efficiency of high-quality video coding is also presented in this paper based on such analysis. The modified encoder scheme allows for an additional stage of processing performed on the transformed prediction to replace selected frequency components of the signal with specifically defined synthetic content. The content is introduced in the signal using feature-dependent lookup tables. The approach is shown to achieve consistent gains against conventional HEVC with up to -5.2% coding gains in terms of bitrate savings. Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Masking of transformed intra-predicted blocks for high quality image and video codingabstractMany professional applications for image and video coding require very high levels of quality of the decoded signal. Under these conditions, even state-of-the-art standards such as High Efficiency Video Coding (HEVC) may not provide sufficient compression efficiency. Large values of the residual samples, especially at high frequency components, are difficult to encode and result in high bitrates of the coded signal. A novel scheme for image and video coding is presented in this paper to enhance compression at high quality levels, based on separate transform to the frequency domain of original and prediction signals. Selected frequency components of the prediction signal are discarded by means of appropriate masking patterns prior to the residual computation. The optimal masking pattern is selected for each transformed block and signalled in the bitstream. The approach is shown achieving gains against conventional HEVC under high quality constraints when coding both still images and video sequences. Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
ICIP | 2 |
| 2014 | Adaptive low complexity colour transform for video codingabstractFor video compression, the RGB signals are usually converted at the source to a perceptual colour space, followed by chroma sub-sampling, for coding efficiency. This is based on the typically higher human visual system sensitivity to the luminance than chrominance of image signals. However, there are specific applications that demand carrying the full RGB signals through the transmission chain, which may also benefit from lossless colour transforms, for efficient coding. In either case, the best colour transform function is noted to be content dependent, although fixed transforms are typically adopted for convenience. This paper presents a method of dynamically adapting this colour transform function for each picture block, using a class of low complexity lifting based schemes. The performance of the proposed algorithm is compared with a number of fixed colour transform schemes and shows a significant compression gain over native RGB coding and YCoCgtransform. Rajitha Weerakkody, Marta Mrak |
MMSP | 2 |
| 2014 | Special issue on advances in high dynamic range video research
Marta Mrak, Robert Cohen, Patrick Le Callet |
Signal Process. Image Commun. | 1 |
| 2013 | Visual masking phenomena with high dynamic range contentabstractHigh Dynamic Range (HDR) technology (capture and display), can offer high levels of immersion through a dynamic range that meets and exceeds that of the Human Visual System (HVS). This increase in immersion comes at the cost of higher bitrate requirements, which necessitate the development of efficient HDR-relevant coding solutions. Efficient perception-based compression of HDR imagery requires models that capture accurately the various masking effects experienced by the HVS under HDR conditions, so that bits are not wasted coding redundant imperceptible information. In this paper we present two psychovisual experiments that we carried out with the aid of a high dynamic range display, in order to determine potential differences between Standard Dynamic Range (SDR) and HDR edge masking (EM) and luminance masking (LM) effects. The EM experimental results indicate that the visibility threshold is higher for the case of HDR content than SDR, especially on the dark background side of an edge. The LM experimental results suggest that the HDR visibility threshold is higher compared to SDR for both dark and bright luminance backgrounds. Yang Zhang 0003, Dimitris Agrafiotis, Matteo Naccari, Marta Mrak, David Bull 0001 |
ICIP | 4 |
| 2013 | Intensity dependent spatial quantization with application in HEVCabstractWith the increase of video resolutions used in multimedia applications, solutions for improved compression are sought. The next generation video compression standard, High Efficiency Video Coding (HEVC) is being developed by ITU-T and ISO/IEC MPEG with the goal to provide significant improvements over H.264/AVC. To further increase compression capabilities of HEVC, this paper proposes an additional bitrate saving mechanism. In this context, an Intensity Dependant Spatial Quantization (IDSQ) perceptual tool is proposed which exploits the intensity masking of the human visual system and perceptually adjusts quantization. The proposed IDSQ allows for adaptation to the video characteristics and its design meets low complexity implementation requirements. The proposed IDSQ has been tested in the HEVC test model and its performance is reported under “just noticeable distortion” or “perceptually lossless” conditions. In this setup, bitrate reductions of up to 25% are reported with an average of 3.4% over an exhaustive test setup. Matteo Naccari, Marta Mrak |
ICME | 2 |
| 2013 | Binary alpha channel compression for coding of supplementary video streamsabstractWhen two video sequences are combined in each frame, an alpha channel signal is commonly used. Alpha channels are usually required in professional and studio applications for video editing and post-production processing. However, in other applications related to video broadcasting, alpha channels may be required also at the receiver side. One example is the insertion of sign language interpreters in broadcasted videos to help people with hearing difficulties to follow the programs. The interpreter may be distracting for normal hearing audiences so a solution where the video of the interpreter is switchable would be preferable. In this case, the frame composition, i.e. main program plus interpreter, is done at receiver and therefore the alpha channel needs also to be transmitted. In such an application scenario, it is important to devise efficient alpha channel coding algorithms which require low computational complexity. This paper proposes a lossless encoding algorithm which compresses homogeneous regions with few bits and involving a low amount of computational complexity for alpha channel coding. When compared the HEVC codec in lossless modality the proposed algorithm provides higher compression ratios (up to 54% improvement) and reduced computational complexity. Moreover, the proposed algorithm also outperforms conventional lossless coding techniques such as the Lempel-Ziv-Welch algorithm. Matteo Naccari, Marta Mrak |
MMSP | 2 |
| 2013 | Mirroring of coefficients for Transform Skipping in Video CodingabstractThis paper studies the Transform Skip option of the emerging High Efficiency Video Coding (HEVC) standard, and presents a modification to the ordering of coefficients within a transform skipped residual block. Depending on the scanning order for the given block, the coefficients are rearranged relative to their original spatial positions. The rearrangement is achieved by mirroring the coefficients across one of the block diagonals. When the diagonal scan is used, mirroring across the anti-diagonal is used, otherwise mirroring across the main diagonal (transpose) is used. In this way the distribution of significant coefficients is closer to the distribution in a case when transform is applied. Therefore the entropy coding may benefit from the new arrangement. The experiments show BD-rate gains of up to 2.4% for Screen Content sequences. Rajitha Weerakkody, Marta Mrak |
MMSP | 2 |
| 2013 | Improving inter prediction in HEVC with residual DPCM for lossless screen content codingabstractVideo content containing computer generated objects is usually denoted as screen content and is becoming popular in applications such as desktop sharing, wireless displays, etc. Screen content images and videos are characterized by high frequency details such as sharp edges and high contrast image areas. On these areas classical lossy encoding tools - spatial transform plus quantization - may significantly compromise their quality and intelligibility. Therefore, lossless coding is used instead and improved coding tools should be specifically devised for screen content. In this context this paper proposes a residual differential pulse code modulation (RDPCM) applied to inter predicted residuals and tested in the context of the HEVC range extension development. The proposed method exploits the spatial correlation present in blocks containing edges or text areas which are poorly predicted by motion compensation. In addition to the baseline inter RDCPM, two improvements to the compression efficiency and the overall throughput are presented and assessed. When compared to HEVC lossless coding as specified in Version 1 of the standard, the proposed algorithm achieves up to 8% average bitrate reduction while not increasing the overall decoding complexity. Matteo Naccari, Saverio G. Blasi, Marta Mrak, Ebroul Izquierdo |
PCS | 3 |
| 2013 | High dynamic range video compression by intensity dependent spatial quantization in HEVCabstractThe Human Visual System (HVS) shows non-linear sensitivity to the distortion introduced by lossy image and video coding. This non-linear sensitivity is due to luminance masking, contrast masking and the spatial and temporal frequency masking phenomena of the HVS. This paper proposes a perception-based quantization method that exploits luminance masking in the HVS in order to enhance the performance of the High Efficiency Video Coding (HEVC) standard for the case of High Dynamic Range (HDR) video content. A profile scaling based on a tone-mapping curve computed for each HDR frame is introduced. The quantization step is then perceptually tuned on a Transform Unit (TU) basis. The proposed method has been integrated into the reference codec considered for the HEVC range extensions and its performance was assessed by measuring the bitrate reduction against the codec without perceptual quantization. The HDR-VDP-2 image quality metric was employed to measure the compressed picture quality. For the same quality level, an average bitrate reduction of 9% is achieved across all tested HDR sequences. Yang Zhang 0003, Matteo Naccari, Dimitris Agrafiotis, Marta Mrak, David Bull 0001 |
PCS | 4 |
| 2013 | Adaptive transform skipping for improved coding of motion compensated residuals
Andrea Gabriellini, Matteo Naccari, Marta Mrak, David Flynn, Glenn Van Wallendael |
Signal Process. Image Commun. | 3 |
| 2012 | Spatial transform skip in the emerging High Efficiency Video Coding standardabstractTo meet the still growing compression efficiency needs for high definition content, ITU and MPEG are now defining the High Efficiency Video Coding (HEVC) standard. The HEVC codec still relies on a hybrid motion compensated predictive video coding architecture as its H.264/AVC ancestor though novel coding tools are introduced. These novel coding tools provide a highly uncorrelated prediction residue for which classical frequency decomposition methods as the discrete cosine transform may not provide an effective energy compaction. Therefore this paper proposes a transform skip mode which allows skipping one or both directions where the transform is applied. The proposed transform skip mode is integrated in the HEVC codec and is able to provide bitrate reductions of up to 5% at the same objective quality when compared with the HEVC reference codec. Andrea Gabriellini, Matteo Naccari, Marta Mrak, David Flynn |
ICIP | 3 |
| 2012 | Multi-loop quality scalability based on high efficiency video codingabstractScalable video coding performance largely depends on the underlying single layer coding efficiency. In this paper, the quality scalability capabilities are evaluated on a base of the new High Efficiency Video Coding (HEVC) standard under development. To enable the evaluation, a multi-loop codec has been designed using HEVC. Adaptive inter-layer prediction is realized by including the lower layer in the reference list of the enhancement layer. As a result, adaptive scalability on frame level and on prediction unit level is accomplished. Compared to single layer coding, 19.4% Bjontegaard Delta bitrate increase is measured over approximately a 30dB to 40dB PSNR range. When compared to simulcast, 20.6% bitrate reduction can be achieved. Under equivalent conditions, the presented technique achieves 43.8% bitrate reduction over Coarse Grain Scalability of the SVC - H.264/AVC-based standard. Glenn Van Wallendael, Jan De Cock, Rik Van de Walle, Marta Mrak |
PCS | 4 |
| 2012 | Scalable Comic-Like Video Summaries and Layout DisturbanceabstractThis paper describes an efficient system for scalable video summarization that exploits comic-like summaries and multi-scale representations to facilitate interactivity and balance between content coverage and compactness. Due to the layout disturbance induced by the transitions between scales, a new heuristic algorithm is proposed to restrict changes to bounded summary segments. Conducted user evaluations show that the proposed methodology improves usability while keeping the summaries compact and informative. Luis Herranz, Janko Calic, José María Martínez Sanchez, Marta Mrak |
IEEE Trans. Multim. | 4 |
| 2011 | Parallel processing for combined intra prediction in high efficiency video codingabstractAdvanced intra prediction is one of the key components in highly efficient video coding since it reduces the bit-rate of the most costly intra frames. Combined Intra Prediction (CIP) is a novel tool that takes into account well established directional prediction from neighboring blocks, as well as local mean prediction from current block. In this way, it reduces bit-rate by further exploiting spatial redundancy. This paper proposes an implementation that will overcome its potential limitation limited parallelization capabilities. In order to make it more computationally parallelizable, adaptive open-loop prediction templates have been designed. Experiments have been performed which show that the proposed design preserves coding gains introduced by CIP, while providing a solution for more parallelizable, and therefore faster, implementations. Marta Mrak, Andrea Gabriellini, David Flynn, Thomas Davies 0002 |
ICIP | 1 |
| 2010 | Adaptive motion-estimation-mode selection for depth video codingabstractAn effective representation of 3D video in future 3D-TV systems consists of monoscopic video (colour component) and associated per-pixel depth information (depth component). As depth component indicates relative distance between objects within the scene and a camera, pixel values change not only when objects move in vertical and horizontal directions but also when they move in a depth direction. Instead of predicting motion of objects in two directions as appearing in traditional video codecs, three-dimensional block matching (3D-BM) achieves more accurate motion estimation in depth video coding. However overall performance of 3D-BM exceeds that of the traditional two-dimensional block matching (2D-BM) only at high bit rate. In this paper, an adaptive 2D-3D BM selection algorithm is introduced to compromise performance of 2DBM and 3D-BM. The Lagrangian optimisation algorithm is applied to select motion estimation mode at a block level. The experiment results reveal that the proposed adaptive motion-estimation-mode selection can improve the performance of 3D-BM at low bit rate while advantages of 3D-BM are preserved at high bit rate. Bunchat Kamolrat, Warnakulasuriya Anil Chandana Fernando, Marta Mrak |
ICASSP | 3 |
| 2010 | Special issue on advances in image and video processing techniques
Mislav Grgic, Marta Mrak, Kresimir Delac |
Multim. Tools Appl. | 2 |
| 2009 | Utilisation of edge adaptive upsampling in compression of depth map videos for enhanced free-viewpoint renderingabstractIn this paper we propose a novel video object edge adaptive upsampling scheme for application in video-plus-depth and Multi-View plus Depth (MVD) video coding chains with reduced resolution. Proposed scheme is for improving the rate-distortion performance of reduced-resolution depth map coders taking into account the rendering distortion induced in free-viewpoint videos. The inherent loss in fine details due to downsampling, particularly at video object boundaries causes significant visual artefacts in rendered free-viewpoint images. The proposed edge adaptive upsampling filter allows the conservation and better reconstruction of such critical object boundaries. Furthermore, the proposed scheme does not require the edge information to be communicated to the decoder, as the edge information used in the adaptive upsampling is derived from the reconstructed colour video. Test results show that as much as 1.2 dB gain in free-viewpoint video quality can be achieved with the utilization of the proposed method compared to the scheme that uses the linear MPEG re-sampling filter. The proposed approach is suitable for video-plus-depth as well as MVD applications, in which it is critical to satisfy bandwidth constraints while maintaining high free-viewpoint image quality. Erhan Ekmekcioglu, Marta Mrak, Stewart Worrall 0001, Ahmet M. Kondoz |
ICIP | 2 |
| 2009 | Depth based object prioritisation for 3D video communication over Wireless LANabstractThis paper presents a content prioritization technique for perceptually enhanced communication of 3D video data over wireless local area network. The prioritisation is performed on the basis of a novel segmentation technique, which utilizes the depth information of 3D content for separating individual objects of the sequence. These objects are then encoded and transmitted separately under different protection levels depending on their perceived importance for perceptual visual quality. The algorithm performance is demonstrated for object based MPEG-4 video transmission over home network and is compared with a basic prioritisation scheme based on splitting a sequence into even and odd frames and encoding them separately with equivalent prioritization as in the proposed scheme. The results demonstrate significant improvement in subjective quality with the proposed scheme. Sabih Nasir, Chaminda Hewage, Marta Mrak, Stewart Worrall 0001, Ahmet M. Kondoz |
ICIP | 3 |
| 2009 | Fast analysis of scalable video for adaptive browsing interfaces
Marta Mrak, Janko Calic, Ahmet M. Kondoz |
Comput. Vis. Image Underst. | 1 |
| 2009 | Special issue on scalable coded media beyond compression
G. Charith K. Abhayaratne, Ebroul Izquierdo, Marta Mrak, Stefano Tubaro |
Signal Process. Image Commun. | 3 |
| 2009 | Object tracking in surveillance videos using compressed domain features from scalable bit-streams
Khurram Mehmood, Marta Mrak, Janko Calic, Ahmet M. Kondoz |
Signal Process. Image Commun. | 2 |
| 2008 | Flexible generation of video summaries from layered video bit-streamsabstractThe work presented in this paper introduces a method for efficient adaptability of video summaries to spatial requirements defined by display size, user’s needs and channel limitations. By utilising compressed domain features and an efficient contour evolution algorithm, a scale space of temporal video descriptors is generated, enabling dynamic video summarisation in real-time. The summary is laid out utilising an unsupervised robust spectral clustering technique and a fast discrete optimisation algorithm. Results show excellent scalability of the video summarisation interface and highly improved efficiency of summary generation. Janko Calic, Marta Mrak, Ahmet M. Kondoz |
ICIP | 2 |
| 2008 | Multi bearer channel resource allocation for optimised transmission of video objectsabstractThis paper presents a novel channel optimisation scheme that enhances the quality of object based video, transmitted over a fixed bandwidth channel. The optimisation methodology is based on an accurate modelling of video packet distortion at the encoder. Video packets are ranked according to their expected distortion and are then mapped to one of a number of different priority radio bearers. In the proposed scheme the video compression technique uses motion compensated prediction and video frames are split into a number of video packets. The algorithm performance is demonstrated for object based MPEG-4 video transmission over a UMTS/FDD system. The results demonstrate that the performance gain achieved with the proposed scheme can reach 2 dB, compared with the equal error protection scheme for video transmission over a fixed bandwidth channel. Sabih Nasir, Stewart Worrall 0001, Marta Mrak, Ahmet M. Kondoz |
ICIP | 3 |
| 2008 | Flexible motion model with variable size blocks for depth frames coding in colour-depth based 3D video codingabstractNew techniques for representing 3D video are often realised using monoscopic video and associated per-pixel depth information. While the compression of both monoscopic and depth video channels can be achieved using traditional video coding techniques, in this paper, specific properties of the depth channel are exploited to further compress the depth information. More precisely, enhanced compression of the depth channel is achieved by using a highly flexible motion model, binary partition tree (BPT) which enables adaptive partitioning of the depth frames leading to better rate-distortion optimisation for inter-frame prediction. Comparing to a conventional approach, limited variable size block (LVSB) which enables only limited frame partitioning options, the presented approach brings a gain of up to 70% in bitrate reduction for the depth information. Bunchat Kamolrat, Warnakulasuriya Anil Chandana Fernando, Marta Mrak, Ahmet M. Kondoz |
ICME | 3 |
| 2008 | Hierarchical motion analysis for fast summarisation of scalable coded videoabstractDue to a high demand for efficient video summarisation and video adaptation technologies, this paper focuses on utilisation of compressed domain feature extraction and hierarchical analysis of motion information in scalable video in order to generate intuitive visual summaries. By combining the analysis of inherently hierarchical motion activity measure and a fast geometrical curve simplification algorithm, a set of the most representative key-frames is generated in a very fast and robust manner. The experimental results show good subjective representation, while the method efficiency enables fast generation of summaries of large-scale video repositories. Marta Mrak, Janko Calic, Giovanni Cordara, Ahmet M. Kondoz |
ICME | 1 |
| 2007 | Spatially Adaptive Wavelet Transform for Video Coding with Multi-Scale Motion CompensationabstractIn this paper a technique that enables efficient synthesis of the prediction signal for application in multi-scale motion compensation is presented. The technique targets prediction of high-pass spatial subbands for motion compensation at higher scales. Since in the targeted framework these subbands are obtained by high-pass filtering of prediction signal in pixel domain, an adaptive approach for filtering is proposed to support decomposition of differently predicted frame areas. In this way an efficient application of different prediction modes at all scales used for compensation is enabled. Experimental results show that for fast sequences where such an application of different prediction modes is crucial, the proposed adaptive transform introduces significant objective and visual improvements. Marta Mrak, Ebroul Izquierdo |
ICIP (2) | 1 |
| 2007 | Optimised Compression Strategy in Wavelet-Based Video Coding using Improved Context ModelsabstractAccurate probability estimation is a key to efficient compression in entropy coding phase of state-of-the-art video coding systems. Probability estimation can be enhanced if contexts in which symbols occur are used during the probability estimation phase. However, these contexts have to be carefully designed in order to avoid negative effects. Methods that use tree structures to model contexts of various syntax elements have been proven efficient in image and video coding. In this paper we use such structure to build optimised contexts for application in scalable wavelet-based video coding. With the proposed approach context are designed separately for intra-coded frames and motion-compensated frames considering varying statistics across different spatio-temporal subbands. Moreover, contexts are separately designed for different bit-planes. Comparison with compression using fixed contexts from embedded ZeroBlock coding (EZBC) has been performed showing improvements when context modelling on tree structures is applied. Toni Zgaljic, Marta Mrak, Ebroul Izquierdo |
ICIP (3) | 2 |
| 2007 | Perceptually adaptive joint deringing-deblocking filtering for scalable video transmission over wireless networks
Shuai Wan, Marta Mrak, Naeem Ramzan, Ebroul Izquierdo |
Signal Process. Image Commun. | 2 |
| 2006 | An Entropy Coding Scheme for Multi-Component Scalable Motion InformationabstractFully scalable video bit-stream requires layered structure of most of its components. For that reason few methods targeting scalability on motion information have been proposed over the last decade. However, layered representation requires new entropy coding strategies able to efficiently handling of redundancies between different layers. In this paper three methods for entropy coding of layered motion information are proposed. The influence of these schemes on the reconstructed video quality has been also studied Toni Zgaljic, Marta Mrak, Nikola Sprljan, Ebroul Izquierdo |
ICASSP (2) | 2 |
| 2006 | Evaluation of Techniques for Modeling of Layered Motion StructureabstractMotion information scalability is important for scalable bit-stream adaptation on low bit-rates, when motion rate occupies a significant portion of the total bit-rate. This type of scalability can be achieved by layered representation of motion block partitioning and predictive coding of associated motion vectors across these layers. So far, several approaches for creating layered motion structure targeting quality scalability have been proposed and in this paper their accuracy is evaluated. For that purpose optimal motion models have been found. It has been shown that simple evaluation of reconstruction error at the encoder side improves suboptimal modeling techniques. Marta Mrak, Nikola Sprljan, Ebroul Izquierdo |
ICIP | 1 |
| 2005 | A Resolution Adaptive Interpolation Technique for Enhanced Decoding of Scalable Coded VideoabstractSubpixel accurate motion compensated temporal filtering introduces a significant coding gain in scalable 3D wavelet video codecs. The influence of the chosen subpixel interpolation technique has not yet been fully analysed in the context of resolution scalability. That problem is addressed in this paper. It is shown that support for increased accuracy and resolution adaptive spatial interpolation needs to be featured in a scalable video decoder, when low resolution sequences are targeted. Using the proposed resolution adaptive filters based on sinc kernels leads to improved decoding performance at low resolution in the sense of achieving higher quality while reducing the complexity of the system. Marta Mrak, Nikola Sprljan, Ebroul Izquierdo |
ICASSP (2) | 1 |
| 2005 | A fast error protection scheme for transmission of embedded coded images over unreliable channels and fixed packet sizeabstractJoint source-channel coding enables efficient transmission of embedded bitstreams over unreliable channels. We address channels with fixed packetisation and decoding without or with minimal delays. The computation of an optimal protection scheme for such bitstreams is generally an exponential complexity problem and hence not applicable in a straightforward implementation. Using the rate-distortion characteristics of the source bitstream and the dynamic programming approach we construct an efficient unequal error protection scheme for the predefined channels. Our algorithm is of linear complexity and thus applicable in real time scenarios. Nikola Sprljan, Marta Mrak, Ebroul Izquierdo |
ICASSP (3) | 2 |
| 2003 | A context modeling algorithm and its application in video compressionabstractA new algorithm for context modeling of binary sources with application to video compression is presented. Our proposed method is based on a tree rearrangement and tree selection process for an optimized modeling of binary context trees. We demonstrate its use for adaptive context-based coding of selected syntax elements in a video coder. For that purpose we apply our proposed technique to the H.264/AVC standard and evaluate its performance for different sources and different quantization parameters. Experimental results show that by using our proposed algorithm coding gains similar or superior to those obtained with the H.264/AVC CABAC algorithm is achieved. Marta Mrak, Detlev Marpe, Thomas Wiegand 0001 |
ICIP (3) | 1 |