Miska M. Hannuksela

dblp:42/1215 · DBLP profile ↗
← Back
123ranked-venue papers
8as first author
27since 2021 · last 2025
0000-0003-3405-0850ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 111 · 7 first-author · 26 since 2021Systems, architecture and hardware · 5Artificial intelligence and machine learning · 3Computer networks · 2Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for Machines
abstract
The proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters.
Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
PCS6
2025 A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec
Nannan Zou, Antti Hallapuro, Francesco Cricri, Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Miska M. Hannuksela, Esa Rahtu
PCS7
2024 IN-Loop Filter for Object Mask Coding in Versatile Video Coding
abstract
This paper explores the challenges and solutions in compressing object mask video, which can provide additional scene context for machine learning applications. Object masks, identifying specific objects or regions in images, are crucial for precise visual analysis. However, compressing them with video codecs tuned for human consumption may lead to detrimental artefacts for machine tasks. To alleviate such artefacts, a modification to the luma mapping and chroma scaling (LMCS) in-loop filter in the Versatile Video Coding standard is proposed, targeting the reconstruction of object mask video. By performing object mask reconstruction within the encoding loop, the encoder has access to correctly reference pictures for prediction. Consequently, coding performance is significantly enhanced. Objective evaluation against post-state-of-the-art processing mask reconstruction confirms a 42.0dB and 66.1dB improvement in Y-PSNR for random access and all intra encoding, respectively, while maintaining coding complexity similar to unmodified encoding without post-processing.
Sebastian Schwarz, Miska M. Hannuksela, Döne Bugdayci Sansli
ICIP2
2024 Feasibility Study of Multi-Layer VVC Coding Scheme for Hybrid Machine-Human Consumption
abstract
The proliferation of machine vision applications necessitates developing more efficient visual data compression schemes for machine consumption. However, numerous automated use cases still require keeping humans in the loop, leading to the need for a machine-optimized video streaming with the option for human supervision. This paper investigates the feasibility of using the multi-layer coding approach of the emerging Versatile Video Coding (VVC) standard to create favorable conditions for hybrid machine-human consumption. We introduce a multi-layer coding scheme, where the base layer (BL) is optimized for machines and the enhancement layer (EL) complements the stream for human vision. Our results demonstrate that the bitrate of the proposed multi-layer stream (BL + EL) is, on average, 11% higher than that of a single-layer VVC. However, the more compact BL yields overall bandwidth savings as long as the EL is required less than 80% of the time.
Jaakko Laitinen, Tero Partanen, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
ICME5
2024 Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for Machines
abstract
Recent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively.
Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
MMSP5
2024 Packed Regions Information SEI Message
abstract
Specific regions of interest (ROIs) within a video are of greater interest than the remainder of the video for many use cases. The Packed Regions Information (PRI) Supplemental Enhancement Information (SEI) message enables packing of rectangular ROIs from an original picture into a smaller resolution picture for video coding, reducing pixel rate and bitrate. The SEI message signals metadata describing the size and position of the ROIs in the coded picture and in the original picture. Decoders may use the metadata to reconstruct target pictures at the original resolution from the decoded pictures containing the packed regions. The PRI SEI message is under consideration for potential inclusion in a future version of the Versatile Supplemental Enhanced Information (VSEI) standard. Experimental results are provided for use of the PRI SEI message with test conditions for machine analysis of coded video content, showing reductions in bitrate and pixel rate.
Jill M. Boyce, Miska M. Hannuksela, Honglei Zhang 0001, Antti Hallapuro
VCIP2
2023 NN-VVC: Versatile Video Coding boosted by self-supervisedly learned image coding for machines
abstract
The recent progress in artificial intelligence has led to an ever-increasing usage of images and videos by machine analysis algorithms, mainly neural networks. Nonetheless, compression, storage and transmission of media have traditionally been designed considering human beings as the viewers of the content. Recent research on image and video coding for machine analysis has progressed mainly in two almost orthogonal directions. The first is represented by end-to-end (E2E) learned codecs which, while offering high performance on image coding, are not yet on par with state-of-the-art conventional video codecs and lack interoperability. The second direction considers using the Versatile Video Coding (VVC) standard or any other conventional video codec (CVC) together with pre- and post-processing operations targeting machine analysis. While the CVC-based methods benefit from interoperability and broad hardware and software support, the machine task performance is often lower than the desired level, particularly in low bitrates. This paper proposes a hybrid codec for machines called NN-VVC, which combines the advantages of an E2E-learned image codec and a CVC to achieve high performance in both image and video coding for machines. Our experiments show that the proposed system achieved up to -43.20% and -26.8% Bjøntegaard Delta rate reduction over VVC for image and video data, respectively, when evaluated on multiple different datasets and machine vision tasks. To the best of our knowledge, this is the first research paper showing a hybrid video codec that outperforms VVC on multiple datasets and multiple machine vision tasks.
Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Antti Hallapuro, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu
ISM7
2023 Optimal Tile Size and Streaming Field of View for VR Streaming
abstract
Virtual reality (VR) video services require a high bitrate, and hence, viewport-adaptive streaming techniques like motion-constrained-tile-set (MCTS) have been found important to reduce streaming-rate and storage demands. The tiling scheme and streaming field of view (FOV) are among the key elements in designing an optimal VR viewport-adaptive streaming solution, in terms of rate-distortion (R-D) performance. The aim of this study is to propose an optimal configuration for the tile grid and streaming FOV, considering different VR viewing situations such as head motion speed, system delay, and head-mounted display FOV. To achieve this, a wide range of tiling schemes and streaming FOVs are examined to study the storage and streaming R-D performance of the MCTS-based technique in both viewport and non-viewport areas using a quality metric called Zonal-cubic PSNR. The findings demonstrate that for VR applications focused on preserving high viewport quality, fine tile grids lead to higher performance. In scenarios featuring small and large HMD FOV, the optimal configuration involves a small and medium streaming FOV, respectively.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
MMSP3
2023 Overfitting NN loop-filters in video coding
abstract
Overfitting is usually regarded as a negative condition since it impairs the generalisation power of a model. Nevertheless, overfitting a Neural Network (NN) on test data may be advantageous to improve the compression efficiency of image/video coding tools and systems. Previous research has demonstrated the benefits of NN overfitting for post-processing operations, i.e. post-filters, but not yet for actual decoding tools. Generally, the NN is overfitted on test data at the encoder end, and the weight update is coded and sent to the decoder end along the image/video bitstream. The proposed approach follows this strategy. In particular, the overfitting of the Low Operation Point (LOP) loop-filter in NN-based Video Coding (NNVC) software is studied. The overall approach yields Bjøntegaard Delta rate (BD-rate) of -7.74%, -13.73% and -12.49%, for the Y, U and V components, respectively. Out of these coding gains, 1.21%, 6.43% and 5.52%, for the Y, U and V components, are attributed to the overfitting. The boost in the coding gains comes with only 1.5% more complexity, due to the multiplier parameters introduced during the overfitting.
Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela, Tapio Elomaa
VCIP7
2022 Comparison of Boundary Artifact Removal Methods in Coding of Generalized Cubemap Projection Using VVC
abstract
Virtual reality applications use 360-degree videos and head mount displays with stereoscopic capabilities to provide full immersion experience. Among existing projection format, Cubemap projection provides and improved compression performance, but similar to other projects, it suffers from visual artifacts in the rendered viewport because of discontinuity of the content in different regions. In this work, we investigated the effect of different methods in Versatile Video Coding standard and other pre- and post-processing algorithms for removing boundary artifacts by introducing a new objective quality metric for systematic comparison. We further improve the current methods by aligning the GCMP’s face boundaries to coding unit boundaries. The observation is that the combination of these method with existing methods in VVC offers the best result.
Kianoush Jafari, Alireza Aminlou, Miska M. Hannuksela
ICASSP3
2022 Content-Adaptive Neural Network Post-Processing Filter with NNR-Coded Weight-Updates
abstract
Neural Network (NN) filters improve the perceptual quality of reconstructed videos by reducing compression artefacts. For content adaptation, a few NN-filters use over-fitting. As the adaptation signal is a weight-update, compression is required to minimise significant bitrate overheads. Most approaches, however, use generic data compression algorithms, which are inadequate for coding NN weight-updates. This work introduces a content-adaptive NN post-processing filter with weight-updates coded using the Neural Network compression and Representation (NNR) standard. The bitrate overhead is further decreased by over-fitting only a subset of weights, selected via energy-based analysis. The proposed filter saved about 4.57% (Y), 10.33% (Cb), 6.53% (Cr) Bjøntegaard Delta rate (BD-rate) on top of the Versatile Video Coding (VVC) Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0, in Random Access (RA) configuration. Compared to the non-over-fitted NN, the performance was doubled; and compared to 7z, NNR reduced the bitrate of the weight-update by ∼64%.
María Santamaría 0001, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela
ICIP6
2022 Bridging the Gap Between Image Coding for Machines and Humans
abstract
Image coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance.
Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu
ICIP7
2022 Stochastic Binary-Ternary Quantization for Communication Efficient Federated Computation
abstract
A stochastic binary-ternary (SBT) quantization approach is introduced for communication efficient federated computation; form of collaborative computing where locally trained models are exchanged between institutes. Communication of deep neural network models could be highly inefficient due to their large size. This motivates model compression in which quantization is an important step. Two well-known quantization algorithms are binary and ternary quantization. The first leads into good compression, sacrificing accuracy. The second provides good accuracy with less compression. To better benefit from trade-off between accuracy and compression, we propose an algorithm to stochastically switch between binary and ternary quantization. By combining with uniform quantization, we further extend the proposed algorithm to a hierarchical method which results in even better compression without sacrificing the accuracy. We tested the proposed algorithm using Neural network Compression Test Model (NCTM) provided by MPEG community. Our results demonstrate that the hierarchical variant of the proposed algorithm outperforms other quantization algorithms in term of compression, while maintaining the accuracy competitive to that provided by other methods.
Rangu Goutham, Homayun Afrabandpey, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Miska M. Hannuksela, Hamed Rezazadegan Tavakoli
ICIP6
2022 Adaptive Multi-Scale Progressive Probability Model for Lossless Image Compression
abstract
Domain adaptation is an efficient technique to improve the performance of a system by adapting a pre-trained model to the given input data. The adaptation technique has been generally applied in conventional video codecs. For neural network-based systems, the encoder may adapt the decoder to the input data by fine-tuning a pre-trained model present at the decoder side. The weight update is then transferred to the decoder and the updated model is used to decode the bitstream. However, due to the large number of parameters in deep neural networks, the overhead of the weight update may diminish the gain from the adaptation technique. In recent years, various methods have been proposed to reduce the overhead without significantly compromising the gain. In this paper, we propose an adaptive multi-scale progressive probability model for lossless image compression. The proposed method uses the data that has already been processed at the inference stage to fine-tune the probability model. Importantly, the decoder can apply the fine-tuning by itself resulting a small adaptation overhead to help the decoder in performing the fine-tuning. The proposed method achieves up to 0.28 bits-per-pixel (BPP) reduction on four benchmark datasets compared to the state-of-the-art method.
Honglei Zhang 0001, Francesco Cricri, Nannan Zou, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela
ICIP5
2022 Optimizing storage and delivery of Omnidirectional Videos in Viewport-dependent streaming
abstract
The OMAF standard makes use of a framework called the viewport-dependent-delivery for the streaming of 360-degree videos. OMAF uses ISOBMFF for storage and MPEG-DASH as one of the delivery mechanisms. In viewport-dependent-streaming videos are spatially divided and encoded into multiple tracks and each track is further segmented for DASH delivery. Segmentation requires additional metadata which adds to bitrate overhead. The main contributor to this overhead is the track fragment run in a box with the four-character code, ‘trun’. The TRUN records the following information of each sample in a track: the size, duration, flags, and time offsets and uses a fixed byte size to record this information. To minimize the bitrate overhead of TRUN, four different representation algorithms have been explored. This paper briefly describes the four TRUN representations and discusses the benefits and drawbacks of each algorithm. For evaluation, the algorithms were implemented in the MP4BOX module of the GPAC suite. The results were evaluated for different segment durations (500ms, 1s, 2s, 4s), different tiling grids (8x4, 9x6), two videos (bip-bop, countertiles) with different packaging techniques (no encryption, encryption of Keyframes, encryption of all frames) The algorithms reduced the bitrate overhead by 59% on average as compared to the original TRUN representation.
Kashyap Kammachi Sreedhar, Miska M. Hannuksela, Emre Aksu, Lauri Ilola, Lukasz Condrad
ISM2
2022 Low-precision post-filtering in video coding
abstract
Neural Networks (NNs) have demonstrated their effectiveness in tackling challenges involving multimedia content. In the video coding field, NNs are actively exploited as novel tools that complement conventional signal processing tools, as well as end-to-end coding solutions. Since NNs use commonly floating-point arithmetic, different results may be generated in different computing environments, leading to discrepancies and even corrupted reconstructions. Accordingly, this issue is solved by employing fixed-point arithmetic instead. This paper studies the quantisation of a 32-bit floating-point (float32) NN post-filter to 32-bit fixed-point (int32) and 16-bit fixed-point (int16). On top of the VVC Test Model (VTM) 11.0 NN-based Video Coding (NNVC) 1.0, the coding gains of the float32 post-filter are 5.01% (Y), 18.95% (Cb) and 17.33% Cr. Compared to the float32 inference, the quantised models produce coding losses: 0.01% (Y, Cb and Cr) for the int32 inference and 0.45% (Y), 1.97% (Cb) and 1.19% (Cr) for the int16 inference. Nevertheless, the fixed-point approaches achieve bit exact matches in different computing environments. Moreover, the decoding time with int16 is about half the decoding time of int32.
Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela
ISM7
2022 The Lottery Ticket Adaptation for Neural Video Coding
abstract
Recently, learning based video compression methods have attracted increasing attention. However, most learning based video codecs are not adaptive to different video contents. Though adaptation at inference time is a solution to tackle this issue, adapting all the codec’s parameters is computationally expensive and brings heavy bitrate overhead. The recently proposed Lottery Ticket Hypothesis (LTH) states that an over-parameterized neural network contains smaller subnetworks (winning tickets) that can match the performance of the original network. In this paper, we present a novel lottery-ticket adaptation technique on decoder-side multiplicative parameters of a neural network, transferring the concept of winning lottery tickets to video compression tasks. At inference time, the winning multiplicative parameters are overfitted, compressed, and signaled together with encoded frames for decoding. We show that our approach outperforms the Versatile Video Coding (VVC) standard in the Multiscale Structural Similarity (MS-SSIM) at a low bitrate on both the UVG and JVET sequences. To the best of our knowledge, this is the first attempt to apply LTH in the video compression domain. Also, this is the first published end-to-end learned video codec working directly on YUV format, which outperforms VVC on UVG and JVET datasets in MS-SSIM.
Nannan Zou, Francesco Cricri, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu
ISM5
2022 Overview of the Neural Network Compression and Representation (NNR) Standard
abstract
Neural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR.
Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer
IEEE Trans. Circuits Syst. Video Technol.8
2021 Learned Enhancement Filters for Image Coding for Machines
abstract
Machine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively.
Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu
ISM7
2021 Content-adaptive convolutional neural network post-processing filter
abstract
Neural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration.
María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj
ISM7
2021 Enhancing Image Coding for Machines with Compressed Feature Residuals
abstract
As computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP.
Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu
ISM8
2021 Adaptation and Attention for Neural Video Coding
abstract
Neural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings.
Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu
ISM8
2021 Coding of volumetric content with MIV using VVC subpictures
abstract
Storage and transport of six degrees of freedom (6DoF) dynamic volumetric visual content for immersive applications requires efficient compression. ISO/IEC MPEG has recently been working on a standard that aims to efficiently code and deliver 6DoF immersive visual experiences. This standard is called the MIV. MIV uses regular 2D video codecs to code the visual data. MPEG jointly with ITU-T VCEG, has also specified the VVC standard. VVC introduced recently the concept of subpicture. This tool was specifically designed to provide independent accessibility and decodability of sub-bitstreams for omnidirectional applications. This paper shows the benefit of using subpictures in the MIV use-case. While different ways in which subpictures could be used in MIV are discussed, a particular case study is selected. Namely, subpictures are used for parallel encoding and to reduce the number of decoder instances. Experimental results show that the cost of using subpictures in terms of bitrate overhead is negligible (0.1% to 0.4%), when compared to the overall bitrate. The number of decoder instances on the other hand decreases by a factor of two.
María Santamaría 0001, Vinod Kumar Malamal Vadakital, Lukasz Kondrad, Antti Hallapuro, Miska M. Hannuksela
MMSP5
2021 VVC Adaptive Loop Filter Optimization for Subpicture-based Viewport-adaptive Streaming
abstract
Virtual reality (VR) systems require delivering high-fidelity 360° video content to immerse viewers to the captured scene. The viewport-adaptive streaming (VAS) methods have been developed to deliver 360° VR content efficiently. The Versatile Video Coding (VVC) standard introduces the subpicture picture partitioning tool, which creates isolated regions suitable for VAS. The usage of Adaptive Loop Filter (ALF) as a VVC in-loop filtering operation is limited in subpicture-based VAS. This paper aims at enabling usage of ALF in subpicture-base VAS through proposing a set of encoding constraints that are standard compliant. While ALF is activated, the proposed constrains guarantee that no coding coordination with respect to sharing of ALF parameters among subpictures is required. This further allows subpicture-based parallel encoding of high-resolution VR content. We study the performance of several methods targeting both single- and multi-thread encoding platforms. The experimental results indicate that VR content encoding can be parallelized at a subpicture-group level, while still preserving most of the ALF gain. The proposed method with subpicture-group encoding parallelization achieves on average -2.4%, -4.0%, and -4.3% Bjøntegaard delta rate reduction for Y, U, and V components respectively, compared to the case where ALF operation is deactivated.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela
MMSP3
2021 Learn to overfit better: finding the important parameters for learned image compression
abstract
For most machine learning systems, overfitting is an undesired behavior. However, overfitting a model to a test image or a video at inference time is a favorable and effective technique to improve the coding efficiency of learning-based image and video codecs. At the encoding stage, one or more neural networks that are part of the codec are finetuned using the input image or video to achieve a better coding performance. The encoder en-codes the input content into a content bitstream. If the finetuned neural network is part (also) of the decoder, the encoder signals the weight update of the finetuned model to the decoder along with the content bitstream. At the decoding stage, the decoder first updates its neural network model according to the received weight update, and then proceeds with decoding the content bitstream. Since a neural network contains a large number of parameters, compressing the weight update is critical to reducing bitrate overhead. In this paper, we propose learning-based methods to find the important parameters to be overfitted, in terms of rate-distortion performance. Based on simple distribution models for variables in the weight update, we derive two objective functions. By optimizing the proposed objective functions, the importance scores of the parameters can be calculated and the important parameters can be determined. Our experiments on lossless image compression codec show that the proposed method significantly outperforms a prior-art method where overfitted parameters were selected based on heuristics. Furthermore, our technique improved the compression performance of the state-of-the-art lossless image compression codec by 0.1 bit per pixel.
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, María Santamaría 0001, Yat-Hong Lam, Miska M. Hannuksela
VCIP6
2021 An Overview of Omnidirectional MediA Format (OMAF)
abstract
During recent years, there have been product launches and research for enabling immersive audio-visual media experiences. For example, a variety of head-mounted displays and 360° cameras are available in the market. To facilitate interoperability between devices and media system components by different vendors, the Moving Picture Experts Group (MPEG) developed the Omnidirectional MediA Format (OMAF), which is arguably the first virtual reality (VR) system standard. OMAF is a storage and streaming format for omnidirectional media, including 360° video and images, spatial audio, and associated timed text. This article provides a comprehensive overview of OMAF.
Miska M. Hannuksela, Ye-Kui Wang
Proc. IEEE1
2021 The High-Level Syntax of the Versatile Video Coding (VVC) Standard
abstract
Versatile Video Coding (VVC), a.k.a. ITU-T H.266 | ISO/IEC 23090-3, is the new generation video coding standard that has just been finalized by the Joint Video Experts Team (JVET) of ITU-T VCEG and ISO/IEC MPEG at its$19^{\mathrm {th}}$meeting ending on July 1, 2020. This paper gives an overview of the VVC high-level syntax (HLS), which forms its system and transport interface. Comparisons to the HLS designs in High Efficiency Video Coding (HEVC) and Advanced Video Coding (AVC), the previous major video coding standards, are included. When discussing new HLS features introduced into VVC or differences relative to HEVC and AVC, the reasoning behind the design differences and the benefits they bring are described. The HLS of VVC enables newer and more versatile use cases such as video region extraction, composition and merging of content from multiple coded video bitstreams, and viewport-adaptive 360° immersive media.
Ye-Kui Wang, Robert Skupin, Miska M. Hannuksela, Sachin Deshpande, Hendry, Virginie Drugeon, Rickard Sjöberg, Byeongdoo Choi, Vadim Seregin, Yago Sánchez de la Fuente, Jill M. Boyce, Wade Wan, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.3
2020 Lossless Image Compression Using a Multi-scale Progressive Statistical Model
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Nannan Zou, Emre Aksu, Miska M. Hannuksela
ACCV (3)6
2020 On Subpicture-based Viewport-dependent 360-degree Video Streaming using VVC
abstract
Virtual reality applications create an immersive experience using 360° video with high resolution and frame rate. However, since the user only views a portion of 360° video according to his/her current viewport, streaming the whole content with high resolution causes bandwidth wastage. To address this issue, viewport-dependent approaches have been proposed such that only the part of the video which falls within user's current viewport is transmitted in high quality while the rest of the content is transmitted in lower quality. The selection of high- and low-quality parts is constantly adapted according to the user's head motion, which requires frequent intra coded frames at switching points, leading to an increment in the overall streaming bitrate. In this paper a viewport-adaptive streaming scheme is introduced, which avoids intra frames at switching points by introducing long intra period for non-changing parts of the content during head motion. This scheme has been realized taking advantage of mixed Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit feature of Versatile Video Coding (VVC) standard. This method reduces bitrate significantly, especially for the sequences with either no or only slow camera motion, which is common for 360° video capturing.
Maryam Homayouni, Alireza Aminlou, Miska M. Hannuksela
ISM3
2020 Efficient Adaptation of Neural Network Filter for Video Compression
abstract
We present an efficient finetuning methodology for neural-network filters which are applied as a postprocessing artifact-removal step in video coding pipelines. The fine-tuning is performed at encoder side to adapt the neural network to the specific content that is being encoded. In order to maximize the PSNR gain and minimize the bitrate overhead, we propose to finetune only the convolutional layers' biases. The proposed method achieves convergence much faster than conventional finetuning approaches, making it suitable for practical applications. The weight-update can be included into the video bitstream generatedby the existing video codecs. We show that our method achieves up to 9.7% average BD-rate gain when compared to the state-of-art Versatile Video Coding (VVC) standard codec on 7 test sequences.
Yat Hong Lam, Alireza Zare, Francesco Cricri, Jani Lainema, Miska M. Hannuksela
ACM Multimedia5
2020 L2C - Learning to Learn to Compress
abstract
In this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments.
Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu
MMSP6
2019 An Overview of the OMAF Standard for 360° Video
abstract
Omnidirectional MediA Format (OMAF) is arguably the first virtual reality (VR) system standard, recently developed by the Moving Picture Experts Group (MPEG). OMAF defines a media format that enables omnidirectional media applications, focusing on 360° video, images, and audio, as well as the associated timed text, supporting three degrees of freedom (3DOF). This paper gives an overview of the first edition of the OMAF standard.
Miska M. Hannuksela, Ye-Kui Wang, Ari Hourunranta
DCC1
2019 Approximating Binarization in Neural Networks
abstract
Binarization of neural networks' activations may be a requirement for some applications. A typical example is end-to-end learned deep image compression systems where the encoder's output is requred to be a binary vector. Binarization is non-differentiable, therefore one needs to approximate it in order to train neural networks with stochastic gradient descent. In this paper, we investigate these training strategies and provide improvements over baselines. We find that during training, constraining the activations in a region that is far away from binary points leads to a better performance at test-time. The above finding provides a counter-intuitive result and leads to re-thinking the binarization approximation problem in neural networks.
Çaglar Aytekin, Francesco Cricri, Jani Lainema, Emre Aksu, Miska M. Hannuksela
IJCNN5
2019 Shared Coded Picture Technique for Tile-Based Viewport-Adaptive Streaming of Omnidirectional Video
abstract
Tile-based viewport-adaptive streaming methods have been used in delivering omnidirectional video for virtual reality applications. In these methods, the 360° video is encoded in multiple quality versions by using the motion constrained tile set (MCTS) technique. A set of high-quality and low-quality tiles, corresponding to viewport and non-viewport areas, respectively, are selected and transmitted to the user. However, these methods require frequent intra random access points to ensure seamless viewport switching capability, very high decoding complexity, or a multi-layer coding scheme. The frequent intra random access points include very high bitrate in viewport switching points. The high decoding complexity and multi-layer decoder requirements are not aligned with the omnidirectional media format (OMAF) standard. Such requirements make these methods sub-optimal or impractical for streaming the omnidirectional video. This paper studies the current tile-based solutions for delivering the omnidirectional content. Moreover, the OMAF-compliant shared coded picture (SCP)-based scheme is proposed in this paper for streaming the omnidirectional video. The core concept of the SCP-based method is to manipulate the switching point pictures in a way that the frequent intra-coded pictures are no longer required for the viewport switching operations between different quality versions of the content. The experiments illustrated that the SCP-based method outperforms the MCTS-based method on average by 11% to 14% in terms of streaming bitrate reduction with only 4% extra decoding complexity.
Ramin Ghaznavi Youvalari, Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.4
2019 6K and 8K Effective Resolution with 4K HEVC Decoding Capability for 360 Video Streaming
abstract
The recent Omnidirectional MediA Format (OMAF) standard, which specifies the delivery of 360° video content, supports only equirectangular projection (ERP) and cubemap projection and their region-wise packing with a limitation on video decoding capability to the maximum resolution of 4K (e.g., 4,096 × 2,048). Streaming of 4K ERP content allows only a limited viewport resolution, which is lower than the resolution of many current head-mounted displays (HMDs). Therefore, to take full advantage of high-resolution HMDs, delivery of 360° video content beyond 4K resolution needs to be enabled. In this regard, we propose two specific mixed-resolution packing schemes of 6K (e.g., 6,144 × 3,072) and 8K (e.g., 8,192 × 4,096) ERP content and their realization in tile-based streaming, while complying with the 4K decoding constraint and the High Efficiency Video Coding standard. The proposed packing schemes offer 6K and 8K effective resolution at the viewport. Using our proposed test methodology, experimental results indicate that the proposed layouts significantly decrease streaming bitrates when compared to mixed-quality viewport-adaptive streaming of 4K ERP. Our results further indicate that 8K-effective packing outperforms 6K-effective packing especially in high-quality videos.
Alireza Zare, Maryam Homayouni, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ACM Trans. Multim. Comput. Commun. Appl.4
2018 2D Video Coding of Volumetric Video Data
abstract
Due to the increased popularity of augmented and virtual reality experiences, the interest in representing the real world in an immersive fashion has never been higher. Distributing such representations enables users all over the world to freely navigate in never seen before media experiences. Unfortunately, such representations require a large amount of data, not feasible for transmission on today's networks. Thus, efficient compression technologies are in high demand. This paper proposes an approach to compress 3D video data utilizing 2D video coding technology. The proposed solution was developed to address the needs of `tele-immersive' applications, such as virtual (VR), augmented (AR) or mixed (MR) reality with Six Degrees of Freedom (6DoF) capabilities. Volumetric video data is projected on 2D image planes and compressed using standard 2D video coding solutions. A key benefit of this approach is its compatibility with readily available 2D video coding infrastructure. Furthermore, objective and subjective evaluation shows significant improvement in coding efficiency over reference technology.
Sebastian Schwarz, Miska M. Hannuksela, Vida Fakour Sevom, Nahid Sheikhi-Pour
PCS2
2017 Perceptual quality assessment of HEVC main profile depth map compression for six degrees of freedom virtual reality video
abstract
The last years have shown significant advances in immersive media. Virtual reality video, in the form of spherical panoramic video, is already widely available. Such technologies will continually evolve towards the ultimate goal of a truly virtual reality experience. The first step in this direction is the support of limited translational head movement for restricted six degrees of freedom virtual reality video. This experience can be achieved by rendering virtual viewports from supplementary depth information. In this context, this paper investigates the effect of depth map compression on the perceptual quality of immersive media. The presented study is focused on near-future applications with readily available hardware. Objective quality assessments indicate possible bit rate savings of 17% through high-quality depth maps. However, these findings could not be confirmed subjectively. On the contrary, subjective evaluation shows a robustness to low-quality depth maps when viewed in a virtual reality scenario.
Sebastian Schwarz, Miska M. Hannuksela
ICIP2
2017 Nested polygonal chain mapping of omnidirectional video
abstract
Conventionally, the equirectangular projection (ERP) has been used for representing 360-degree panorama videos. However, ERP suffers from over stretching in polar regions and thus increases the bitrate and the encoding/decoding complexity. This paper presents the nested polygonal chain packing method where the top and bottom stripes of an ERP picture are resampled sample-row-wise. The sampling ratio is a linear function of the sample row being processed. The height of the top or bottom stripe can be for example a quarter of the height of the ERP picture. The resampled sample rows are arranged in nested polygonal chains. According to the presented experimental results, the proposed nested polygonal chain packing provides on average 6.1% bitrate reduction compared to ERP.
Kashyap Kammachi Sreedhar, Miska M. Hannuksela
ICIP2
2017 Virtual reality content streaming: Viewport-dependent projection and tile-based techniques
abstract
Virtual reality (VR) head-mounted display (HMD) requires spherical panoramic contents with high-spatial and temporal fidelity to immerse the viewers into the captured scene. Hereby, VR contents are extremely bandwidth intensive and impose technical challenges for the design of a VR streaming system. A bandwidth-efficient VR streaming system can be achieved using the viewport-aware adaptation techniques, in which part of the sphere within the viewer's field of view is presented at higher quality. In this paper, two recently emerged viewport-adaptive streaming methods so-called tile-based method and truncated square pyramid (TSP) projection, a well-studied viewport-dependent projection, are compared using a proposed quality assessment methodology. The comparison is made in terms of storage and streaming bitrate performances. The simulation results indicate that the tile-based approach has slightly lower streaming performance, while offering a significant storage and encoding time saving at the server side, compared to TSP-based streaming.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela
ICIP3
2017 Comparison of HEVC coding schemes for tile-based viewport-adaptive streaming of omnidirectional video
abstract
Virtual reality applications make use of 360-degree panoramic or omnidirectional video with high resolution and high frame rate in order to create the immersive experience to the user. The user views only a portion of the captured 360-degree scene at each time instant, hence streaming the whole omnidirectional video in highest quality is not efficient. In order to alleviate the problem of bandwidth wastage, viewport-adaptive encoding and streaming schemes have been proposed. In these schemes, part of the captured scene that is within the viewer's field of view is delivered at highest quality while the rest of the scene in a lower quality. In this work, three tile-based viewport-adaptive methods using motion-constrained tile sets (MCTS), region-of-interest scalability and simulcast approach have been studied for streaming omnidirectional content. In the performed experiments with various tiling arrangements, MCTS-based scheme required highest bitrate compared to other methods. The scalable coding scheme provided the highest performance in terms of streaming bitrate saving on average up to 53% and 35% compared to streaming the whole omnidirectional video and MCTS-based method, respectively.
Ramin Ghaznavi Youvalari, Alireza Zare, Huameng Fang, Alireza Aminlou, Qingpeng Xie, Miska M. Hannuksela, Moncef Gabbouj
MMSP6
2017 Standardization status of 360 degree video coding and delivery
abstract
The emergence of consumer level capturing and display devices for 360 degree video creates new and promising segments in entertainment, education, professional training, and other markets. In order to avoid market fragmentation and ensure interoperability of 360 degree video ecosystems, industry and academia cooperate in standardization efforts in this field. In the video coding domain, 360 degree video invalidates many established procedures, e.g., concerning evaluation of the visual quality, while the specific content characteristics offer potential for higher compression efficiency beyond the current standards. Likewise, 360 degree video puts stricter demands on the system level aspects of transmission but may also offer the potential to enhance existing transport schemes. The Joint Collaborative Team on Video Coding (JCT-VC) as well as the Joint Video Exploration Team (JVET) already started investigations into 360 degree video coding while numerous activities in the Systems subgroup of the Moving Picture Experts Group (MPEG) started to investigate application requirements and delivery aspects of 360 degree video. This paper reports on the current status of the outlined standardization efforts.
Robert Skupin, Yago Sánchez de la Fuente, Ye-Kui Wang, Miska M. Hannuksela, Jill M. Boyce, Mathias Wien
VCIP4
2017 Row-Interleaved Sampling for Depth-Enhanced 3D Video Coding for Polarized Displays
abstract
Passive stereoscopic displays create the illusion of three dimensions by employing orthogonal polarizing filters and projecting two images onto the same screen. In this article, a coding scheme targeting depth-enhanced stereoscopic video coding for polarized displays is introduced. We propose to use asymmetric row-interleaved sampling for texture and depth views prior to encoding. The performance of the proposed scheme is compared with several other schemes, and the objective results confirm the superior performance of the proposed method. Furthermore, subjective evaluation proves that no quality degradation is introduced by the proposed coding scheme compared to the reference method.
Maryam Homayouni, Payman Aflaki, Miska M. Hannuksela, Moncef Gabbouj
ACM Trans. Appl. Percept.3
2016 Fisheye video coding using elastic motion compensated reference frames
abstract
Fisheye cameras have become extremely popular in applications where the goal is to capture large fields of view with only one camera. However, the wide-angle fisheye imagery has special characteristics that may not be very well suited for modern video codecs that employ block-based translational motion model. This model fails to describe complex deformable motion which is often present in fisheye videos. In this paper, we advocate for the usage of elastic motion model in compensating such a complex motion. The presented design enables the re-use of existing codecs, such as HEVC, without modifications in low-level coding tools. Experimental results show that a savings in bit rate of up to 6.54% is achievable over standalone HEVC if the elastic motion compensated prediction is used as an additional reference frame.
Ashek Ahmmed, Miska M. Hannuksela, Moncef Gabbouj
ICIP2
2016 HEVC still image coding and high efficiency image file format
abstract
The High Efficiency Video Coding (HEVC) standard includes support for a large range of image representation formats and provides an excellent image compression capability. The High Efficiency Image File Format (HEIF) offers a convenient way to encapsulate HEVC coded images, image sequences and animations together with associated metadata into a single file. This paper discusses various features and functionalities of the HEIF file format and compares the compression efficiency of HEVC still image coding to that of JPEG 2000. According to the experimental results HEVC provides about 25% bitrate reduction compared to JPEG 2000, while keeping the same objective picture quality.
Jani Lainema, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital, Emre Aksu
ICIP2
2016 RTP/RTCP Reception Hint Tracks for Video Call Recording and Playback
abstract
This paper proposes to use RTP (Real-time Transport Protocol) Reception Hint Tracks for convenient recording and playback of a video call in MP4 format. The feasibility of RTP Reception Hint Tracks is validated as a part of an implemented end-to-end video call system. The proposed approach records a bidirectional Linphone video call and multiplexes it as MP4 RTP Reception Hint Tracks with the L-SMASH software library. It also stores RTCP (RTP Control Protocol) Reception Hint Tracks for additional timing information. Playback of the MP4 file is performed with VLC Media Player that is made compatible with RTP Reception Hint Tracks. The proposed proof-of-concept setup meets particularly well the needs of multi-codec solutions where different audio and video codecs can be used for a video call recording and playback. According to our analysis, recording RTP reception Hint Tracks increases the Linphone CPU time by under 1% and the bitrate by under 2% over the bare bitrate of the recorded RTP packets.
Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital
ISM5
2016 Standard-Compliant Multiview Video Coding and Streaming for Virtual Reality Applications
abstract
Virtual reality (VR) systems employ multiview cameras or camera rigs to capture a scene from the entire 360-degree perspective. Due to computational or latency constraints, it might not be possible to stitch multiview videos into a single video sequence prior to encoding. In this paper we investigate the coding and streaming of multiview VR video content. We present a standard-compliant method where we first divide the camera views into two types: Primary views represent a subset of camera views with lower resolution and non-overlapping (minimally overlapping) content which cover the entire 360-degree field-of-view to guarantee immediate monoscopic viewing during very rapid head movements. Auxiliary views consist of remaining camera views with higher resolution which produce overlapping content with the primary views and are additionally used for stereoscopic viewing. Based on this categorization, we propose a coding arrangement in which, the primary views are independently coded in the base layer and the additional auxiliary views are coded as an enhancement layer, using inter-layer prediction from primary views. The proposed system not only meets the low latency requirements of VR systems, but also conforms to the existing multilayer extensions of the High Efficiency Video Coding standard. Simulation results show that the coding and streaming performance of the proposed scheme is significantly improved compared to earlier methods.
Kashyap Kammachi Sreedhar, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ISM3
2016 Viewport-Adaptive Encoding and Streaming of 360-Degree Video for Virtual Reality Applications
abstract
Virtual reality applications use 360-degree videos and head mount displays (HMDs) with stereoscopic capabilities to provide full immersion experience. In these applications it is also common to use 4K resolution or higher per view for 360-degree videos. Consequently, this leads to technical challenges in handling the bandwidth requirements while keeping the system latency to the minimal. When the content is viewed with a HMD, a subset of the entire 360-degree video is displayed at a single point of time. To improve the resolution and picture quality of the displayed content, viewport based coding is desirable. In this regard, we investigated various viewport dependent projection schemes including the existing variants of Pyramidal projection. In this regard we propose the multi-resolution versions of Equirectangular and Cubemap projections. Additionally, we developed a methodology for comparing the rate-distortion performance of these projections. Based on the simulation results, it was observed that multi-resolution projections of Equirectangle and Cubemap outperform other projection schemes, significantly.
Kashyap Kammachi Sreedhar, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ISM3
2016 Efficient Coding of 360-Degree Pseudo-Cylindrical Panoramic Video for Virtual Reality Applications
abstract
Pseudo-cylindrical panoramas represent the data distribution of spherical coordinates closely in two-dimensional domain due to the equidistant sampling of 360-degree scene. Therefore, unlike the cylindrical projections, they do not suffer from the over stretching in the polar areas. However, due to the non-rectangular format in effective picture area and sharp edges at its borders, the compression performance is inefficient. In this paper, we propose two methods which improve the compression performance of both intra-frame and inter-frame coding of pseudo-cylindrical panoramic content and meanwhile reduce the coding artifacts. In the intra-frame coding method, border edges are smoothed by modifying the content of the image in the non-effective picture area, which are cropped at the receiver side. In the inter-frame coding method, gaining the benefit of 360-degree property of the content, non-effective picture area of reference frames at border is filled with the content of the effective picture area from the opposite border to enhance the performance of motion compensation.
Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ISM3
2016 HEVC-compliant Tile-based Streaming of Panoramic Video for Virtual Reality Applications
abstract
Delivering wide-angle and high-resolution spherical panoramic video content entails a high streaming bitrate. This imposes challenges when panorama clips are consumed in virtual reality (VR) head-mounted displays (HMD). The reason is that the HMDs typically require high spatial and temporal fidelity contents and strict low-latency in order to guarantee the user's sense of presence while using them. In order to alleviate the problem, we propose to store two versions of the same video content at different resolutions, each divided into multiple tiles using the High Efficiency Video Coding (HEVC) standard. According to the user's present viewport, a set of tiles is transmitted in the highest captured resolution, while the remaining parts are transmitted from the low-resolution version of the same content. In order to enable randomly choosing different combinations, the tile sets are encoded to be independently decodable. We further study the trade-off in the choice of tiling scheme and its impact on compression and streaming bitrate performances. The results indicate streaming bitrate saving from 30% to 40%, depending on the selected tiling scheme, when compared to streaming the entire video content.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ACM Multimedia3
2016 Analysis of regional down-sampling methods for coding of omnidirectional video
abstract
In order to compress omnidirectional video clips, a projection onto a two-dimensional image plane is necessary. The most commonly used projection format is the equirectangular panoramic projection, which results into a significant amount of redundant samples in the polar areas. The redundant samples incur extra bitrate and increase the encoding/decoding time. In this paper, we study regional down-sampling (RDS) for achieving better compression and smaller encoding/decoding time for omnidirectional content. We extend the persistent RDS method applied equally to all pictures to be applied to selected pictures only in our proposed temporal RDS method and then compare the persistent and temporal RDS methods. The simulation results indicate that both the persistent and temporal RDS improve the rate-distortion (RD) performance compared to the conventional coding of equirectangular panoramas, while the temporal RDS method has less sequence-wise RD performance variation and slightly better RD performance on average when compared to the persistent RDS technique. Alongside the coding methods, we study spherical quality measurement methods for VR images/video and analyze the coding methods with these quality metrics. Moreover, we propose a uniformly sampled spherical quality metric in order to evaluate the coding distortion of omnidirectional videos.
Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela
PCS3
2016 HEVC-compliant viewport-adaptive streaming of stereoscopic panoramic video
abstract
Virtual reality (VR) provides unprecedented immersive experience using high-resolution spherical stereoscopic panoramic video. Such an experience is achieved by using head-mounted display (HMD) which has very strict latency bounds in order to respond promptly to user movements. Conventional streaming of VR video requires large bandwidth because the entire captured panorama is transmitted. However, only a limited field-of-view (FOV) is displayed by an HMD, resulting in wastage of bandwidth. To alleviate the problem, this paper proposes a High Efficiency Video Coding (HEVC) compliant approach for efficient coding and streaming of stereoscopic VR content. The proposed method is based on partitioning video pictures into tiles, where only the required tiles corresponding to the primary viewport are transmitted in high resolution, while the remaining parts are transmitted in low resolution. Furthermore, this method enables coding stereoscopic video contents using a conventional HEVC codec, while still achieving significant compression gain by means of adopting inter-view prediction only in intra random access point (IRAP) pictures. Using this method, the predicted view can be decoded independently of the main view, hence allowing simultaneous decoding instances. Experimental results demonstrate that the proposed approach is able to substantially improve compression efficiency and streaming bitrate performance.
Alireza Zare, Kashyap Kammachi Sreedhar, Vinod Kumar Malamal Vadakital, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
PCS5
2015 Who is moving - user or device?: experienced quality of mobile 3d video in vehicles
abstract
'Viewing while commuting' is a typical use case for mobile video. However, experimental and behavioral influences of watching three-dimensional (3D) video in vibrating vehicles have not been widely researched. The goal of this study is 1) to explore the influence of video presentation modes (two-dimensional and stereoscopic 3D) on the quality of experience and 2) to understand the nature of the movement patterns that users perform to maintain an optimal viewing position while viewing videos on a mobile device in three commuting contexts and in a controlled laboratory environment. A hybrid method for quality evaluation was used for combining quantitative preference ratings, qualitative descriptions of quality, situational audio/video data-collection, and sensors. The high-quality and heterogeneous audiovisual stimuli were viewed on a mobile device equipped with a parallax barrier display. The results showed that the stereoscopic 3D (S3D) video presentation mode provided more satisfying quality of experience than the two-dimensional presentation mode in all studied contexts. To maintain an optimal viewing position in the vehicles, the users moved the device in their hands to the directions around the vertical and the horizontal axes in a leaned sitting position. This movement behavior was guided by the contexts but not by the quality, indicating the general importance of these results for mobile video viewing in vibrating vehicles.
Satu Jumisko-Pyykkö, Panos Markopoulos 0001, Miska M. Hannuksela
Advances in Computer Entertainment3
2015 Overview of the multiview high efficiency video coding (MV-HEVC) standard
abstract
This paper reviews the multiview extension (MV-HEVC) of the High Efficiency Video Coding (HEVC) standard. MV-HEVC is capable of multiview video coding with or without accompanying depth views. The key design concepts and design elements of MV-HEVC are described in the paper. Furthermore, the features and characteristics of MV-HEVC compared to other standardized video codec extensions for three-dimensional (3D) video coding are reviewed.
Miska M. Hannuksela, Xuehui Huang, Houqiang Li
ICIP1
2015 Upsampled-view distortion optimization for mixed resolution 3D Video Coding
abstract
The MVC+D extension of the Advanced Video Coding (H.264/AVC) standard enables multiview-and-depth 3D video coding but specifies that all views are coded at equal spatial resolution. In mixed resolution 3D video coding some of the views are coded at reduced resolution. This paper proposes an improvement for the mode decisions in depth encoding in the mixed resolution scenario. We modify the distortion calculation for rate-distortion optimized depth coding. The proposed solution optimizes the depth data compression with assumption that it will be used not only for view synthesis but also for depth-based super resolution in the post-processing stage. The algorithm is implemented on top of the mixed resolution 3D video encoder based on the 3DV-ATM reference software. Evaluation of the proposed solution, tested under the JCT-3V common test conditions, is done against the mixed resolution MVC+D coding with the view synthesis distortion enabled. The results show 2.64%dBR gain for coded views and 0.64% gain for synthesized views.
Michal Joachimiak, Miska M. Hannuksela, Payman Aflaki, Moncef Gabbouj
ICIP2
2015 Seamless switching of H.265/HEVC-coded dash representations with open GOP prediction structure
abstract
The Dynamic Adaptive Streaming over HTTP (DASH) enables bitrate adaptation through different representations of the same content. It is common to encode random access point (RAP) pictures at segment boundaries to support representation switching. As an open group of pictures (GOP) results into a temporary discontinuity of the video playback due to the inability to decode some pictures when switching representations, closed GOP prediction structures are normally used in DASH. This paper proposes two similar methods for using the open GOP prediction structure in DASH representations while a full picture rate is maintained also during representation switching. The first method is enabled with straightforward changes in the decoding of the High Efficiency Video Coding (H.265/HEVC) standard, whereas the second method utilizes the adaptive resolution change feature of the scalable (SHVC) extension of H.265/HEVC. Experiments show that the proposed methods outperform the use of closed GOPs by 5.6% on average in terms of Bjontegaard delta bitrate (BD-rate).
Miska M. Hannuksela, Houqiang Li
ICIP2
2015 Disparity-compensated inter-layer motion prediction using standardized HEVC extensions
abstract
The scalable and multiview extensions of the High Efficiency Video Coding share the same high-level syntax coding structure. For the scalable extension, the motion field of the inter-layer reference picture is modified through Motion Field Mapping before used for motion vector prediction. However, the motion field of the inter-layer reference picture is used without modification in the multiview extension. In this paper, a disparity-compensated inter-layer motion prediction is proposed for multiview video coding to achieve disparity compensation in inter-layer motion prediction using the scaled reference layer offset. The experimental results show that comparing with the multiview extension anchor, the proposed method achieves an average of 1.1% bitrate reduction.
Miska M. Hannuksela, Houqiang Li
ISCAS2
2015 Improved downstream rate-distortion performance of SHVC in DASH using sub-layer-selective interlayer prediction
abstract
Dynamic Adaptive Streaming over HTTP (DASH) has gained wide acceptance due to its ability of bitrate adaptation in diverse network conditions. In the meanwhile, it is asserted that the Scalable Extension (SHVC) of High Efficiency Video Coding (HEVC) can bring more efficient storage and caching compared with traditional single-layer video coding. However, using of SHVC in DASH raises the downstream bitrate relative to encoding each representation independently. This paper proposes a method to improve the downstream rate-distortion performance of SHVC in DASH. Experimental results show that the proposed method can bring up to 16.9 %-unit downstream bitrate reduction compared with directly use SHVC in DASH, while the efficiency in storage space is maintained.
Xuehui Huang, Miska M. Hannuksela, Houqiang Li
MMSP2
2015 Scalable Bit Allocation Between Texture and Depth Views for 3-D Video Streaming Over Heterogeneous Networks
abstract
In the multiview video plus depth (MVD) coding format, both texture and depth views are jointly compressed to represent the 3-D video content. The MVD format enables synthesis of virtual views through depth-image-based rendering; hence, distortion in the texture and depth views affects the quality of the synthesized virtual views. Bit allocation between texture and depth views has been studied with some promising results. However, to the best of our knowledge, most of the existing bit-allocation methods attempt to allocate a fixed amount of total bit rate between texture and depth views; that is, to select appropriate pair of quantization parameters for texture and depth views to maximize the synthesized view quality subject to a fixed total bit rate. In this paper we propose a scalable bit-allocation scheme, where a single ordering of texture and depth packets is derived and used to obtain optimal bit allocation between texture and depth views for any total target rates. In the proposed scheme, both texture and depth views are encoded using the quality scalable coding method; that is, medium grain scalable (MGS) coding of the Scalable Video Coding (SVC) extension of the Advanced Video Coding (H.264/AVC) standard. For varying target total bit rates, optimal bit truncation points for both texture and depth views can be obtained using the proposed scheme. Moreover, we propose to order the enhancement layer packets of the H.264/SVC MGS encoded depth view according to their contribution to the reduction of the synthesized view distortion. On one hand, this improves the depth view packet ordering when considered the rate-distortion performance of synthesized views, which is demonstrated by the experimental results. On the other hand, the information obtained in this step is used to facilitate optimal bit allocation between texture and depth views. Experimental results demonstrate the effectiveness of the proposed scalable bit-allocation scheme for texture and depth views.
Jimin Xiao, Miska M. Hannuksela, Tammam Tillo, Moncef Gabbouj, Ce Zhu, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2014 Flexible depth map spatial resolution in depth-enhanced multiview video coding
abstract
Multiview video plus depth (MVD) has proved to be a promising format enabling various 3D applications. One approach to achieve better MVD compression is to adjust the spatial resolution of depth map based on the content and the application. In this research two schemes are considered: first, multiview video coding accompanied by depth maps to improve the texture coding performance. Second, multiview video plus depth (MVD) coding targeting highest quality for synthesized views. Two algorithms to select the best spatial resolution for each scheme are proposed and the results show 10.8% and 16.5% bitrate reduction compared to the anchor case where depth map resolution is fixed.
Payman Aflaki, Miska M. Hannuksela, Moncef Gabbouj
ICASSP2
2014 Weighted-prediction-based color gamut scalability extension for the H.265/HEVC video codec
abstract
Color gamut scalability refers to coding a video in a layered manner where the base and enhancement layers are coded in different color gamut spaces. Color gamut scalability and its relationship with spatial scalability are currently being studied for the scalable extension of HEVC (SHVC) to enable coding of ultra-high definition content having BT.2020 color gamut with 10-bit precision as an enhancement layer and high definition content having BT.709 color gamut with 8-bit precision as the base layer. In this paper, we propose to use the weighted prediction tool of the SHVC standard to map the color gamut of the base layer to the enhancement layer. In addition, we also propose a high-precision bit-depth mapping of the base layer to the enhancement layer that jointly performs upsampling with a bit-depth increase. Simulation results show that these two schemes improve the coding efficiency of the All Intra and Random Access configurations by about 6.8% and 3.6% on average, respectively, compared to a basic scheme where the bit-depth of the base layer is increased by simple bit-shifting. These gains are achieved by imposing no changes to the SHVC standard; hence make the proposed method very useful for practical use-cases as well.
Alireza Aminlou, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
ICASSP3
2014 Backward compatible enhancement of chroma format in HEVC
abstract
First version of the latest video coding standard, High Efficiency Video Coding (HEVC), only supports coding of video in YUV 4:2:0 chroma format. An extension of the standard that will support other chroma formats is currently under development, however, version 1 decoders will not be able to handle the bitstreams created using this extension. In this paper, we propose a novel method to create scalable bitstreams that involve a backward compatible base layer in 4:2:0 format that can be handled by HEVC version 1 decoders and code additional layers to enhance the chroma resolution. The proposal codes 4:2:0 video in the base layer and the high resolution chroma components as auxiliary pictures as separate enhancement layers. The high resolution chroma components could optionally be predicted from the upsampled 4:2:0 chroma components of base layer. The simulations show that the proposed method achieves scalability with 9.5% coding efficiency penalty on average compared to single layer coding of 4:4:4 video. When compared to simulcast of 4:2:0 and 4:4:4 video, proposed method provides 38% gain on average. Proposed method makes services using high chroma fidelity easier to be deployed, due to the backwards compatibility to existing HEVC implementations with high coding efficiency.
Döne Bugdayci Sansli, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
ICIP3
2014 Adaptive Spatial Resolution Selection for Stereoscopic Video Compression with MV-HEVC: A Frequency Based Approach
abstract
One approach for stereoscopic video compression is to down sample the content prior to encoding and up sample it to the original spatial resolution after decoding. In this study it is shown that the ratio by which the content should be rescaled is sequence dependent. Hence, a frequency based method is introduced enabling fast and accurate estimation of the best down sampling ratio for different stereoscopic video clips. It is shown that exploiting this approach can bring 3.38% delta bitrate reduction over five camera-captured sequences.
Payman Aflaki, Miska M. Hannuksela, Moncef Gabbouj
ISM2
2014 Joint depth and texture filtering targeting MVD compression
abstract
To meet the requirement of current bandwidth and storage facilities, it is required to further decrease the bitrate of 3D video content. This paper presents a scheme to partially filter different regions of the texture views and depth maps taking into account the characteristics of both of them. A combination of the distance of objects from the cameras and the spatial details of texture within the objects defines the location and strength of the smoothing filter that should be applied to that image. A series of subjective tests was conducted to confirm that the perceived quality of the filtered content remains intact and no quality degradations is introduced by the filtering steps. Finally, objective measurements show a Bjontegaard delta bitrate reduction up to 23.77% with an average of 15.2%.
Payman Aflaki, Miska M. Hannuksela, Maryam Homayouni, Moncef Gabbouj
VCIP2
2014 Improved weighted prediction based color gamut scalability in SHVC
abstract
One use case that the scalable extension (SHVC) of the state-of-the-art High Efficiency Video Coding (HEVC) standard aims for is to support Ultra High Definition (UHD) TV broadcast in a backwards compatible way with the existing High Definition (HD) TV broadcast. However, since UHD content typically has higher bit-depth and wider color gamut in addition to increased spatial resolution, the compression efficiency is highly affected by the inter-layer processing applied on the base layer picture. This paper proposes an improvement for the weighted prediction based color gamut scalability to have a better mapping between the color gamuts of the base and enhancement layers. The proposed method aims at capturing the nonlinear characteristics of the color gamut mapping using a piecewise linear model, whose parameters are signaled through weighted prediction mechanism and multiple inter-layer reference pictures. Compared to other existing methods for color gamut mapping in SHVC, such as the 3D Look Up Table (LUT) method, the proposed weighted prediction based approach is less complex, as it does not require any changes to the decoder. The simulation results show up to 3.8% Bjontegaard delta bitrate gain in luma for all intra and 3.0% for random access configurations compared to the existing weighted prediction based scalability method in SHVC.
Döne Bugdayci Sansli, Alireza Aminlou, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
VCIP4
2014 Simultaneous 2D and 3D perception for stereoscopic displays based on polarized or active shutter glasses
Payman Aflaki, Miska M. Hannuksela, Hamed Sarbolandi, Moncef Gabbouj
J. Vis. Commun. Image Represent.2
2014 Overview of the MVC + D 3D video coding standard
Ying Chen 0011, Miska M. Hannuksela, Teruhiko Suzuki, Shinobu Hattori
J. Vis. Commun. Image Represent.2
2014 Differential Coding Using Enhanced Inter-Layer Reference Picture for the Scalable Extension of H.265/HEVC Video Codec
abstract
Differential coding methods improve coding efficiency of scalable video codecs by adding the high-frequency component present in the previously coded enhancement layer (EL) pictures to the base layer (BL) picture. This paper proposes a method to enable differential coding in a scalable codec design without affecting the core coding tools, thus allowing a practical implementation to reuse single-layer hardware or software components. This is achieved by creating an additional reference picture called enhanced inter-layer reference (EILR) and inserting it to the EL decoded picture buffer and reference picture lists. An EILR picture is generated by adding differential information to the current inter-layer reference picture. The differential information is calculated using the previously decoded pictures of the BL and EL and the motion information of the BL picture. The proposed method reduces luma total bitrate on average by 2.2% and 2.8% for random access and low-delay test cases, respectively. The improvements are more significant for chroma components with the average bitrate reduction of 6.5%. The measured decoding time increase for a reference software implementation is 16% with negligible overhead on encoding time.
Alireza Aminlou, Jani Lainema, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.4
2013 De-noising of distance maps sensed by time-of-flight devices in poor sensing environment
abstract
We propose a non-local de-noising approach aimed at filtering range data sensed by Photonic Mixer Device sensors. We address specifically the case of poor sensing environment when the reflected signal amplitude is low. In our approach, signal components of phase-delay and amplitude of the sensed signal are regarded as components of a complex-valued variable and processed together in a single step. This imposes better filter adaptivity and similarity weighting. The complex-domain filtering provides additional feedback in the form of improved noise-level confidence, which can be utilized in iterative de-noising schemes. Pre-filtering of individual components is proposed to suppress structural artifacts. Our approach compares favorably with state of the art approaches.
Mihail Georgiev, Atanas P. Gotchev, Miska M. Hannuksela
ICASSP3
2013 Coding of mixed-resolution multiview video in 3D video application
abstract
The emerging MVC+D standard specifies the coding of Multiview Video plus Depth (MVD) data for enabling advanced 3D video applications. MVC+D specifications define the coding of all views of MVD at equal spatial resolution and apply a conventional MVC technique for coding the multiview texture and the depth independently. This paper presents a modified MVC+D coding scheme, where only the base view is coded at the original resolution whereas dependent views are coded at reduced resolution. To enable inter-view prediction, the base view is downsampled within the MVC coding loop to provide a relevant reference for dependent views. At the decoder side, the proposed scheme consists of a post-processing scheme which upsamples of the decoded views to their original resolution. The proposed scheme is compared against the original MVC+D scheme and an average of 4% delta bitrate reduction (dBR) in the coded views and 14.5% of dBR in the synthesized views are reported.
Payman Aflaki, Wenyi Su, Michal Joachimiak, Dmytro Rusanovskyy, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
ICIP5
2013 A fast and accurate re-calibration technique for misaligned stereo cameras
abstract
In this paper, we propose a practical approach for robust rectification of stereo camera setups without use of calibration pattern. Our solution simplifies the process to a non-general case of rectification to avoid explicit use of Fundamental Matrix estimation. The solution shows better or comparable robustness than some of recent solutions, but for much lower computational cost and code complexity.
Mihail Georgiev, Atanas P. Gotchev, Miska M. Hannuksela
ICIP3
2013 Efficient video resolution adaptation using scalable H.265/HEVC
abstract
Dynamically changing the spatial resolution in a video conferencing session is useful for seamlessly adapting the bitrate to changing network conditions and for improving the user experience. Similar to earlier standards, the emerging High Efficiency Video Coding (H.265/HEVC) standard does not allow prediction across different resolutions, so an Instantaneous Decoding Refresh (IDR) picture must be sent to reinitialize the stream when a resolution change happens. IDR pictures take significantly more bits compared to predictively coded pictures. Thus, using them for resolution switching significantly reduces coding efficiency and increases the delay. In this paper we propose a method to support efficient adaptive resolution change using the emerging scalable H.265/HEVC standard. The proposed approach utilizes the inter-layer predicted random access pictures at the enhancement layer for resolution switching, instead of IDR pictures. The experimental results show that when the proposed method was used, the bitrate was reduced at the switching point by 34% on average for the tested video sequences. In addition, visual examples are shown demonstrating the improved visual quality with the proposed method.
Hoda Roodaki, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
ICIP3
2013 Influence Of camera imaging pipeline on stereo-matching quality: An experimental study
abstract
The paper aims at characterizing the role of camera capture process in depth quality estimation by stereo-matching methods. An evaluation software application has been developed which integrates a modular image processing pipeline (IPP) in to the depth estimation process. The application allows for simulating the presence of camera-specific processing artifacts and their influence on the captured stereo imagery in typical imaging scenarios. Furthermore, it allows for benchmarking and optimizing both the depth estimation module and camera capture settings for a jointly-optimal performance.
Mihail Georgiev, Atanas P. Gotchev, Miska M. Hannuksela
ISCAS3
2013 Flexible Coding Order for 3D video extension of H.265/HEVC
abstract
This paper presents a novel Multi-View Video plus Depth (MVD) data coding configuration for the 3D video extension of the High Efficiency Video Coding standard (3D-HEVC). To improve compression efficiency of dependent views, 3D-HEVC utilizes disparity information which is derived from the motion information of the neighboring blocks. However, in the case of MVD data coding, disparity information can be derived directly from the depth data. To utilize this approach, we propose to code the depth maps of the dependent views before their associated texture pictures. Consequently, disparity can be derived from coded depth and can be utilized for coding of texture data. The proposed method improves the compression performance of 3D-HEVC by 0.6% in delta bit rate on average. In addition to this, coded depth maps can be used for Backward View Synthesis Prediction (B-VSP) to further reduce bitrate of dependent texture views. Not only the rate distortion performance is improved, but also the proposed configuration reduces the complexity of decoding by over 8%.
Srikanth Gopalakrishna, Miska M. Hannuksela, Moncef Gabbouj
PCS2
2013 Multiview-video-plus-depth coding and inter-component prediction in high-level-syntax extension of H.265/HEVC
abstract
The multiview-video-plus-depth (MVD) format has become popular in representing 3D video content in a manner that enables free-viewpoint capability in the decoder side. The scalable (SHVC) and multiview (MV-HEVC) extensions of the High Efficiency Video Coding standard (H.265/HEVC) enable their functionality without additional coding tools and use the same high level syntax. In this paper, it is proposed to use the same coding approach as taken in SHVC and MV-HEVC for coding of MVD data. Furthermore, an inter-component motion vector prediction (ICP) method, which is realized through the temporal motion vector prediction (TMVP) mechanism of H.265/HEVC, is introduced to exploit the redundancy between texture and depth views. The experimental results show that 1.0% and 3.6% bitrate reduction can be achieved by ICP compared to independent coding of texture and depth for synthesized views and depth views, respectively, when the ICP method is applied to a 3-view coding scenario intended for multi-view autostereoscopic displays.
Houqiang Li, Miska M. Hannuksela
PCS3
2013 Nonlinear Depth Map Resampling for Depth-Enhanced 3-D Video Coding
abstract
Depth-enhanced 3-D video coding includes coding of texture views and associated depth maps. It has been observed that coding of depth map at reduced resolution provides better rate-distortion performance on synthesized views comparing to utilization of full resolution (FR) depth maps in many coding scenarios based on the Advanced Video Coding (H.264/AVC) standard. Conventional techniques for down and upsampling do not take typical characteristics of depth maps, such as distinct edges and smooth regions within depth objects, into account. Hence, more efficient down and upsampling tools, capable of preserving edges better, are needed. In this letter, novel non-linear methods to down and upsample depth maps are presented. Bitrate comparison of synthesized views, including texture and depth map bitstreams, is presented against a conventional linear resampling algorithm. Objective results show an average bitrate reduction of 5.29% and 3.31% for the proposed down and upsampling methods with ratio ½, respectively, comparing to the anchor method. Moreover, a joint utilization of the proposed down and upsampling brings up to 20% and on average 7.35% bitrate reduction.
Payman Aflaki, Miska M. Hannuksela, Dmytro Rusanovskyy, Moncef Gabbouj
IEEE Signal Process. Lett.2
2013 Multiview-Video-Plus-Depth Coding Based on the Advanced Video Coding Standard
abstract
This paper presents a multiview-video-plus-depth coding scheme, which is compatible with the advanced video coding (H.264/AVC) standard and its multiview video coding (MVC) extension. This scheme introduces several encoding and in-loop coding tools for depth and texture video coding, such as depth-based texture motion vector prediction, depth-range-based weighted prediction, joint inter-view depth filtering, and gradual view refresh. The presented coding scheme is submitted to the 3D video coding (3DV) call for proposals (CfP) of the Moving Picture Experts Group standardization committee. When measured with commonly used objective metrics against the MVC anchor, the proposed scheme provides an average bitrate reduction of 26% and 35% for the 3DV CfP test scenarios with two and three views, respectively. The observed bitrate reduction is similar according to an analysis of the results obtained for the subjective tests on the 3DV CfP submissions.
Miska M. Hannuksela, Dmytro Rusanovskyy, Wenyi Su, Lulu Chen, Ri Li, Payman Aflaki, Deyan Lan, Michal Joachimiak, Houqiang Li, Moncef Gabbouj
IEEE Trans. Image Process.1
2012 Joint view filtering for multiview depth map sequences
abstract
Multi-view depth data representing the real-world 3D scenery can be inconsistent in the inter-view direction. Such inconsistency of depth data can distort the performance of the inter-view motion-compensated prediction (MCP), if the Multiview Video Coding extension of the Advanced Video Coding standard (H.264/MVC) is used for coding of multiview depth map data. In this paper we propose novel joint-view depth filtering (JVDF), which improves inter-view consistency of the multiview depth map data. Depth map images from all available views are warped to the same viewpoint, which results in multiple estimates of a real-world depth value. Estimates that belong to a confidence interval are fused together through a weighted average to produce a “noise-free” depth value estimate. In the provided simulation results the JVDF applied to 3-view depth map data significantly improves the performance of H.264/MVC coding, providing up to 13% of bitrate reduction.
Ri Li, Dmytro Rusanovskyy, Miska M. Hannuksela, Houqiang Li
ICIP3
2012 Texture denoising utilized in depth-enhanced multiview video coding
abstract
In this study, we applied a locally adaptive filtering in 3D DCT domain to 3 views as a preprocessing stage for encoding. Afterwards, these views were encoded and the decoded views were used by a Depth-Image-Based Rendering algorithm (DIBR) to produce virtual intermediate views. A stereopair at a suitable separation for viewing on a stereoscopic display was selected among the synthesized views. A large-scale subjective assessment of the selected synthesized stereopair was performed. A bitrate reduction of 9% on average and up to 21.8% was achieved with almost no penalties on subjective perceived quality.
Payman Aflaki, Dmytro Rusanovskyy, Timo Utriainen, Emilia Pesonen, Miska M. Hannuksela, Satu Jumisko-Pyykkö, Moncef Gabbouj
PCS5
2012 Gradual view refresh in depth-enhanced multiview video
abstract
Depth-enhanced multiview video, such as the multiview video plus depth (MVD) format, can be used to provide displaying-time view adjustment capability through depth-image-based rendering (DIBR) and additional compression improvement compared to the Multiview Video Coding standard. In this paper a gradual view refresh (GVR) method is presented to code random access points and provide fast startup in streaming for MVD bitstreams. When decoding is started from a GVR point, a subset of the views can be accurately decoded, while the remaining views can be approximately reconstructed using DIBR. Perfect reconstruction of all views can be reached at a subsequent random access point. The GVR coding was found to be effective with up to 10% bitrate reduction for sequences with static camera arrangement. The use of GVR for fast startup in video streaming was found to be clearly superior to transmitting a MVD bitstream conventionally from rate-distortion point of view. It has been found in earlier studies that there seems to be a delay from stimulus onset until depth is fully perceived, hence giving a reason to believe that accurate reconstruction of all views might not be necessary immediately after starting decoding. Furthermore, it was observed in this paper that the objective picture quality reduction during GVR was only moderate, verifying the applicability of the presented GVR method.
Miska M. Hannuksela, Lulu Chen, Dmytro Rusanovskyy, Houqiang Li
PCS1
2012 Depth-based motion vector prediction in 3D video coding
abstract
There are several data formats available for 3D video, among which is the Multiview Video plus Depth (MVD) representation that enables depth-image-based rendering (DIBR). In addition to DIBR, the MVD data format can enable more efficient compression for texture video coding. In this paper, we propose advanced motion vector prediction (MVP) techniques which utilize the availability of depth information in the MVD format for more efficient coding of the corresponding texture pictures. The proposed MVP was implemented on top of the Multiview Video Coding (MVC) extension of the Advanced Video Coding (H.264/AVC) standard and its reference software and tested under the simulation conditions of the Call for Proposals of 3D video coding technology issued by the Moving Picture Experts Group (MPEG). In the provided simulation results the proposed scheme significantly outperformed the conventional MVP (by up to 9.9% and on average by 7.5% in Bjontegaard delta bitrate) when it was applied for coding of 3-view MVD data.
Wenyi Su, Dmytro Rusanovskyy, Miska M. Hannuksela, Houqiang Li
PCS3
2012 Intra coding for depth maps using adaptive boundary location
abstract
Depth maps, an essential part in the new generation of 3D video coding, allow rendering of arbitrary viewpoints of a video scene. Depth maps are characterized by sharp object boundaries, which significantly affect the rendering quality and account for the most bitrate for depth map coding. This paper proposes a novel intra coding method for depth maps based on a two-step adaptive boundary location process. By extracting a series of sub-blocks along a depth boundary and refining the boundary within sub-blocks, accurate predictions for blocks with arbitrary edge shapes can be realized. Experimental results show that the proposed scheme achieves bitrate reductions of up to 28% and 13% on average for seven test sequences of MPEG 3DV compared to original intra coding of H.264/AVC considering the same quality of synthesized views. Besides, subjective quality of virtual views is improved owning to well preserved boundary information.
Lulu Chen, Miska M. Hannuksela, Houqiang Li
VCIP2
2012 Rate adaptation for dynamic adaptive streaming over HTTP in content distribution network
Imed Bouazizi, Miska M. Hannuksela, Moncef Gabbouj
Signal Process. Image Commun.3
2012 System Layer Integration of High Efficiency Video Coding
abstract
This paper describes the integration of High Efficiency Video Coding (HEVC) into end-to-end multimedia systems, formats, and protocols such as Real-time transport Protocol, the transport stream of the MPEG-2 standard suite, and dynamic adaptive streaming over the Hypertext Transport Protocol. This paper gives a brief overview of the high-level syntax of HEVC and the relation to the Advanced Video Coding standard (H.264/AVC). A section on HEVC error resilience concludes the HEVC overview. Furthermore, this paper describes applications of video transport and delivery such as broadcast, television over the Internet Protocol, Internet streaming, video conversation, and storage as provided by the different system layers.
Thomas Schierl, Miska M. Hannuksela, Ye-Kui Wang, Stephan Wenger
IEEE Trans. Circuits Syst. Video Technol.2
2012 Overview of HEVC High-Level Syntax and Reference Picture Management
abstract
The increasing proportion of video traffic in telecommunication networks puts an emphasis on efficient video compression technology. High Efficiency Video Coding (HEVC) is the forthcoming video coding standard that provides substantial bit rate reductions compared to its predecessors. In the HEVC standardization process, technologies such as picture partitioning, reference picture management, and parameter sets are categorized as “high-level syntax.” The design of the high-level syntax impacts the interface to systems and error resilience, and provides new functionalities. This paper presents an overview of the HEVC high-level syntax, including network abstraction layer unit headers, parameter sets, picture partitioning schemes, reference picture management, and supplemental enhancement information messages.
Rickard Sjöberg, Ying Chen 0011, Akira Fujibayashi, Miska M. Hannuksela, Jonatan Samuelsson, Thiow Keng Tan, Ye-Kui Wang, Stephan Wenger
IEEE Trans. Circuits Syst. Video Technol.4
2010 Inter-view-predicted redundant pictures for viewpoint switching in multiview video streaming
abstract
Interactive selection of the desired viewpoint is one important feature in multiview video streaming. When receivers differ in the number of simultaneously displayed views, it is challenging to optimize the multiview coding structure particularly when it comes to the use of inter-view prediction and anchor picture frequency. This paper presents a coding and transmission scheme for streaming to clients of different display capabilities. The scheme saves transmission bandwidth by sending only those views that are displayed in a receiver. Furthermore, the scheme enables viewpoint switching through inter-view-predicted redundant pictures representing different sets of views being switched from and switched to. When switching happens at a switch point, an appropriate redundant picture is adaptively chosen based on which views were transmitted before and will be transmitted after the switch point. Experimental results show that the proposed scheme improves the rate-distortion (RD) performance compared to simple anchor picture insertion.
Miska M. Hannuksela, Houqiang Li
ICASSP2
2010 Subjective study on compressed asymmetric stereoscopic video
abstract
Asymmetric stereoscopic video coding takes advantage of the binocular suppression of the human vision by representing one of the views with a lower quality. This paper describes a subjective quality test with asymmetric stereoscopic video. Different options for achieving compressed mixed-quality and mixed-resolution asymmetric stereo video were studied and compared to symmetric stereo video. The bitstreams for different coding arrangements were simulcast-coded according to the Advanced Video Coding (H.264/AVC) standard. The results showed that in most cases, resolution-asymmetric stereo video with the downsampling ratio of 1/2 along both coordinate axes provided similar quality as symmetric and quality-asymmetric full-resolution stereo video. These results were achieved under same bitrate constrain while the processing complexity decreased considerably. Moreover, in all test cases, the symmetric and mixed-quality full-resolution stereoscopic video bitstreams resulted in a similar quality at the same bitrates.
Payman Aflaki, Miska M. Hannuksela, Jukka Häkkinen, Paul Lindroos, Moncef Gabbouj
ICIP2
2010 Congestion-aware transmission rate control using Medium Grain Scalability of Scalable Video Coding
abstract
In packet-oriented networks, packet losses occur mainly due to queue overflows in congested network elements. An increase of the packet rate therefore raises the likelihood of congestion and expected number of lost packets. However, an increase of the packet size has typically a negligible impact on the packet loss rate or congestion in wired packet-switched networks. Scalable Video Coding (SVC), an extension of the Advanced Video Coding (H.264/AVC), provides different types of scalability, one of which is Medium Grain Quality Scalability (MGS). When MGS is in use, layer representations can be pruned unevenly without affecting the decoding of the remaining bitstream. The article proposes a packetization algorithm to improve the quality of the reconstructed video stream while the expected packet loss rate remains unchanged compared to conventional packetization. The algorithm is based on appending MGS enhancement layer data into conventionally generated packet payloads until the Maximum Transmission Unit (MTU) is reached. The simulation results show that the proposed algorithm provides 0.3 to 0.5 dB gain in average luma Peak Signal-to-Noise Ratio when compared with a conventional packetization method while the packet rate remains unchanged.
Miska M. Hannuksela, Haibo Zhu, Houqiang Li, Moncef Gabbouj
ICIP1
2010 Joint multiview video plus depth coding
abstract
Multiview video plus depth (MVD), where per-pixel depth map sequences are associated with multiview texture video, is a promising approach for three-dimensional (3D) video solutions requiring view synthesis. The Multiview Video Coding (MVC) standard has been typically used to compress the MVD representation, resulting into the multiview texture video and the multiview depth video being in two separate MVC bitstreams. Such a coding scheme does not, however, utilize the similarities of the motion information of the multiview texture video and the multiview depth video. The joint multiview video plus depth coding (JMVDC) scheme presented in this paper uses inter-view prediction similarly to MVC as well as the inter-layer motion prediction tool of the Scalable Video Coding (SVC) standard to exploit the correlation of the motion in the texture and depth video sequences. The simulation results show that the JMVDC method achieves about 10 to 20% saving in depth bitrate compared with conventional MVC-based coding of MVD representations.
Miska M. Hannuksela, Houqiang Li
ICIP2
2010 Depth-level-adaptive view synthesis for 3D video
abstract
In the multiview video plus depth (MVD) representation for 3D video, a depth map sequence is coded for each view. In the decoding end, a view synthesis algorithm is used to generate virtual views from depth map sequences. Many of the known view synthesis algorithms introduce rendering artifacts especially at object boundaries. In this paper, a depth-level-adaptive view synthesis algorithm is presented to reduce the amount of artifacts and to improve the quality of the synthesized images. The proposed algorithm introduces awareness of the depth level so that no pixel value in the synthesized image is derived from pixels of more than one depth level. Improvements on objective quality of the synthesized views were achieved in five out of eight test cases, while the subjective quality of the proposed method was similar to or better than that of the view synthesis method used by Moving Picture Experts Group (MPEG).
Ying Chen 0011, Weixing Wan, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
ICME3
2010 Time-variable camera separation for compression of stereoscopic video
abstract
This paper presents a hypothesis that stereoscopic perception requires a short adjustment period after a scene change before it is fully effective. A compression method based on this hypothesis is proposed - instead of coding pictures from the left and right views conventionally, a view in the middle of the left and right view is coded for a limited period after a scene change. The coded middle view can be utilized in two alternative ways in rendering. First, it can be rendered as such, which causes an abrupt change from conventional monoscopic video to stereoscopic video. Second, the layered depth video (LDV) coding scheme can be used to associate depth, background texture, and background depth to the middle view, enabling view synthesis and gradual view disparity increase in rendering. Subjective experiments were conducted to evaluate and validate the presented hypothesis and compare the two rendering methods. The results indicate that when the maximum disparity between the left and right views was relatively small, the presented time-variable camera separation method was imperceptible. A compression gain, the magnitude of which depended on the scene duration, was achieved with half of the sequences having a suitable disparity for the presented coding method.
Maosheng Ji, Miska M. Hannuksela, Moncef Gabbouj, Houqiang Li
VCIP2
2010 Perceptual-based quality assessment for audio-visual services: A survey
Junyong You, Ulrich Reiter, Miska M. Hannuksela, Moncef Gabbouj, Andrew Perkis
Signal Process. Image Commun.3
2010 Multiple Description Video Coding With H.264/AVC Redundant Pictures
abstract
Multiple description coding offers interesting solutions for error resilient multimedia communications as well as for distributed streaming applications. In this letter, we propose a scheme based on H.264/AVC for encoding of image sequences into multiple descriptions. The pictures are split into multiple coding threads. Redundant pictures are inserted periodically in order to increase the resilience to loss and to reduce the error propagation. They are produced with different reference frames than the corresponding primary pictures. We show, given the channel conditions, how to optimally allocate the rates to primary and redundant pictures, such that the total distortion at the receiver is minimized. Extensive experiments demonstrate that the proposed scheme outperforms baseline solutions based on loss and content-adaptive intra coding. Finally, we show how to further reduce the distortion by efficient combination of primary and redundant pictures, if both are available at the decoder.
Ivana Radulovic, Pascal Frossard, Ye-Kui Wang, Miska M. Hannuksela, Antti Hallapuro
IEEE Trans. Circuits Syst. Video Technol.4
2009 An objective video quality metric based on spatiotemporal distortion
abstract
This paper proposes an objective video quality metric based on an analysis of spatial and temporal distortions. Spatial quality features extracted from the spatiotemporal region of reference and distorted videos are used to express the spatial distortion. Temporal distortion, caused by frame freezing resulting from a packet loss, is derived from the spatial distortion before and after the frozen frames. The overall quality is predicted according to the weighted combination of qualities over all the temporal regions. The experimental results with respect to the subjective measurements demonstrate the fast computation and promising performance of the proposed model compared with existing methods.
Junyong You, Miska M. Hannuksela, Moncef Gabbouj
ICIP2
2009 Regionally Adaptive Filtering for Asymmetric Stereoscopic Video Coding
abstract
In asymmetric stereoscopic video coding, one view can be coded in a lower resolution of the other. In this scenario, stereoscopic video can be compressed with only moderately increased bandwidth and complexity compared to 2D monoview video coding. The subjective quality degradation of this scenario can be negligible compared to coding two views with original resolution. The low-resolution view can be predicted from the high-resolution view to achieve higher coding efficiency. In this paper, a regionally adaptive filtering algorithm is proposed to generate a predictor of a macroblock (MB) or MB partition of the low-resolution view from the high-resolution view. Different filters are applied for different picture regions. Disparity motion matching and clustering are applied in the encoder for generation of regionally adaptive filters. Simulation results show that the proposed algorithm results in up to 27% bit-rate saving compared with methods without adaptive filtering.
Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela
ISCAS4
2009 Joint Texture and Depth Map Video Coding based on the Scalable Extension of H.264/AVC
abstract
Depth-Image-Based Rendering (DIBR) is widely used for view synthesis in 3D video applications. Compared with traditional 2D video applications, both the texture video and its associated depth map are required for transmission in a communication system that supports DIBR. To efficiently utilize limited bandwidth, coding algorithms, e.g. the Advanced Video Coding (H.264/AVC) standard, can be adopted to compress the depth map using the 4:0:0 chroma sampling format. However, when the correlation between texture video and depth map is exploited, the compression efficiency may be improved compared with encoding them independently using H.264/AVC. A new encoder algorithm which employs Scalable Video Coding (SVC), the scalable extension of H.264/AVC, to compress the texture video and its associated depth map is proposed in this paper. Experimental results show that the proposed algorithm can provide up to 0.97 dB gain for the coded depth maps, compared with the simulcast scheme, wherein texture video and depth map are coded independently by H.264/AVC.
Siping Tao, Ying Chen 0011, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj, Houqiang Li
ISCAS3
2009 Perceptual quality assessment based on visual attention analysis
abstract
Most existing quality metrics do not take the human attention analysis into account. Attention to particular objects or regions is an important attribute of human vision and perception system in measuring perceived image and video qualities. This paper presents an approach for extracting visual attention regions based on a combination of a bottom-up saliency model and semantic image analysis. The use of PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural SIMilarity) in extracted attention regions is analyzed for image/video quality assessment, and a novel quality metric is proposed which can exploit the attributes of visual attention information adequately. The experimental results with respect to the subjective measurement demonstrate that the proposed metric outperforms the current methods.
Junyong You, Andrew Perkis, Miska M. Hannuksela, Moncef Gabbouj
ACM Multimedia3
2009 Coding techniques in Multiview Video Coding and Joint Multiview Video Model
abstract
Since early 2006, Joint Video Team has been devoting on the development of Multiview Video Coding (MVC) standard as an extension of H.264/AVC. This MVC standard has been finalized in 2008. During the standardization of MVC, there was also a project namely Joint Multiview Video Model (JMVM), which focused on the advanced coding tools that are potentially useful. Those coding tools adopted into JMVM, including illumination compensation and motion skip, have not been added into MVC specification. In this paper, coding techniques in MVC as well as the tools in JMVM are described and discussed, focusing on the coding efficiency.
Ying Chen 0011, Miska M. Hannuksela, Antti Hallapuro, Moncef Gabbouj, Houqiang Li
PCS2
2009 Efficient hierarchical inter picture coding for H.264/AVC baseline profile
abstract
Bi-predictive (B) slices are not supported in the Baseline profile of the Advanced Video Coding (H.264/AVC) standard, which results in a decreased coding efficiency compared with other profiles supporting B slices. However, many application standards, such as the mobile multimedia services specified by the Third Generation Partnership Project (3GPP), use only the Baseline profile for H.264/AVC. Therefore, it is worth investigating H.264/AVC coding when only intra (I) and inter (P) slices are supported. In this paper, a content-adaptive Quantization Parameter (QP) cascading scheme for the hierarchical P coding method compatible with Baseline profile of H.264/AVC is proposed. The proposed method is based on a picture-level QP optimization. The proposed method has a significantly better rate-distortion performance than the traditional IPPP coding structure and outperforms hierarchical P coding methods using fixed delta QP settings between temporal levels noticeably with up to 0.53 dB gain in average luminance Peak Signal-to-Noise Ratio (PSNR).
Weixing Wan, Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
PCS4
2009 Error Resilient Coding and Error Concealment in Scalable Video Coding
abstract
Scalable video coding (SVC), which is the scalable extension of the H.264/AVC standard, was developed by the Joint Video Team (JVT) of ISO/IEC MPEG (Moving Picture Experts Group) and ITU-T VCEG (Video Coding Experts Group). SVC is designed to provide adaptation capability for heterogeneous network structures and different receiving devices with the help of temporal, spatial, and quality scalabilities. It is challenging to achieve graceful quality degradation in an error-prone environment, since channel errors can drastically deteriorate the quality of the video. Error resilient coding and error concealment techniques have been introduced into SVC to reduce the quality degradation impact of transmission errors. Some of the techniques are inherited from or applicable also to H.264/AVC, while some of them take advantage of the SVC coding structure and coding tools. In this paper, the error resilient coding and error concealment tools in SVC are first reviewed. Then, several important tools such as loss-aware rate-distortion optimized macroblock mode decision algorithm and error concealment methods in SVC are discussed and experimental results are provided to show the benefits from them. The results demonstrate that PSNR gains can be achieved for the conventional inter prediction (IPPP) coding structure or the hierarchical bi-predictive (B) picture coding structure with large group of pictures size, for all the tested sequences and under various combinations of packet loss rates, compared with the basic joint scalable video model (JSVM) design applying no error resilient tools at the encoder and only picture copy error concealment method at the decoder.
Ying Chen 0011, Ye-Kui Wang, Houqiang Li, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.5
2009 Error Resilient Video Coding Using Redundant Pictures
abstract
This paper presents several error resilient video coding methods based on redundant pictures. We combine redundant picture coding with reference picture selection and reference picture list reordering to prevent error propagation in motion compensated video coding. A hierarchical redundant picture allocation method is employed to make a tradeoff between error resilience and coding efficiency. For improved end-to-end rate-distortion performance in packet loss environment, three adaptive redundant picture allocation methods are further developed, utilizing characteristics of the input video content. Simulation results show that the adaptive redundant picture coding methods can achieve average luma peak-signal-to-noise improvements up to 3.5 dB compared to the loss-aware rate distortion optimized intra macroblock refresh algorithm implemented in the H.264/AVC Joint Model (JM). The proposed redundant picture coding methods are standard-compliant and do not introduce any additional end-to-end delay, therefore suit for low-delay applications such as video telephony and video conferencing. Due to the good error resilience performance, some of the proposed redundant picture coding methods have been adopted and integrated into the JM.
Chunbo Zhu, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li
IEEE Trans. Circuits Syst. Video Technol.3
2008 Picture-level adaptive filter for asymmetric stereoscopic video
abstract
In asymmetric stereoscopic video coding, one view is coded in a quarter of the resolution of the other and the low- resolution view is predicted from the high-resolution view. This way, stereoscopic video effect could be achieved with only moderately increased bandwidth and complexity. Inter-view prediction tools for generating the predictor of a maroblock (MB) or MB partition in the low-resolution view from the high-resolution view play a vital role for coding efficiency in asymmetric video coding. In this paper, we propose a method that applies an adaptive filter to generate picture-level adaptive inter-view predictors for MBs or MB partitions. At the encoder, a low complexity preprocessing module is built to find out the filters. Simulation results show that the proposed method provides a bit-rate saving of 26% at maximum and 5% on average.
Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj
ICIP3
2008 Low-complexity asymmetric multiview video coding
abstract
Multiview video coding (MVC) is currently under development by the Joint Video Team (JVT) as an extension to Advanced Video Coding (H264/AVC). Based on the suppression theory in binocular vision, the fidelity of one of the two views of a stereoscopic display can be reduced without noticeable degradation of subjective quality. Thus, in MVC, a subset of views can be coded with lower spatial resolution at negligible cost to subjective quality. Due to different resolutions, a downsampling process is required in an MVC decoder in order to enable motion compensation (MC) between views. In this paper, a low-complexity MC algorithm is proposed for MVC to enable inter-view prediction between pictures with different resolutions. It requires lower memory consumption and lower computational complexity compared with the conventional downsampled inter-view prediction, while providing comparable efficiency, as shown by the simulation results.
Ying Chen 0011, Shujie Liu 0001, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
ICME4
2008 Single-loop decoding for multiview video coding
abstract
Multiview video coding (MVC) is currently being standardized by the Joint Video Team as an extension of H264/AVC. When an MVC bitstream is decoded, some views (named target views) are to be displayed; some other views (named dependent views) may not be displayed but are needed for inter-view prediction of the target views. The original MVC design requires pictures of the dependent views to be fully decoded and stored. This entails both high decoding complexity and high memory consumption for the pictures in the views which are not intended for display, particularly when the number of dependent views is large. In this paper, a single-loop decoding (SLD) scheme is introduced to address these disadvantages. SLD requires only partial decoding of pictures in dependent views and thus significantly reduces decoding complexity and memory consumption. The proposed method is based on the so-called motion skip, wherein inter-view motion and coding mode prediction is exploited. Experimental results show that compared to coding schemes that require comparable complexity, significant compression gain can be achieved. For example, 25% bit-rate saving on average can be obtained compared to simulcast. Simulation results also show that the proposed SLD scheme provides a substantial reduction of complexity and memory size, at the expense of only a minor compression efficiency loss, compared with multiple-loop decoding MVC schemes.
Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj
ICME3
2008 Frame loss error concealment for multiview video coding
abstract
The Multiview Video Coding (MVC) standard is currently under development by the Joint Video Team as an extension of the Advanced Video Coding (H.264/AVC) standard. An MVC encoder compresses more than one viewpoint of a scene captured by different cameras. Redundancies between views can be used for inter-view prediction in encoding as well as error concealment in decoding. In this paper, a new algorithm utilizing motion information of pictures from other views to conceal a lost picture is proposed. The algorithm first derives motion information for a lost picture based on motion fields of pictures in adjacent views. Then, traditional motion compensation is invoked within the view containing the lost picture to derive a concealed frame. Experimental results show that the proposed algorithm can improve video quality with a negligible computational complexity overhead compared to simple temporal error concealment algorithms.
Shujie Liu 0001, Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela, Houqiang Li
ISCAS5
2008 Does context matter in quality evaluation of mobile television?
abstract
Subjective quality evaluation is used to optimize the produced audiovisual quality from fundamental signal processing algorithms to consumer services. These studies typically follow the basic principles of controlled psychoperceptual experiments. However, when compromising compression and transmission parameters for consumer services, the ecological validity of conventional quality evaluation methods can be questioned. To tackle this, we firstly present a novel user-oriented quality evaluation method for mobile television in its usage contexts. Secondly, we present the results of an experiment conducted with 30 participants comparing acceptability and satisfaction of quality as well as goals of viewing in three mobile contexts and under four different residual transmission error rates, when the participants also performed simultaneous assessment tasks. Finally, we compare the results with a previous laboratory experiment. The studied error rates impacted negatively on all measured tasks with some contextual differences. Moreover, the evaluations were more favorable and less discriminate in the mobile contexts compared to the laboratory.
Satu Jumisko-Pyykkö, Miska M. Hannuksela
Mobile HCI2
2008 Semi-Fuzzy Rate Controller for Variable Bit Rate Video
abstract
A novel semi-fuzzy (SF) rate control algorithm (RCA) for variable bit rate (VBR) video applications is proposed. The proposed RCA is optimized to provide high quality compressed video bit streams in a wide operating range from constant quality to nearly constant bit rate. Thanks to a low degree of computational complexity, it is suitable for real-time applications of VBR video. The proposed RCA operates under given buffer size, delay and quality constraints. It provides a VBR video bit stream by controlling the quantization parameter (QP) on a picture basis. The QP is mainly controlled by a fuzzy rate controller and a deterministic quality controller, which are optimized such that they minimize the variation of quality to provide encoded video with high and stable visual quality. The proposed RCA has been implemented in an H.264/AVC video codec and the experimental results show that it provides a high-level average quality for encoded video while strictly obeying the buffering delay and quality constraints.
Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.2
2007 System and Transport Interface of SVC
abstract
Scalable video coding (SVC) and transmission has been a research topic for many years. Among other objectives, it aims to support different receiving devices, perhaps connected through a heterogeneous network structure, using a single bit stream. Earlier attempts of standardized scalable video coding, for example in MPEG-2, H.263, or MPEG-4 Visual, have not been commercially successful. Nevertheless, the Joint Video Team has recently focused on the development of the scalable video extensions of H.264/AVC, known as SVC. Some of the key problems of older scalable compression techniques have been solved in SVC and, at the same time, new and compelling use cases for SVC have been identified. While it is certainly important to develop coding tools targeted at high coding efficiency, the design of the features of the interface between the core coding technologies and the system and transport are also of vital importance for the success of SVC. Only through this interface, and novel mechanisms defined therein, applications can take advantage of the scalability features of the coded video signal. This paper provides an overview of the system interface features defined in the SVC specification. We discuss, amongst other features, bit stream structure, extended network abstraction layer (NAL) unit header, and supplemental enhancement information (SEI) messages related to scalability information.
Ye-Kui Wang, Miska M. Hannuksela, Stéphane Pateux, Alexandros Eleftheriadis, Stephan Wenger
IEEE Trans. Circuits Syst. Video Technol.2
2006 Low-Complexity Fuzzy Video Rate Controller for Streaming
abstract
In this paper we propose a low-complexity fuzzy video rate control algorithm with buffer constraint designed for real-time streaming applications. While in low delay video communications bit streams with constant bitrate are required, in streaming application more delay and variation in bitrate is acceptable. The described video rate control algorithm (RCA) provides a variable bitrate video by control of the quantization scale (QS) on picture basis. The QS is mainly controlled by a fuzzy controller such that it minimizes the variation of QS to provide encoded video with high visual quality so as to utilize the variable bitrate benefits as much as possible. The proposed rate control algorithm (RCA) has been implemented in the MPEG-4, H.263 and H.264/AVC standard video codecs and the experimental results show that it provides high level average quality for encoded video while it strictly obeys streaming constraints
Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj
ICASSP (2)2
2006 Fuzzy Rate Controller for Variable Bitrate Video in Mobile Applications
abstract
In this paper we propose a low-complexity fuzzy video rate controller designed for real-time variable bitrate applications with buffer constraints. The algorithm is optimized for streaming application in mobile devices. Furthermore, the proposed algorithm can be used for local recording application so as the recorded video may be streamed in future. Today, many mobile phones include a digital camera that can be used to capture and encode video in real-time. We assume that no memory for storage of uncompressed video is available in mobile phones. Therefore, look-ahead and multi-pass rate controls are not possible. Furthermore, considering the processing power and, more importantly, battery life constraints in mobile devices, the proposed algorithm needs to be as simple as possible. The described variable bitrate (VBR) bit rate control algorithm controls the quantization scale (QS) on picture basis. The QS is mainly controlled by a fuzzy controller such that it minimizes the variation of QS to provide encoded video with high visual quality so as to utilize the variable bitrate benefits as much as possible. The proposed rate control algorithm (RCA) has been implemented in the MPEG-4, H.263 and H.264/AVC standard video codecs and the experimental results show that it provides high level average quality for encoded video while it strictly obeys buffering constraints.
Mehdi Rezaei, Alireza Akhbardeh, Miska M. Hannuksela, Moncef Gabbouj
ICC3
2006 Video Splicing and Fuzzy Rate Control in IP Multi-Protocol Encapsulator for Tune-In Time Reduction in IP Datacasting (IPDC) over DVB-H
abstract
A novel video splicing and rate control method is proposed which minimizes the tune-in time in IPDC over DVB-H. DVB-H uses a time-sliced transmission scheme to reduce the power consumption used for radio reception. One of the significant factors in tune-in time is the time from the start of media decoding to the start of correct output from decoding, which would be minimized when a time-slice is started with a random access point picture such as an instantaneous decoding refresh (IDR) picture in H.264/AVC. In IPDC over DVB-H, the encapsulation to time-slices is performed independently of encoding in a network element called IP encapsulator. At the time of encoding, time-slice boundaries are not known exactly, and it is impossible to govern the location of IDR pictures relative to time-slices. It is proposed that an additional stream consisting of IDR pictures only is transmitted to the IP encapsulator, which replaces pictures in a normal bitstream with IDR pictures according to time-slice boundaries in order to achieve the minimum tune-in time. It has to be ensured that the "spliced" bit stream resulting from the operation of the IP encapsulator complies with the hypothetical reference decoder (HRD) specification of H.264/AVC. A video rate control system utilizing a fuzzy controller is proposed to satisfy the HRD requirements for the spliced bit stream. Simulation results show that the proposed splicing method and rate control system can provide standard bit streams with good average quality of decoded video and with minimum tune-in time.
Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj
ICIP2
2006 System and Transport Interface of H.264/AVC Scalable Extension
abstract
The scalable extension of H.264/AVC, known as scalable video coding or SVC, has recently been the main focus of the Joint Video Team. The higher level syntax of SVC follows the design principles of H.264/AVC. This allows the work towards an optimized file format and RTF payload format to be conducted in parallel with the core SVC specification. This paper provides an overview of the system and transport interface design of SVC, including the SVC high-level syntax and functionalities, SVC file format, and SVC RTF payload format.
Ye-Kui Wang, Stephan Wenger, Miska M. Hannuksela
ICIP3
2006 Error Resilient Video Coding using Redundant Pictures
abstract
Coding of redundant pictures is supported in the latest international video coding standard H.264 (also known as MPEG-4 part 10 or AVC). This paper proposes a standard-compliant way to encode and decode redundant pictures for improved error resilience. The method is based on a combination of picture-level reference picture selection, reference picture list ordering, and a hierarchical allocation of redundant pictures, which can efficiently prevent temporal error propagation without relying on feedback information. Simulation results show that the method outperforms the optimal loss-aware rate-distortion optimized intra refresh method. The proposed algorithm has been adopted into the H.264 joint model.
Chunbo Zhu, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li
ICIP3
2006 Video Encoding and Splicing for Tune-in Time Reduction in IP Datacasting (IPDC) Over DVB-H
abstract
A novel video encoding and splicing method is proposed which minimizes the tune-in time of "channel zapping", i.e. changing from one audiovisual service to another, in IPDC over digital video broadcasting for handheld terminals (DVB-H). DVB-H uses a time-sliced transmission scheme to reduce the power consumption used for radio reception. Tune-in time in DVB-H refers to the time between the start of the reception of a broadcast signal and the start of the media rendering. One of the significant factors in tune-in time is the time from the start of media decoding to the start of correct output from decoding, which is minimized when a time-slice is started with a random access point picture such as an independent decoding refresh (IDR) picture in H.264/AVC. In IPDC over DVB-H, encapsulation to time-slices is performed independently from encoding in a network element called IP encapsulator. At the time of encoding, time-slice boundaries are not known exactly, and it is therefore impossible to govern the location of IDR pictures relative to time-slices. It is proposed that an additional stream consisting of IDR pictures only is transmitted to the IP encapsulator, which replaces pictures in a normal bitstream with IDR pictures according to time-slice boundaries in order to achieve the minimum tune-in time. It has to be ensured that the "spliced" stream resulting from the operation of the IP encapsulator complies with the hypothetical reference decoder (HRD) specification of H.264/AVC. A video encoding and rate control system is proposed to satisfy the HRD requirements for the spliced stream. Simulation results show that in addition to fulfilling HRD compliancy, good average quality of decoded video is achieved with minimum tune-in time
Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj
ICME2
2006 Spliced Video and Buffering Considerations for Tune-In Time Minimization in DVB-H for Mobile TV
abstract
A novel video splicing method is proposed which minimizes the tune-in time of mobile TV in Digital Video Broadcasting for Handheld terminals (DVB-H). DVB-H uses a time-sliced transmission scheme to reduce the power consumption used for radio reception. Tune-in time in DVB-H refers to the time between the start of the reception of a broadcast signal and the start of the media rendering. One of the significant factors in tune-in time is the time from the start of media decoding to the start of correct output from decoder, which can be minimized when a time-slice is started with a random access point picture such as an independent decoding refresh (IDR) picture in H.264/AVC. In IP datacasting (IPDC) over DVB-H, the encapsulation to time-slices is performed independently from encoding in a network element called IP encapsulator. At the time encoding, time-slice boundaries are not known exactly, and it is impossible to govern the location of IDR pictures relative to time-slice boundaries. It is proposed that an additional stream consisting of IDR pictures only is transmitted to the IP encapsulator, which replaces pictures in a normal bitstream with IDR pictures according to time-slice boundaries in order to achieve the minimum tune-in time. It has to be ensured that the "spliced" stream resulting from the operation of the IP encapsulator complies with the Hypothetical Reference Decoder (HRD) specification of H.264/AVC.
Mehdi Rezaei, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital, Moncef Gabbouj
PIMRC2
2006 Method for Unequal Error Protection in DVB-H for Mobile Television
abstract
This paper introduces a method for unequal error protection (UEP) of media data in a time-sliced DVB-H channel. Media datagrams are assigned priorities using some a-priori knowledge. Datagrams covering a certain period of playback time are first grouped based on the priority assignment. Each group is then protected using Reed-Solomon forward error correction (FEC) codes and packed into multi-protocol encapsulation (MPE) FEC frames as defined by the DVB-H standard. All MPE-FEC frames for a certain period of playback time are then sent back to back without any delay between these MPE-FEC frames. This method of UEP is generic and can be tuned according to the priority assignment algorithm. Simulations using H.264/AVC video were conducted to evaluate the performance of the proposed method. It used a simple priority assignment algorithm. The resulting rate distortion graphs show good performance and an average luma peak signal-to-noise ratio (PSNR) improvement of up to 0.8 dB was achieved
Vinod Kumar Malamal Vadakital, Miska M. Hannuksela, Mehdi Rezaei, Moncef Gabbouj
PIMRC2
2005 On Datacasting of H.264/AVC over DVB-H
abstract
This paper investigates the performance of H. 264/AVC video codec in a DVB-H (Handheld) datacasting environment. DVB-H was designed to provide point-to-multipoint (PTM) broadcast/multicast type transmission to handheld, battery operated devices. Reed-Solomon (RS) forward error correcting (FEC) codes are applied to Multi-Protocol Encapsulation (MPE) section payloads, termed MPE-FEC, to deliver data over DVB-H. The bitrate overheads incurred due to the additional FEC and packetization headers are analyzed. The requirement of additional data protection in the form of MPE-FEC is illustrated with the help of simulation results.
Vinod Kumar Malamal Vadakital, Miska M. Hannuksela, Harri Pekkonen, Moncef Gabbouj
MMSP2
2004 Isolated regions in video coding
abstract
Different types of prediction are applied in modern video coding. While predictive coding improves compression efficiency, the propagation of transmission errors becomes more likely. In addition, predictive coding brings difficulties to other aspects of video coding, including random access, parallel processing, and scalability. In order to combat the negative effects, video coding schemes introduce mechanisms such as slices and intracoding, to limit and break the prediction. This paper proposes the use of the isolated regions coding tool that jointly limits in-picture prediction and interprediction on a region-of-interest basis. The tool can be used to provide random access points from non-intrapictures and to respond to intrapicture update requests. Furthermore, it can be applied as an error-robust macroblock mode decision method and can be used in combination with unequal error protection. Finally, it enables mixing of scenes, which is useful in coding of masked scene transitions.
Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj
IEEE Trans. Multim.1
2003 Random access using isolated regions
abstract
Random access is a desirable feature in many video communication systems. Intra pictures is conventionally used as random access points, but correct picture content is recovered gradually within a range of pictures starting from a non-intra random access point. This paper proposes the use of the isolated regions technique for gradual decoder refresh and presents how the proposed method can be used in the upcoming ITU-T recommendation H.264, also known as MPEG-4 part 10 or advanced video coding. The presented simulations reveal that the proposed method outperforms intra-picture-based random access points in error-prone network conditions. It is also shown that the proposed method is more flexible and suits packet-based transmission better compared to progressively located intra-coded slices.
Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj
ICIP (3)1
2003 H.264/AVC in wireless environments
abstract
Video transmission in wireless environments is a challenging task calling for high-compression efficiency as well as a network friendly design. Both have been major goals of the H.264/AVC standardization effort addressing "conversational" (i.e., video telephony) and "nonconversational" (i.e., storage, broadcast, or streaming) applications. The video compression performance of the H.264/AVC video coding layer typically provides a significant improvement. The network-friendly design goal of H.264/AVC is addressed via the network abstraction layer that has been developed to transport the coded video data over any existing and future networks including wireless systems. The main objective of this paper is to provide an overview over the tools which are likely to be used in wireless environments and discusses the most challenging application, wireless conversational services in greater detail. Appropriate justifications for the application of different tools based on experimental results are presented.
Thomas Stockhammer, Miska M. Hannuksela, Thomas Wiegand 0001
IEEE Trans. Circuits Syst. Video Technol.2
2002 H.26L/JVT coding network abstraction layer and IP-based transport
abstract
The JVT/H.26L video coding scheme conceptually consists of a video coding layer (VCL) responsible mainly for coding efficiency and a network abstraction layer (NAL) that supports video specific transport features for a variety of networks. This paper describes the H.26L/JVT NAL in general, including the parameter set concept and the transport over packet-based and bit-stream oriented networks. Encapsulation of coded video data to RTP/UDP/IP transport is discussed, and applications of H.26L in fixed and wireless IP-based transmission environments are presented.
Thomas Stockhammer, Miska M. Hannuksela, Stephan Wenger
ICIP (2)2
2002 Coding of faded scene transitions
abstract
Coding of a scene transition is often a challenging problem, from the compression efficiency point of view, because motion compensation may not be a powerful enough method to represent changes between pictures in the transition. This paper proposes a overlay coding technique for coding faded scene transitions. As shown by extensive simulations, over 50% bit-rate savings in both cross-fades and through-black fades compared to earlier techniques can be achieved. Overlay coding suits situations where video is edited manually or automatically.
Dong Tian, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj
ICIP (2)2
2002 Sub-picture: ROI coding and unequal error protection
abstract
Region-of-interest coding and unequal error protection are two important tools in video communication systems to improve the received visual quality. One common property of the two techniques is that unequal coding or transmission is applied to improve the quality of the most important parts of images. The proposed sub-picture coding technique facilitates both region-of-interest coding and unequal error protection by partitioning images to regions of interest and separating the corresponding coded data units from each other. Simulation results show that the overall subjective quality is considerably improved compared to the conventional coding schemes.
Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj
ICIP (3)2
2002 The error concealment feature in the H.26L test model
abstract
This paper presents the error concealment (EC) feature implemented by the authors in the test model of the draft ITU-T video coding standard H.26L. The selected EC algorithms are based on weighted pixel value averaging for INTRA. pictures and boundary-matching-based motion vector recovery for INTER pictures. The specific concealment strategy and some special methods, including handling of B-pictures, multiple reference frames and entire frame losses, are described. Both subjective and objective results are given based on simulations under Internet conditions. The feature was adopted and is now included in the latest H.26L reference software TML-9.0.
Ye-Kui Wang, Miska M. Hannuksela, Viktor Varsa, Ari Hourunranta, Moncef Gabbouj
ICIP (2)2