EDBT 2026 Demo / reviewers in the wild / expert
Jani Lainema
dblp:89/6811
· DBLP profile ↗
39ranked-venue papers
4as first author
9since 2021 · last 2025
0009-0001-4873-8499ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 4 first-author · 9 since 2021Systems, architecture and hardware · 3Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Convolutional Cross-Component Models for Chroma Prediction in Video CodingabstractIn this paper we present two novel approaches for improving intra and inter chroma prediction in video coding. Our research demonstrates that treating the cross-component predictor as a two-dimensional convolutional model can significantly enhance chroma prediction performance. The proposed two convolutional models incorporate multiple spatial neighbors, a bias term, and a nonlinear term. For intra-coded blocks, we derive the model coefficients on the reconstructed neighborhood of the block, while for inter-coded blocks, the model coefficients are determined using prediction samples. To evaluate our methods, we implemented them on top of the ECM software that is currently under exploration by the ITU-T/ISO/IEC Joint Video Experts Team. Our intra cross-component predictor achieves BD-rate savings of {−1.47%, −2.90%, −3.02%}, {−0.92%, −2.04%, −2.32%} (Y, U, V) for the all intra and the random access configurations over ECM-5.0, respectively. Our inter cross-component predictor achieves BD-rate savings of {−0.09%, −1.25%, −1.46%}, {−0.04%, −3.42%, −3.85%} for the random access and the low-delay B configurations over ECM-9.0, respectively. Both proposed methods have been adopted into the ECM software. Pekka Astola, Alireza Aminlou, Ramin Ghaznavi Youvalari, Jani Lainema |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Overfitting NN loop-filters in video codingabstractOverfitting is usually regarded as a negative condition since it impairs the generalisation power of a model. Nevertheless, overfitting a Neural Network (NN) on test data may be advantageous to improve the compression efficiency of image/video coding tools and systems. Previous research has demonstrated the benefits of NN overfitting for post-processing operations, i.e. post-filters, but not yet for actual decoding tools. Generally, the NN is overfitted on test data at the encoder end, and the weight update is coded and sent to the decoder end along the image/video bitstream. The proposed approach follows this strategy. In particular, the overfitting of the Low Operation Point (LOP) loop-filter in NN-based Video Coding (NNVC) software is studied. The overall approach yields Bjøntegaard Delta rate (BD-rate) of -7.74%, -13.73% and -12.49%, for the Y, U and V components, respectively. Out of these coding gains, 1.21%, 6.43% and 5.52%, for the Y, U and V components, are attributed to the overfitting. The boost in the coding gains comes with only 1.5% more complexity, due to the multiplier parameters introduced during the overfitting. Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela, Tapio Elomaa |
VCIP | 5 |
| 2022 | Content-Adaptive Neural Network Post-Processing Filter with NNR-Coded Weight-UpdatesabstractNeural Network (NN) filters improve the perceptual quality of reconstructed videos by reducing compression artefacts. For content adaptation, a few NN-filters use over-fitting. As the adaptation signal is a weight-update, compression is required to minimise significant bitrate overheads. Most approaches, however, use generic data compression algorithms, which are inadequate for coding NN weight-updates. This work introduces a content-adaptive NN post-processing filter with weight-updates coded using the Neural Network compression and Representation (NNR) standard. The bitrate overhead is further decreased by over-fitting only a subset of weights, selected via energy-based analysis. The proposed filter saved about 4.57% (Y), 10.33% (Cb), 6.53% (Cr) Bjøntegaard Delta rate (BD-rate) on top of the Versatile Video Coding (VVC) Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0, in Random Access (RA) configuration. Compared to the non-over-fitted NN, the performance was doubled; and compared to 7z, NNR reduced the bitrate of the weight-update by ∼64%. María Santamaría 0001, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela |
ICIP | 3 |
| 2022 | Low-precision post-filtering in video codingabstractNeural Networks (NNs) have demonstrated their effectiveness in tackling challenges involving multimedia content. In the video coding field, NNs are actively exploited as novel tools that complement conventional signal processing tools, as well as end-to-end coding solutions. Since NNs use commonly floating-point arithmetic, different results may be generated in different computing environments, leading to discrepancies and even corrupted reconstructions. Accordingly, this issue is solved by employing fixed-point arithmetic instead. This paper studies the quantisation of a 32-bit floating-point (float32) NN post-filter to 32-bit fixed-point (int32) and 16-bit fixed-point (int16). On top of the VVC Test Model (VTM) 11.0 NN-based Video Coding (NNVC) 1.0, the coding gains of the float32 post-filter are 5.01% (Y), 18.95% (Cb) and 17.33% Cr. Compared to the float32 inference, the quantised models produce coding losses: 0.01% (Y, Cb and Cr) for the int32 inference and 0.45% (Y), 1.97% (Cb) and 1.19% (Cr) for the int16 inference. Nevertheless, the fixed-point approaches achieve bit exact matches in different computing environments. Moreover, the decoding time with int16 is about half the decoding time of int32. Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela |
ISM | 5 |
| 2021 | Content-adaptive convolutional neural network post-processing filterabstractNeural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration. María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj |
ISM | 4 |
| 2021 | Adaptation and Attention for Neural Video CodingabstractNeural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 6 |
| 2021 | Quantization and Entropy Coding in the Versatile Video Coding (VVC) StandardabstractThe paper provides an overview of the quantization and entropy coding methods in the Versatile Video Coding (VVC) standard. Special focus is laid on techniques that improve coding efficiency relative to the methods included in the High Efficiency Video Coding (HEVC) standard: The inclusion of trellis-coded quantization, the advanced context modeling for entropy coding of transform coefficient levels, the arithmetic coding engine with multi-hypothesis probability estimation, and the joint coding of chroma residuals. Beside a description of the design concepts, the paper also discusses motivations and implementation aspects. The effectiveness of the quantization and entropy coding methods specified in VVC is validated by experimental results. Heiko Schwarz, Muhammed Z. Coban, Marta Karczewicz, Tzu-Der Chuang, Frank Bossen, Alexander Alshin, Jani Lainema, Christian R. Helmrich, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Regression-Based Motion Vector Field for Video CodingabstractIn this paper, we study a method for compensating the non-translational motion behavior in video coding. The proposed method models the motion field of a prediction block based on the motion information of the neighboring blocks by using a linear regression approach. In order to provide a finer granularity of motion vectors the Regression-based Motion Vector Field (RMVF) method derives the motion field in 4×4 sub-block accuracy. Such approach generates a smooth and more realistic motion vector field inside the prediction block. The motion field generated with RMVF is then used as a new merge mode along with other merge modes in VTM-2.0 test model of the Versatile Video Coding (H.266/VVC) standard. The conducted experiments with JVET CTC sequences illustrate that the proposed RMVF method provides 0.77%, 0.19% and 0.41% bitrate reductions with random access (RA), low delay B (LDB) and low delay P (LDP) configurations, respectively. Furthermore, this method provides on average 0.65% bitrate saving for the 360° sequences in equirectangular projection format (ERP) with RA configuration. Ramin Ghaznavi Youvalari, Alireza Aminlou, Jani Lainema |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Transform Coding in the VVC StandardabstractIn the past decade, the development of transform coding techniques has achieved significant progress and several advanced transform tools have been adopted in the new generation Versatile Video Coding (VVC) standard. In this paper, a brief history of transform coding development during VVC standardization is presented, and the transform coding tools in the VVC standard are described in detail together with their initial design, incremental improvements and implementation aspects. To improve coding efficiency, four new transform coding techniques are introduced in VVC, which are namely Multiple Transform Selection (MTS), Low-Frequency Non-separable Secondary Transform (LFNST) and Sub-Block Transform (SBT), as well as a large (64-point) type-2 DCT. The experimental results on VVC reference software (VTM-9.0) show that average 4.5% and 3.6% overall coding gain can be achieved by the VVC transform coding tools for All Intra and Random Access configurations, respectively. Xin Zhao 0003, Seung-Hwan Kim 0001, Yin Zhao, Hilmi E. Egilmez, Moonmo Koo, Shan Liu 0001, Jani Lainema, Marta Karczewicz |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | Ccalf Coefficient Derivation Using Combined NeighborsabstractLoop filtering is widely used in recent video coding standards to reduce coding artifacts introduced by prediction and residual coding process. Versatile Video Coding (VVC) standard employs Cross Component Adaptive Loop Filter (CCALF) to improve chroma fidelity. This paper describes CCALF coefficient derivation that uses combined luma neighboring pixels. With the proposed combination, more correlated pixels provide less contribution to CCALF process resulting in improving performance. Simulation results show the proposed method has combined BD-rate gains of -1.51%, -2.01%, -2.40% for all intra (AI), random Fig. 1: Block diagram of VVC showing interactions between access (RA), and low delay (LB) conditions, respectively, CCALF and other loop filtering tools [1] compared to VTM-7.0 reference software. Krit Panusopone, Seungwook Hong, Limin Wang 0009, Jani Lainema |
ICIP | 4 |
| 2020 | Efficient Adaptation of Neural Network Filter for Video CompressionabstractWe present an efficient finetuning methodology for neural-network filters which are applied as a postprocessing artifact-removal step in video coding pipelines. The fine-tuning is performed at encoder side to adapt the neural network to the specific content that is being encoded. In order to maximize the PSNR gain and minimize the bitrate overhead, we propose to finetune only the convolutional layers' biases. The proposed method achieves convergence much faster than conventional finetuning approaches, making it suitable for practical applications. The weight-update can be included into the video bitstream generatedby the existing video codecs. We show that our method achieves up to 9.7% average BD-rate gain when compared to the state-of-art Versatile Video Coding (VVC) standard codec on 7 test sequences. Yat Hong Lam, Alireza Zare, Francesco Cricri, Jani Lainema, Miska M. Hannuksela |
ACM Multimedia | 4 |
| 2020 | Joint Cross-Component Linear Model For Chroma Intra PredictionabstractThe Cross-Component Linear Model (CCLM) is an intra prediction technique that is adopted into the upcoming Versatile Video Coding (VVC) standard. CCLM attempts to reduce the inter-channel correlation by using a linear model. For that, the parameters of the model are calculated based on the reconstructed samples in luma channel as well as neighboring samples of the chroma coding block. In this paper, we propose a new method, called as Joint Cross-Component Linear Model (J-CCLM), in order to improve the prediction efficiency of the tool. The proposed J-CCLM technique predicts the samples of the coding block with a multi-hypothesis approach which consists of combining two intra prediction modes. To that end, the final prediction of the block is achieved by combining the conventional CCLM mode with an angular mode that is derived from the co-located luma block. The conducted experiments in VTM-8.0 test model of VVC illustrated that the proposed method provides on average more than 1.0% BD-Rate gain in chroma channels. Furthermore, the weighted YCbCr bitrate savings of 0.24% and 0.54% are achieved in 4:2:0 and 4:4:4 color formats, respectively. Ramin Ghaznavi Youvalari, Jani Lainema |
MMSP | 2 |
| 2020 | L2C - Learning to Learn to CompressabstractIn this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu |
MMSP | 5 |
| 2020 | General Video Coding Technology in Responses to the Joint Call for Proposals on Video Compression With Capability Beyond HEVCabstractAfter the development of the High-Efficiency Video Coding Standard (HEVC), ITU-T VCEG and ISO/IEC MPEG formed the Joint Video Exploration Team (JVET), which started exploring video coding technology with higher coding efficiency, including development of a Joint Exploration Model (JEM) algorithm and a corresponding software implementation. The technology explored in the last version of the JEM further increases the compression capabilities of the hybrid video coding approach by adding new tools, reaching up to 30% bit rate reduction compared to HEVC based on the Bjøntegaard delta bit rate (BD-rate) metric, and further improvement beyond that in terms of subjective visual quality. This provided enough evidence to issue a joint Call for Proposals (CfP) for a new standardization activity now known as Versatile Video Coding (VVC). All technology proposed in the responses to the CfP was based on the classic block-based hybrid video coding design, extending it by new elements of partitioning, intra- and inter-picture prediction, prediction signal filtering, transforms, quantization/scaling, entropy coding, and in-loop filtering. This article provides an overview of technology that was proposed in the responses to the CfP, with a focus on techniques that were not already explored in the JEM context. Benjamin Bross, Kenneth Andersson, Max Bläser, Virginie Drugeon, Seung-Hwan Kim 0001, Jani Lainema, Shan Liu 0001, Jens-Rainer Ohm, Gary J. Sullivan, Ruoyang Yu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Wide Angular Intra Prediction for Versatile Video CodingabstractThis paper presents a technical overview of Wide Angular Intra Prediction (WAIP) that was adopted into the test model of Versatile Video Coding (VVC) standard. Due to the adoption of flexible block partitioning using binary and ternary splits, a Coding Unit (CU) can have either a square or a rectangular block shape. However, the conventional angular intra prediction directions, ranging from 45 degrees to -135 degrees in clockwise direction, were designed for square CUs. To better optimize the intra prediction for rectangular blocks, WAIP modes were proposed to enable intra prediction directions beyond the range of conventional intra prediction directions. For different aspect ratios of rectangular block shapes, different number of conventional angular intra prediction modes were replaced by WAIP modes. The replaced intra prediction modes are signaled using the original signaling method. Simulation results reportedly show that, with almost no impact on the run-time, on average 0.31% BD-rate reduction is achieved for intra coding using VVC test model (VTM). Xin Zhao 0003, Shan Liu 0001, Xiang Li 0003, Jani Lainema, Gagan Rath, Fabrice Urban, Fabien Racapé |
DCC | 5 |
| 2019 | Approximating Binarization in Neural NetworksabstractBinarization of neural networks' activations may be a requirement for some applications. A typical example is end-to-end learned deep image compression systems where the encoder's output is requred to be a binary vector. Binarization is non-differentiable, therefore one needs to approximate it in order to train neural networks with stochastic gradient descent. In this paper, we investigate these training strategies and provide improvements over baselines. We find that during training, constraining the activations in a region that is far away from binary points leads to a better performance at test-time. The above finding provides a counter-intuitive result and leads to re-thinking the binarization approximation problem in neural networks. Çaglar Aytekin, Francesco Cricri, Jani Lainema, Emre Aksu, Miska M. Hannuksela |
IJCNN | 3 |
| 2019 | Inter-Component Transform for Color Video CodingabstractIn natural digital images and videos, correlations between color components can be observed. These correlations can be exploited to achieve additional coding gain in modern block-based hybrid video coding. To this end, we propose the use of a block-wise, rotational inter-component transform (ICT) applied to the two residual chroma signals that result from conventional intra or inter-picture prediction. Different ICT parameterizations in terms of number and quantization of the rotational angles as well as resulting components signaled in the coded bitstream are investigated. An implementation into the currently developed Versatile Video Coding (VVC) reference software provides average bitrate savings of up to 0.7% (All Intra configuration) with negligible increases in implementation complexity and runtime. Our proposal has been adopted into the VVC draft specification text. Christian Rudat, Christian R. Helmrich, Jani Lainema, Tung Nguyen 0001, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 3 |
| 2016 | HEVC still image coding and high efficiency image file formatabstractThe High Efficiency Video Coding (HEVC) standard includes support for a large range of image representation formats and provides an excellent image compression capability. The High Efficiency Image File Format (HEIF) offers a convenient way to encapsulate HEVC coded images, image sequences and animations together with associated metadata into a single file. This paper discusses various features and functionalities of the HEIF file format and compares the compression efficiency of HEVC still image coding to that of JPEG 2000. According to the experimental results HEVC provides about 25% bitrate reduction compared to JPEG 2000, while keeping the same objective picture quality. Jani Lainema, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital, Emre Aksu |
ICIP | 1 |
| 2014 | Differential Coding Using Enhanced Inter-Layer Reference Picture for the Scalable Extension of H.265/HEVC Video CodecabstractDifferential coding methods improve coding efficiency of scalable video codecs by adding the high-frequency component present in the previously coded enhancement layer (EL) pictures to the base layer (BL) picture. This paper proposes a method to enable differential coding in a scalable codec design without affecting the core coding tools, thus allowing a practical implementation to reuse single-layer hardware or software components. This is achieved by creating an additional reference picture called enhanced inter-layer reference (EILR) and inserting it to the EL decoded picture buffer and reference picture lists. An EILR picture is generated by adding differential information to the current inter-layer reference picture. The differential information is calculated using the previously decoded pictures of the BL and EL and the motion information of the BL picture. The proposed method reduces luma total bitrate on average by 2.2% and 2.8% for random access and low-delay test cases, respectively. The improvements are more significant for chroma components with the average bitrate reduction of 6.5%. The measured decoding time increase for a reference software implementation is 16% with negligible overhead on encoding time. Alireza Aminlou, Jani Lainema, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Interpolation filter design in HEVC and its coding efficiency - complexity analysisabstractCoding efficiency gains in the High Efficiency Video Coding (H.265/HEVC) standard are achieved by improving many aspects of the traditional hybrid coding framework. Motion compensated prediction, and in particular the interpolation filter, is one of the areas that was improved significantly over H.264/AVC. This paper presents the details of the motion compensation interpolation filter design of the H.265/HEVC standard and its improvements over the interpolation filter design of H.264/AVC. These improvements include discrete cosine transform based filter coefficient design, utilizing longer filter taps for luma and chroma interpolation and using higher precision operations in the intermediate computations. The computational complexity of HEVC interpolation filter is also analyzed both from theoretical and practical perspectives. Experimental results show that a 4.5% average bitrate reduction for the luma component and 13.0% average bitrate reduction for the chroma components are achieved compared to interpolation filter of H.264/AVC. The coding efficiency gains are significant for some video sequences and can reach up to 21.7%. Kemal Ugur, Alexander Alshin, Elena Alshina, Frank Bossen, Woojin Han 0001, Jeong-Hoon Park, Jani Lainema |
ICASSP | 7 |
| 2012 | Complexity analysis of next-generation HEVC decoderabstractThis paper analyzes the complexity of the HEVC video decoder being developed by the JCT-VC community. The HEVC reference decoder HM 3.1 is profiled with Intel VTune on Intel Core 2 Duo processor. The analysis covers both Low Complexity (LC) and High Efficiency (HE) settings for resolutions varying from WQVGA (416 × 240 pixels) up to 1600p (2560 × 1600 pixels). The yielded cycle-accurate results are compared with the respective results of H.264/AVC Baseline Profile (BP) and High Profile (HiP) reference decoders. HEVC offers significant improvement in compression efficiency over H.264/AVC: the average BD-rate saving of LC is around 51% over BP whereas the BD-rate gain of HE is around 45% over HiP. However, the average decoding complexities of LC and HE are increased by 61% and 87% over BP and HiP, respectively. In LC, the most complex functions are motion compensation (MC) and loop filtering (LF) that account on average for 50% and 14% of the decoder complexity. The decoding complexity of HE configuration is on average 42% higher than that of the LC configuration. Majority of the difference is caused by extra LF stages. In HE, the complexities of MC and LF are 37% and 32%, respectively. In practice, a standard 3 GHz dual core processor is expected to be able to decode 1080p HEVC content in real-time. Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Moncef Gabbouj, Jani Lainema |
ISCAS | 5 |
| 2012 | Intra Coding of the HEVC StandardabstractThis paper provides an overview of the intra coding techniques in the High Efficiency Video Coding (HEVC) standard being developed by the Joint Collaborative Team on Video Coding (JCT-VC). The intra coding framework of HEVC follows that of traditional hybrid codecs and is built on spatial sample prediction followed by transform coding and postprocessing steps. Novel features contributing to the increased compression efficiency include a quadtree-based variable block size coding structure, block-size agnostic angular and planar prediction, adaptive pre- and postfiltering, and prediction direction-based transform coefficient scanning. This paper discusses the design principles applied during the development of the new intra coding methods and analyzes the compression performance of the individual tools. Computational complexity of the introduced intra prediction algorithms is analyzed both by deriving operational cycle counts and benchmarking an optimized implementation. Using objective metrics, the bitrate reduction provided by the HEVC intra coding over the H.264/advanced video coding reference is reported to be 22% on average and up to 36%. Significant subjective picture quality improvements are also reported when comparing the resulting pictures at fixed bitrate. Jani Lainema, Frank Bossen, Woojin Han 0001, Junghye Min, Kemal Ugur |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Fast motion estimation with dual search window for stereo 3d video encodingabstractStereoscopic 3D video is becoming a reality in many application areas, ranging from high quality entertainment to mobile video services. Due to the need to process two views, the complexity of 3D video applications is significantly higher than traditional 2D counterparts. In order to enable real-time 3D video services in mobile devices, this paper proposes a novel algorithm which reduces the complexity of stereo video encoding with improvement of coding efficiency. A novel search window center prediction method is proposed that exploits the correlation between two views. Experimental results show that the average encoding time of the second view can be decreased by 80% with an increase in coding efficiency of up to 2%. The state-of-art fast motion estimation methods for stereoscopic 3D video encoding show coding efficiency decrease, whereas proposed method achieves the speed-up with increase in coding efficiency, making it suitable for high quality 3D video applications. Michal Joachimiak, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ICME | 3 |
| 2011 | Prediction Signal Aided Spatially Varying TransformabstractSpatially Varying Transform (SVT) is a technique introduced earlier to improve the coding efficiency of video coders [1][2]. SVT allows the position of the transform block within the macroblock to vary in order to better localize the underlying residual signal. The coding gains of SVT come with increased encoding complexity due to the additional need in the encoder to search for the best Location Parameter (LP) which indicates the position of the transform. In this paper, a new technique called Prediction Signal Aided Spatially Varying Transform (PSASVT) is proposed that utilizes the gradient of prediction signal to eliminate the unlikely LPs. As the number of candidate LPs is reduced, a smaller number of LPs are searched by encoder, which reduces the encoding complexity. In addition, less overhead bits are needed to code the selected LP and thus the coding efficiency can be improved. Experimental results show that the number of LPs to be tested in RDO is reduced on average by more than 20%. This reduction in encoding complexity is achieved with a slight increase in coding efficiency, as the number of candidate LPs is reduced. The decoding complexity increase is only a little. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
ICME | 3 |
| 2011 | Angular intra prediction in High Efficiency Video Coding (HEVC)abstractNew video coding solutions, such as the HEVC (High Efficiency Video Coding) standard being developed by JCT-VC (Joint Collaborative Team on Video Coding), are typically designed for high resolution video content. Increasing video resolution creates two basic requirements for practical video codecs; those need to be able to provide compression efficiency superior to prior video coding solutions and the computational requirements need to be aligned with the foreseeable hardware platforms. This paper proposes an intra prediction method which is designed to provide high compression efficiency and which can be implemented effectively in resource constrained environments making it applicable to wide range of use cases. When designing the method, special attention was given to the algorithmic definition of the prediction sample generation, in order to be able to utilize the same reconstruction process at different block sizes. The proposed method outperforms earlier variations of the same family of technologies significantly and consistently across different classes of video material, and has recently been adopted as the directional intra prediction method for the draft HEVC standard. Experimental results show that the proposed method outperforms the H.264/AVC intra prediction approach on average by 4.8 %. For sequences with dominant directional structures, the coding efficiency gains become more significant and exceed 10 %. Jani Lainema, Kemal Ugur |
MMSP | 1 |
| 2011 | Video Coding Using Spatially Varying TransformabstractIn this paper, a novel algorithm called spatially varying transform (SVT) is proposed to improve the coding efficiency of video coders. SVT enables video coders to vary the position of the transform block, unlike state-of-art video codecs where the position of the transform block is fixed. In addition to changing the position of the transform block, the size of the transform can also be varied within the SVT framework, to better localize the prediction error so that the underlying correlations are better exploited. It is shown in this paper that by varying the position of the transform block and its size, characteristics of prediction error are better localized, and the coding efficiency is thus improved. The proposed algorithm is implemented and studied in the H.264/AVC framework. We show that the proposed algorithm achieves 5.85% bitrate reduction compared to H.264/AVC on average over a wide range of test set. Gains become more significant at medium to high bitrates for most tested sequences and the bitrate reduction may reach 13.50%, which makes the proposed algorithm very suitable for future video coding solutions focusing on high fidelity video applications. The gain in coding efficiency is achieved with a similar decoding complexity which makes the proposed algorithm easy to be incorporated in video codecs. However, the encoding complexity of SVT can be relatively high because of the need to perform a number of rate distortion optimization (RDO) steps to select the best location parameter (LP), which indicates the position of the transform. In this paper, a novel low complexity algorithm is also proposed, operating on a macroblock and a block level, to reduce the encoding complexity of SVT. Experimental results show that the proposed low complexity algorithm can reduce the number of LPs to be tested in RDO by about 80% with only a marginal penalty in the coding efficiency. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Efficient SIMD-based implementation of adaptive filterabstractDirectional Adaptive Interpolation Filtering (DAIF) is a novel interpolation technique that was proposed recently for hybrid video coding. It was reported, that this technique outperforms the standard H.264/AVC interpolation in terms of coding gain whereas requiring smaller number of arithmetic operations. In this publication we present an optimized implementation of DAIF on a modern computing platform exploiting the Single Instruction Multiple Data (SIMD) parallelism. In addition, we provide a complexity analysis in which the computational complexity is estimated as number of clock cycles per output sample. Proposed SIMD-based implementation of DAIF has lower or comparable interpolation complexity, compared to the highly optimized SIMD-based implementation of the H.264/AVC interpolations. Considering significantly better coding gain provided by DAIF, we believe this approach will play a significant role in future video coding standards. Antti Hallapuro, Dmytro Rusanovskyy, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ISCAS | 4 |
| 2010 | Intra picture coding with planar representationsabstractIn this paper we introduce a novel concept for Intra coding of pictures especially suitable for representing smooth image segments. Traditional block based transform coding methods cause visually annoying blocking artifacts for image segments with gradually changing smooth content. The proposed solution overcomes this drawback by defining a fully continuous surface of sample values approximating the original image. The gradient of the surface is indicated by transmitting values for selected control points within the image segment and the surface itself is obtained by interpolating sample values in-between the control points. This approach is found to provide up to 30 percent bitrate reductions in the case of natural imagery and it has also been adopted to the initial HEVC codec design by JCT-VC. Jani Lainema, Kemal Ugur |
PCS | 1 |
| 2010 | Low complexity video coding and the emerging HEVC standardabstractThis paper describes a low complexity video codec with high coding efficiency. It was proposed to the High Efficiency Video Coding (HEVC) standardization effort of MPEG and VCEG, and has been partially adopted into the initial HEVC Test Model under Consideration design. The proposal utilizes a quad-tree structure with a support of large macroblocks of size 64×64 and 32×32, in addition to macroblocks of size 16×16. The entropy coding is done using a low complexity variable length coding based scheme with improved context adaptation over the H.264/AVC design. In addition, the proposal includes improved interpolation and deblocking filters, giving better coding efficiency while having low complexity. Finally, an improved intra coding method is presented. The subjective quality of the proposal is evaluated extensively and the results show that the proposed method achieves similar visual quality as H.264/AVC High Profile anchors with around 50% and 35% bit rate reduction for low delay and random-access experiments respectively at high definition sequences. This is achieved with less complexity than H.264/AVC Baseline Profile, making the proposal especially suitable for resource constrained environments. Kemal Ugur, Kenneth Andersson, Arild Fuldseth, Gisle Bjøntegaard, Lars Petter Endresen, Jani Lainema, Antti Hallapuro, Justin Ridge, Dmytro Rusanovskyy, Cixun Zhang, Andrey Norkin, Clinton Priddle, Thomas Rusert, Jonatan Samuelsson, Rickard Sjöberg, Zhuangfei Wu |
PCS | 6 |
| 2010 | High Performance, Low Complexity Video Coding and the Emerging HEVC StandardabstractThis paper describes a low complexity video codec with high coding efficiency. It was proposed to the high efficiency video coding (HEVC) standardization effort of moving picture experts group and video coding experts group, and has been partially adopted into the initial HEVC test model under consideration design. The proposal utilizes a quadtree-based coding structure with support for macroblocks of size 64$\,\times\,$64, 32$\,\times\,$32, and 16$\,\times\,$16 pixels. Entropy coding is performed using a low complexity variable length coding scheme with improved context adaptation compared to the context adaptive variable length coding design in H.264/AVC. The proposal's interpolation and deblocking filter designs improve coding efficiency, yet have low complexity. Finally, intra-picture coding methods have been improved to provide better subjective quality than H.264/AVC. The subjective quality of the proposed codec has been evaluated extensively within the HEVC project, with results indicating that similar visual quality to H.264/AVC High Profile anchors is achieved, measured by mean opinion score, using significantly fewer bits. Coding efficiency improvements are achieved with lower complexity than the H.264/AVC Baseline Profile, particularly suiting the proposal for high resolution, high quality applications in resource-constrained environments. Kemal Ugur, Kenneth Andersson, Arild Fuldseth, Gisle Bjøntegaard, Lars Petter Endresen, Jani Lainema, Antti Hallapuro, Justin Ridge, Dmytro Rusanovskyy, Cixun Zhang, Andrey Norkin, Clinton Priddle, Thomas Rusert, Jonatan Samuelsson, Rickard Sjöberg, Zhuangfei Wu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2009 | Video coding using Variable Block-Size Spatially Varying TransformsabstractIn our previous work, we introduced Spatially Varying Transforms (SVT) for video coding, where the location of the transform block within the macroblock is not fixed but varying. In this paper, we extend this concept and present a novel method, called Variable Block-size Spatially Varying Transforms (VBSVT). VBSVT utilizes Variable Block-size Transforms (VBT) in the SVT framework, and is shown to be more preferable for coding prediction error with different characteristics than fixed block-size SVT and also the standard methods that use fixed or adaptive block sizes at fixed spatial locations. In addition, VBSVT has similar decoding complexity with fixed block-size SVT and lower decoding complexity compared to standard methods as only a portion of the prediction error needs to be decoded. Experimental results show that, VBSVT achieves 4.1% gain over H.264/AVC on average over a wide range of test set. Gains become more significant at high quality levels and go up to 13.5%, which makes the proposed algorithm very suitable for future video coding solutions focusing on high fidelity applications. Cixun Zhang, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ICASSP | 3 |
| 2009 | Low complexity algorithm for Spatially Varying TransformsabstractIn our previous work, we introduced spatially varying transforms (SVT) for video coding, where the location of the transform block within the macroblock is not fixed but varying. SVT has lower decoding complexity compared to standard methods as only a portion of the prediction error needs to be decoded. However, the encoding complexity of SVT can be relatively high because of the need to perform rate distortion optimization (RDO) for each candidate location parameter (LP). In this work, we propose a low complexity algorithm operating on macroblock and block level to reduce the encoding complexity of SVT. The proposed low complexity algorithm includes selection of available candidate LP based on motion difference and a hierarchical search algorithm. Experimental results show that the proposed low complexity algorithm can reduce around 80% of the candidate LP tested in RDO with only marginal penalty in coding efficiency. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
PCS | 3 |
| 2009 | Video Coding Using Spatially Varying Transform
Cixun Zhang, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
PSIVT | 3 |
| 2009 | Video Coding With Low-Complexity Directional Adaptive Interpolation FiltersabstractA novel adaptive interpolation filter structure for video coding with motion-compensated prediction is presented in this letter. The proposed scheme uses an independent directional adaptive interpolation filter for each sub-pixel location. The Wiener interpolation filter coefficients are computed analytically for each inter-coded frame at the encoder side and transmitted to the decoder. Experimental results show that the proposed method achieves up to 1.1 dB coding gain and a 15% average bit-rate reduction for high-resolution video materials compared to the standard nonadaptive interpolation scheme of H.264/AVC, while requiring 36% fewer arithmetic operations for interpolation. The proposed interpolation can be implemented in exactly 16-bit arithmetic, thus it can have important use-cases in mobile multimedia environments where the computational resources are severely constrained. Dmytro Rusanovskyy, Kemal Ugur, Antti Hallapuro, Jani Lainema, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Video coding using pruned transforms and interleaving of multiple blockabstractTechnologies used in today's video coding standards have been designed and optimized mainly for standard definition (SD) resolutions and below. When moving to higher resolutions, data in video frames tend to become more correlated spatially. In this paper, we study how to take advantage of this phenomenon to lower computational requirements for high definition (HD) video coding. A coding method based on low complexity pruned transforms and interleaving of multiple transform coefficient blocks is proposed. An example implementation of this method in the context of H.264/AVC is also presented. Experimental results show that the proposed method maintains the high compression efficiency of H.264/AVC while significantly lowering the coding complexity of the codec. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
ICASSP | 3 |
| 2008 | Video coding with pixel-aligned directional adaptive interpolation filtersabstractIn this paper a novel adaptive interpolation filter structure is proposed to improve the coding efficiency of video coders. Proposed scheme utilizes one dimensional directional adaptive filter for every of the sub-pixel location, whose coefficients are calculated analytically for every frame by minimizing the prediction error energy. The direction of the interpolation filter is different for every sub-pixel position and it is determined based on the alignment of the corresponding sub- pixel with integer pixel samples. Experimental results show that, the proposed method achieves up-to 1.1 dB gain compared to the standard non-adaptive interpolation scheme of H.264/AVC, requiring less number of operations for interpolation. Compared to two-dimensional non-separable adaptive interpolation, proposed scheme has practically the same coding efficiency with approximately 3 times less complexity. Since significant coding efficiency is achieved without increasing the complexity, it is believed that proposed method has important use-cases in mobile multimedia environments where the resources are severely constrained. Dmytro Rusanovskyy, Kemal Ugur, Moncef Gabbouj, Jani Lainema |
ISCAS | 4 |
| 2007 | Adaptive Interpolation Filter with Flexible Symmetry for Coding High Resolution High Quality VideoabstractIn this work, a novel sub-pixel interpolation algorithm is proposed for video coders targeted towards high resolution and high fidelity use cases. Proposed scheme is based on adapting the interpolation filter's symmetry assumptions in a rate-distortion-optimized fashion, taking into account the coding rate and the statistical properties of each image of the video sequence. Experimental results show that, using the proposed algorithm a gain of up-to 1.1 dB is achieved compared to non-adaptive sub-pixel interpolation of H.264/AVC. Compared to other state-of-the-art adaptive sub-pixel interpolation methods, a gain of up-to 0.5 dB is achieved. Proposed scheme outperforms H.264/AVC for all test cases; however, improvement is more significant at high bitrates and at high resolutions. This is especially important for future video coding solutions targeting high fidelity video applications. Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ICASSP (1) | 2 |
| 2006 | Generating H.264/AVC Compliant Bitstreams for Lightweight Decoding Operation Suitable for Mobile Multimedia SystemsabstractIn this work, we propose novel encoder algorithms for the state-of-the-art video coding standard H.264, to generate decoder friendly video bitstreams. Using the proposed algorithms, it is possible to generate bitstreams requiring significantly less decoding complexity, with negligible effect on picture quality. This is achieved by using novel algorithms for mode decision and motion estimation that bias easy-to-decode motion vectors in a rate-distortion optimized fashion. Experimental results show that, more than 15% decoding complexity reduction is achieved with less than a 0.1 dB penalty on the average video quality. We believe that this approach has potential in various use cases especially in mobile multimedia systems, where the video decoder operation is often dominating the handsets power consumption Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
ICASSP (2) | 2 |
| 2003 | Adaptive deblocking filterabstractThis paper describes the adaptive deblocking filter used in the H.264/MPEG-4 AVC video coding standard. The filter performs simple operations to detect and analyze artifacts on coded block boundaries and attenuates those by applying a selected filter. Peter List 0001, Anthony Joch, Jani Lainema, Gisle Bjøntegaard, Marta Karczewicz |
IEEE Trans. Circuits Syst. Video Technol. | 3 |