VLDB 2026 Research / reviewers in the wild / expert
Jingning Han
dblp:15/8011
· DBLP profile ↗
62ranked-venue papers
22as first author
20since 2021 · last 2025
0000-0001-7168-2254ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 21 first-author · 18 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SELIC: Semantic-Enhanced Learned Image Compression via High-Level Textual GuidanceabstractLearned image compression (LIC) techniques have achieved remarkable progress; however, effectively integrating high-level semantic information remains challenging. In this work, we present a Semantic-Enhanced Learned Image Compression framework, termed SELIC, which leverages high-level textual guidance to improve rate-distortion performance. Specifically, SELIC employs a text encoder to extract rich semantic descriptions from the input image. These textual features are transformed into fixed-dimension tensors and seamlessly fused with the image-derived latent representation. By embedding the SELIC tensor directly into the compression pipeline, our approach enriches the bitstream without requiring additional inputs at the decoder, thereby maintaining fast and efficient decoding. Extensive experiments on benchmark datasets (e.g., Kodak) demonstrate that integrating semantic information substantially enhances compression quality. Our SELIC-guided method outperforms a baseline LIC model without semantic integration by approximately 0.1-0.15 dB across a wide range of bit rates in PSNR and achieves a 4.9% BD-rate improvement over VVC. Moreover, this improvement comes with minimal computational overhead, making the proposed SELIC framework a practical solution for advanced image compression applications. Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han |
ICME | 4 |
| 2024 | Learned Image Compression with Dual-Branch Encoder and Conditional Information CodingabstractRecent advancements in deep learning-based image compression are notable. However, prevalent schemes that employ a serial context-adaptive entropy model to enhance rate-distortion (R-D) performance are markedly slow. Furthermore, the complexities of the encoding and decoding networks are substantially high, rendering them unsuitable for some practical applications. In this paper, we propose two techniques to balance the trade-off between complexity and performance. First, we introduce two branching coding networks to independently learn a low-resolution latent representation and a high-resolution latent representation of the input image, discriminatively representing the global and local information therein. Second, we utilize the high-resolution latent representation as conditional information for the low-resolution latent representation, furnishing it with global information, thus aiding in the reduction of redundancy between low-resolution information. We do not utilize any serial entropy models. Instead, we employ a parallel channel-wise auto-regressive entropy model for encoding and decoding low-resolution and high-resolution latent representations. Experiments demonstrate that our method is approximately twice as fast in both encoding and decoding compared to the parallelizable checkerboard context model, and it also achieves a 1.2% improvement in R-D performance compared to state-of-the-art learned image compression schemes. Our method also outperforms classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods in rate-distortion performance, as validated by both PSNR and MS-SSIM metrics on the Kodak dataset. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han |
DCC | 6 |
| 2024 | A Saliency Map Approach to Optimize VMAF for Video and Image CompressionabstractThe Video Multi-method Assessment Fusion (VMAF) has demonstrated a better correlation with Human Visual System than the conventional objective metrics, and has gradually gained adoption in the industry that needs to monitor the visual quality of compressed videos. However, due to its machine learning nature, it can not be expressed through a simple and explicit mathematical formula, which makes it difficult to incorporate VMAF into the rate-distortion optimization framework in video compression. In this work, we propose a new perspective that decomposes the VMAF as a superposition of spatial and temporal factors. The spatial factor, which also directly applies to image quality evaluation, is approximated by a saliency map. It in conjunction with the temporal factor approximated by the motion quantities allows a simple analytical formula that translates the mean squared distortion at each pixel to its impact to the overall VMAF metric. The proposed hypothesis is embedded into the rate-distortion optimization framework, and is experimentally shown to provide considerable coding gains in VMAF for both image and video compression. Jingning Han, Yaowu Xu |
DCC | 2 |
| 2024 | WeConvene: Learned Image Compression with Wavelet-Domain Convolution and Entropy Model
Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han, Feng Liang 0001, Guohe Zhang |
ECCV (50) | 4 |
| 2024 | Efficient Learned Image Compression with Selective Kernel Residual Module and Channel-Wise Causal Context ModelabstractRecently, learning-based image compression approaches have achieved superior performance over classical image compression methods. However, their complexities remain quite high. In this paper, we propose two efficient modules to reduce the complexity. First, we introduce a selective kernel residual module into the core network, which effectively expands the receptive field and captures global information. Second, we present an improved channel-wise causal context model, designed to not only reduce encoding and decoding time but also ensure rate-distortion performance. Experimental results demonstrate that our proposed method achieves better tradeoff than recent leading learned image compression methods, and also outperforms the latest H.266/VVC (4:4:4) in terms of PSNR and MS-SSIM metrics. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han |
ICASSP | 6 |
| 2024 | Fast and High-Performance Learned Image Compression With Improved Checkerboard Context Model, Deformable Residual Module, and Knowledge DistillationabstractDeep learning-based image compression has made great progresses recently. However, some leading schemes use serial context-adaptive entropy model to improve the rate-distortion (R-D) performance, which is very slow. In addition, the complexities of the encoding and decoding networks are quite high and not suitable for many practical applications. In this paper, we propose four techniques to balance the trade-off between the complexity and performance. We first introduce the deformable residual module to remove more redundancies in the input image, thereby enhancing compression performance. Second, we design an improved checkerboard context model with two separate distribution parameter estimation networks and different probability models, which enables parallel decoding without sacrificing the performance compared to the sequential context-adaptive model. Third, we develop a three-pass knowledge distillation scheme to retrain the decoder and entropy coding, and reduce the complexity of the core decoder network, which transfers both the final and intermediate results of the teacher network to the student network to improve its performance. Fourth, we introduce$L_{1}$regularization to make the numerical values of the latent representation more sparse, and we only encode non-zero channels in the encoding and decoding process to reduce the bit rate. This also reduces the encoding and decoding time. Experiments show that compared to the state-of-the-art learned image coding scheme, our method can be about 20 times faster in encoding and 70-90 times faster in decoding, and our R-D performance is also 2.3% higher. Our method achieves better rate-distortion performance than classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods, as measured by both PSNR and MS-SSIM metrics on the Kodak and Tecnick-40 datasets. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han |
IEEE Trans. Image Process. | 7 |
| 2023 | Multi-Rate Adaptive Transform Coding for Video CompressionabstractContemporary lossy image and video coding standards rely on transform coding, the process through which pixels are mapped to an alternative representation to facilitate efficient data compression. Despite impressive performance of end-to-end optimized compression with deep neural networks, the high computational and space demands of these models has prevented them from superseding the relatively simple transform coding found in conventional video codecs. In this study, we propose learned transforms and entropy coding that may either serve as (non)linear drop-in replacements, or enhancements for linear transforms in existing codecs. These transforms can be multi-rate, allowing a single model to operate along the entire rate-distortion curve. To demonstrate the utility of our framework, we augmented the DCT with learned quantization matrices and adaptive entropy coding to compress intra-frame AV1 block prediction residuals. We report substantial BD-rate and perceptual quality improvements over more complex nonlinear transforms at a fraction of the computational cost. Lyndon R. Duong, Bohan Li 0006, Jingning Han |
ICASSP | 4 |
| 2023 | ROI-Based Deep Image Compression with Swin TransformersabstractEncoding the Region Of Interest (ROI) with better quality than the background has many applications including video conferencing systems, video surveillance and object-oriented vision tasks. In this paper, we propose a ROI-based image compression framework with Swin transformers as main building blocks for the autoencoder network. The binary ROI mask is integrated into different layers of the network to provide spatial information guidance. Based on the ROI mask, we can control the relative importance of the ROI and non-ROI by modifying the corresponding Lagrange multiplier λ for different regions. Experimental results show our model achieves higher ROI PSNR than other methods and modest average PSNR for human evaluation. When tested on models pre-trained with original images, it has superior object detection and instance segmentation performance on the COCO validation dataset. Jie Liang 0001, Haisheng Fu, Jingning Han |
ICASSP | 4 |
| 2023 | Learned Image Compression Guided Adaptive Quantization for Perceptual QualityabstractNeural network based image compression has made significant progress in recent years. The learned image codecs are commonly reported to outperform their conventional counterparts in perceptual quality. Despite the superior performance, the learned image codecs are much more complex to decode, which hinders their usage in practice. Without a significant advance in hardware capability, the conventional image codec will likely remain a primary component for large scale image services. It is therefore desirable to improve the quality of conventional image codecs. In this paper, we present an adaptive quantization approach to the conventional image codec with the help of learned image codecs to improve its perceptual quality. It exploits the bit allocation of the neural network based image codec to adapt the quantizers on a block basis. It is experimentally shown that the proposed method provides considerable perceptual quality improvements over other leading contenders. Ruiqi Geng, Bohan Li 0006, Maryla Ustarroz-Calonge, Frank Galligan, Jingning Han, Yaowu Xu |
ICIP | 6 |
| 2023 | Asymmetric Learned Image Compression With Multi-Scale Residual Block, Importance Scaling, and Post-Quantization FilteringabstractRecently, deep learning-based image compression has made significant progresses, and has achieved better rate-distortion (R-D) performance than the latest traditional method, H.266/VVC, in both MS-SSIM metric and the more challenging PSNR metric. However, a major problem is that the complexities of many leading learned schemes are too high. In this paper, we propose an efficient and effective image coding framework, which achieves similar R-D performance with lower complexity than the state of the art. First, we develop an improved multi-scale residual block (MSRB) that can expand the receptive field and capture global information more efficiently, which further reduces the spatial correlation of the latent representations. Second, an importance scaling network is introduced to directly scale the latents to achieve content-adaptive bit allocation without sending side information, which is more flexible than previous importance map methods. Third, we apply a post-quantization filter (PQF) to reduce the quantization error, motivated by the Sample Adaptive Offset (SAO) filter in video coding. Moreover, our experiments show that the performance of the system is less sensitive to the complexity of the decoder. Therefore, we design an asymmetric paradigm, in which the encoder employs three stages of MSRBs to improve the learning capacity, whereas the decoder only uses one stage of MSRB, which reduces the decoder complexity and still yields satisfactory performance. Experimental results show that compared to the state-of-the-art method, the encoding and decoding time of the proposed method are about 17 times faster, and the R-D performance is only reduced by about 1% on both Kodak and Tecnick-40 datasets, which is still better than H.266/VVC(4:4:4) and other leading learning-based methods. Our source code is publicly available athttps://github.com/fengyurenpingsheng. Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jingning Han |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Learned Image Compression With Gaussian-Laplacian-Logistic Mixture Model and Concatenated Residual ModulesabstractRecently deep learning-based image compression methods have achieved significant achievements and gradually outperformed traditional approaches including the latest standard Versatile Video Coding (VVC) in both PSNR and MS-SSIM metrics. Two key components of learned image compression are the entropy model of the latent representations and the encoding/decoding network architectures. Various models have been proposed, such as autoregressive, softmax, logistic mixture, Gaussian mixture, and Laplacian. Existing schemes only use one of these models. However, due to the vast diversity of images, it is not optimal to use one model for all images, even different regions within one image. In this paper, we propose a more flexible discretized Gaussian-Laplacian-Logistic mixture model (GLLMM) for the latent representations, which can adapt to different contents in different images and different regions of one image more accurately and efficiently, given the same complexity. Besides, in the encoding/decoding network design part, we propose a concatenated residual blocks (CRB), where multiple residual blocks are serially connected with additional shortcut connections. The CRB can improve the learning ability of the network, which can further improve the compression performance. Experimental results using the Kodak, Tecnick-100 and Tecnick-40 datasets show that the proposed scheme outperforms all the leading learning-based methods and existing compression standards including VVC intra coding (4:4:4 and 4:2:0) in terms of the PSNR and MS-SSIM. The source code is available at https://github.com/fengyurenpingsheng. Haisheng Fu, Feng Liang 0001, Bing Li 0022, Jie Liang 0001, Guohe Zhang, Dong Liu 0002, Chengjie Tu, Jingning Han |
IEEE Trans. Image Process. | 10 |
| 2022 | An Efficient Scheme of Multi-Hypothesis Motion Compensated Prediction for Video Coding ApplicationsabstractPrior research has demonstrated that the multi-hypothesis motion compensated prediction (MCP) can theoretically provide a better prediction quality than single-reference MCP, thereby improving the compression efficiency in video coding. However, the existing multi-hypothesis MCP methods typically require either additional rate cost to transmit the motion vectors, or significant decoding complexity to conduct the motion search at the decoder end, which is usually expensive. In this work, we propose a novel scheme to materialize the multi-hypothesis MCP that requires no additional rate cost, nor extra motion search on either the encoder or decoder side. Various approaches to synthesize these available multiple references to form the inter prediction are presented. We experimentally demonstrate that the proposed scheme provides considerable and consistent coding gains across a wide range of operating points. Bohan Li 0006, Jingning Han, Yaowu Xu |
ICIP | 2 |
| 2022 | Differential Contrast Based Adaptive Quantization for Perceptual Quality Optimization in Image CodingabstractWe consider the perceptual quality optimization in image coding through adaptive quantization. A differential contrast model is proposed to measure the visual sensitivity to the quantization distortions, and thereby deriving the spatially adaptive quantization strategy. A complementary quantitative approach is provided as a means to efficiently calculate the proposed differential contrast model. The resulting visual quality improvement is experimentally demonstrated. Jingning Han, Frank Galligan, Pascal Massimino, Paul Wilkins, Wan-Teh Chang, Yannis Guyon, Yaowu Xu, Jim Bankoski |
ICIP | 1 |
| 2022 | Probability Model Estimation for M-Ary Random VariablesabstractThe entropy coding system in AV1 processes syntax elements as M-ary random variables. In comparison to the binarization approach used in its predecessor VP9 that converts an M-ary random variables into a series of binary symbols for entropy coding, the M-ary random variable approach provides higher throughput for hardware decoders. The non-binary probability table associated with the M-ary random variable, however, poses new challenges in the probability model estimation process beyond the binary case. This paper provides a retrospect of the probability model estimation for M-ary random variables used in AV1, and proposes new algorithms for the probability estimation process to improve the compression efficiency. Its efficacy is experimentally demonstrated under various testing conditions. Jingning Han, Yaowu Xu |
ICIP | 1 |
| 2022 | Region-of-interest and channel attention-based joint optimization of image compression and computer vision
Linwei Ye, Jie Liang 0001, Yang Wang 0003, Jingning Han |
Neurocomputing | 5 |
| 2021 | Learned Bi-Resolution Image Coding using Generalized Octave ConvolutionsabstractLearned image compression has recently shown the potential to outperform the standard codecs. State-of-the-art rate-distortion (R-D) performance has been achieved by context-adaptive entropy coding approaches in which hyperprior and autoregressive models are jointly utilized to effectively capture the spatial dependencies in the latent representations. However, the latents are feature maps of the same spatial resolution in previous works, which contain some redundancies that affect the R-D performance. In this paper, we propose a learned bi-resolution image coding approach that is based on the recently developed octave convolutions to factorize the latents into high and low resolution components. Therefore, the spatial redundancy is reduced, which improves the R-D performance. Novel generalized octave convolution and octave transposed-convolution architectures with internal activation layers are also proposed to preserve more spatial structure of the information. Experimental results show that the proposed scheme outperforms all existing learned methods as well as standard codecs such as the next-generation video coding standard VVC (4:2:0) in both PSNR and MS-SSIM. We also show that the proposed generalized octave convolution can improve the performance of other auto-encoder-based schemes such as semantic segmentation and image denoising. Jie Liang 0001, Jingning Han, Chengjie Tu |
AAAI | 3 |
| 2021 | Adaptive GOP Size Decision for Multi-Pass Video Coding Based on Hidden Markov ModelabstractMulti-pass coding is a widely utilized technique to improve the compression efficiency in video coding, where frame statistics are collected from the previous passes and then analyzed to provide better encoder decisions, such as rate control parameters, prediction mode selection, motion estimation, etc. In this paper, a novel method to determine the size of each group of picture (GOP) using the multi-pass information is presented. In particular, we propose to categorize frames into regions with different natures, including stationary, high-variance, blending, and scene cut, through analyzing the frame statistics generated from the previous passes using a hidden Markov model. The GOP size is then determined based on the region types and the inter frame correlations. It is experimentally shown that the proposed adaptive GOP size decision provides considerable coding performance improvements over conventional fixed GOP length. Bohan Li 0006, Jingning Han, Yaowu Xu |
ICASSP | 2 |
| 2021 | A Temporal Filtering Approach Based on Optical Flow Estimation for Video CodingabstractVideo coding uses motion compensated prediction to exploit temporal correlations for compression efficiency. Prior works have demonstrated that substantial coding gains can be achieved by decomposing a long-term reference frame into a synthetic reference-only (non-displayable) frame and an overlay displayable frame that resembles the original frame. The source of the reference-only frame is typically generated by temporal filtering along the motion trajectories across nearby frames, where the motion trajectories are built using block matching algorithms (BMAs), thereby reducing the noise level within this synthetic frame. Noting that the efficacy of the conventional BMAs are limited to capturing translational motion activities, this paper proposes a novel approach that uses a per-pixel motion field generated by an optical flow estimation to form the motion trajectory for more efficient temporal filtering. It is experimentally shown that the proposed method better captures non-translational motion activities, which translates into considerable coding gains for video signals with such complicate motion patterns. Bohan Li 0006, Lauren Partin, Jingning Han, Yaowu Xu |
MMSP | 3 |
| 2021 | A Technical Overview of AV1abstractThe AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than a 30% reduction in bit rate compared to its predecessor VP9 for the same decoded video quality. This article provides a technical overview of the AV1 codec design that enables the compression performance gains with considerations for hardware feasibility. Jingning Han, Bohan Li 0006, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, Yue Chen 0040, Yunqing Wang, Paul Wilkins, Yaowu Xu, Jim Bankoski |
Proc. IEEE | 1 |
| 2021 | Learned Multi-Resolution Variable-Rate Image Compression With Octave-Based Residual BlocksabstractRecently deep learning-based image compression has shown the potential to outperform traditional codecs. However, most existing methods train multiple networks for multiple bit rates, which increase the implementation complexity. In this paper, we propose a new variable-rate image compression framework, which employs generalized octave convolutions (GoConv) and generalized octave transposed-convolutions (GoTConv) with built-in generalized divisive normalization (GDN) and inverse GDN (IGDN) layers. Novel GoConv- and GoTConv-based residual blocks are also developed in the encoder and decoder networks. Our scheme also uses a stochastic rounding-based scalar quantization. To further improve the performance, we encode the residual between the input and the reconstructed image from the decoder network as an enhancement layer. To enable a single model to operate with different bit rates and to learn multi-rate image features, a new objective function is introduced. Experimental results show that the proposed framework trained with variable-rate objective function outperforms the standard codecs such as H.265/HEVC-based BPG and state-of-the-art learning-based variable-rate methods. Jie Liang 0001, Jingning Han, Chengjie Tu |
IEEE Trans. Multim. | 3 |
| 2020 | Video Denoising for the Hierarchical Coding Structure in Video CodingabstractModern video codecs explore the temporal and spatial correlations of video signal to achieve the goal of compression. The noise in video signal corrupts such temporal and spatial correlations and thus is difficult to compress. Denoising of video signal is a potential solution to this problem. Despite the significant progress in video denoising in recent years, there is few research exploring the feasibility of denoising for video compression. In this work, we demonstrate that video denoising is able to significantly reduce bit rates while maintaining the subjective and objective quality when appropriately incorporated into the hierarchical coding structure of video coding. We present a temporal filtering algorithm for denoising and apply it to AV1 for lossy video compression. We obtain a significant compression efficiency improvement over videos of different resolutions, types, and noise. Jingning Han, Yaowu Xu |
DCC | 2 |
| 2020 | Online Probability Model Estimation for Video CompressionabstractModern video codec uses arithmetic coding for entropy coding. The arithmetic coding asymptotically achieves the entropy bound provided the true probability distribution. Hence the compression efficiency heavily relies on the ability to capture the time-variant probability model in video signals. Variants of first-order linear probability model update schemes have been used in recent generation video codecs. Built on top of those, a multimodal estimation scheme that forms a higher order probability model update has been proposed in this work. We experimentally demonstrate its coding efficiency. Jingning Han, Yaowu Xu |
DCC | 2 |
| 2020 | An Adaptive Linear Estimator Based Approach to Bi-Directional Motion Compensated PredictionabstractBi-directional motion compensated prediction is widely utilized in video coding. Conventionally, the encoder searches for two motion vectors pointing to reference frames in both directions, and transmits these motion vectors to the decoder. Recognizing that the two reference frames are already available to the decoder, prior work proposed decoder-side motion estimation to extract motion information or optical flow, at the cost of dramatic increase in decoder complexity. This paper proposes a novel bi-directional motion compensation mode that efficiently utilizes the motion information that is already available to the decoder, without recourse to extensive search. An estimation theory based approach is proposed and utilized to provide a high quality prediction, which adaptively combines contributions from multiple motion-compensated references. Experimental results show that the proposed method, while yielding a greatly reduced decoder side complexity, introduces a significant coding gain for a diverse set of video sequences. Bohan Li 0006, Jingning Han, Kenneth Rose |
ICASSP | 2 |
| 2020 | A Non-local Mean Temporal Filter for Video CompressionabstractModern video codecs exploit the temporal and spatial correlations of video signal to achieve compression. The noise in video signal corrupts such correlations and impairs the coding efficiency. Prior works in VP8, VP9, and HEVC exploit the use of temporal filtering to remove certain noise from the source signal. They typically compare a pair of pixels along a motion trajectory and decide the filter coefficients based on the pixel value difference. It is observed that such noise removal allows better rate-distortion performance trade off and hence improves the objective compression efficiency. Note that the compression distortion is evaluated against the original video signal in all cases. This work proposes a non-local mean temporal filter for noise removal. Instead of comparing a pair of pixels along the motion trajectory, it compares two pixel blocks surrounding the pixels of interest. Their distance in L2 norm is then normalized by the frame noise level, which is used to determine the temporal filter coefficients in a non-parametric model. It is experimentally shown that the proposed non-local mean filter approach achieves improved compression efficiency over other contenders. Jingning Han, Yaowu Xu |
ICIP | 2 |
| 2020 | Learned Variable-Rate Image Compression With Residual Divisive NormalizationabstractRecently deep learning-based image compression has shown the potential to outperform traditional codecs. However, most existing methods train multiple networks for multiple bit rates, which increases the implementation complexity. In this paper, we propose a variable-rate image compression framework, which employs more Generalized Divisive Normalization (GDN) layers than previous GDN-based methods. Novel GDN-based residual sub-networks are also developed in the encoder and decoder networks. Our scheme also uses a stochastic rounding-based scalar quantization. To further improve the performance, we encode the residual between the input and the reconstructed image from the decoder network as an enhancement layer. To enable a single model to operate with different bit rates and to learn multi-rate image features, a new objective function is introduced. Experimental results show that the proposed framework trained with variable-rate objective function outperforms all standard codecs such as H.265/HEVC-based BPG and state-of-the-art learning-based variable-rate methods. Jie Liang 0001, Jingning Han, Chengjie Tu |
ICME | 3 |
| 2020 | VMAF Based Rate-Distortion Optimization for Video CodingabstractVideo Multi-method Assessment Fusion (VMAF) is a machine-learning based video quality metric. It is experimentally shown to provide higher correlation with human visual system as compared to conventional metrics like peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) in many scenarios and has drawn considerable interest as an alternative metric to evaluate the perceptual quality. This work proposes a systematic approach to improve the video compression performance in VMAF. It is composed of multiple components including a pre-processing stage with a complement automatic filter parameter selection, and a modified rate-distortion optimization framework tailored for VMAF metric. The proposed scheme achieves on average 37% BD-rate reduction in VMAF, as compared to conventional video codec optimized for PSNR. Sai Deng, Jingning Han, Yaowu Xu |
MMSP | 2 |
| 2020 | Optical Flow Based Co-Located Reference Frame for Video CompressionabstractThis paper proposes a novel bi-directional motion compensation framework that extracts existing motion information associated with the reference frames and interpolates an additional reference frame candidate that is co-located with the current frame. The approach generates a dense motion field by performing optical flow estimation, so as to capture complex motion between the reference frames without recourse to additional side information. The estimated optical flow is then complemented by transmission of offset motion vectors to correct for possible deviation from the linearity assumption in the interpolation. Various optimization schemes specifically tailored to the video coding framework are presented to further improve the performance. To accommodate applications where decoder complexity is a cardinal concern, a block-constrained speed-up algorithm is also proposed. Experimental results show that the main approach and optimization methods yield significant coding gains across a diverse set of video sequences. Further experiments focus on the trade-off between performance and complexity, and demonstrate that the proposed speed-up algorithm offers complexity reduction by a large factor while maintaining most of the performance gains. Bohan Li 0006, Jingning Han, Yaowu Xu, Kenneth Rose |
IEEE Trans. Image Process. | 2 |
| 2019 | A Multi-Pass Coding Mode Search Framework For AV1 Encoder OptimizationabstractThe AV1 codec recently released by the Alliance of Open Media provides nearly 30% BDrate reduction over its predecessor VP9. It substantially extends the available coding block sizes and supports a wide range of prediction modes. There are also a large variety of transform kernel types and sizes. The combination provides an extremely wide range of flexible coding options. To translate such flexibility into compression efficiency, the encoder needs to conduct an extensive search over the space of coding modes. Optimization of the encoder complexity and compression efficiency trade-off is critical to productionizing AV1. Many research efforts have been devoted to devising feature space based pruning methods ranging from decision rules based on some simple observations to more complex neural network models. A multi-pass coding mode search framework is proposed in this work to provide a structural approach to reduce the search volume. It decomposes the original high dimensional space search into cascaded stages of lower dimensional space searches. To retain a near optimal search result, the scheme departs from conventional dimension reduction approach in which one retains a single winner at each stage, and uses that winner for the next stage (dimension). Instead, this framework retains a subset of the states that are the most likely winners at each stage, which are then fed into the next stage to find the next subset of winners. The subset size at each stage is determined by the likelihood that the optimal route will be captured in the current stage. Changing this likelihood parameter tunes the encoder for speed and compression performance trade-off. This framework can integrate with most existing feature based methods at its various stages. The framework provides 60% encoding time reduction at the expense of 0.6% compression loss in libaom AV1 encoder. Ching-Han Chiang, Jingning Han, Yaowu Xu |
DCC | 2 |
| 2019 | DSSLIC: Deep Semantic Segmentation-based Layered Image CompressionabstractDeep learning has revolutionized many computer vision fields in the last few years, including learning-based image compression. In this paper, we propose a deep semantic segmentation-based layered image compression (DSSLIC) framework in which the segmentation map of the input image is obtained and encoded as the base layer of the bit-stream. A compact representation of the input image is also generated and encoded as the first enhancement layer. The segmentation map and the compact version of the image are then employed to obtain a coarse reconstruction of the image. The residual between the input and the coarse reconstruction is additionally encoded as another enhancement layer. Experimental results show that the proposed framework outperforms the H.265/HEVC-based BPG and other codecs in both PSNR and MS-SSIM metrics in RGB domain. Besides, since semantic map is included in the bit-stream, the proposed scheme can facilitate many other tasks such as image search and object-based adaptive image compression1. Jie Liang 0001, Jingning Han |
ICASSP | 3 |
| 2019 | JND-based Perceptual Rate Distortion Optimization for AV1 EncoderabstractAV1 is the next-generation open video coding format, and it can achieve significant coding efficiency with novel coding tools. It supports Lagrangian rate distortion optimization (RDO) method to optimize the coding performance. However, the distortion and the Lagrangian multiplier used in RDO ignore the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. To solve this problem, a perceptual RDO scheme based on the Just Noticeable Distortion (JND) threshold of HVS is proposed. The JND for each pixel is first measured according to three perceptual features: luminance adaptation, masking effects and structure sensitivity. Based on the observation that the regions with smaller distortion visibility thresholds are more sensitive to HVS, a JND-based Lagrangian multiplier is derived to adaptively adjust the rate-distortion (RD) performance for each coding block. Experiments demonstrate that the proposed method can achieve an average SSIM-based -3.93% BD-Rate saving compared with the original AV1 encoder, which effectively improve the coding performance. Li Song 0001, Rong Xie 0004, Jingning Han, Yaowu Xu |
PCS | 4 |
| 2018 | Co-located Reference Frame Interpolation Using Optical Flow Estimation for Video CompressionabstractThe hierarchical coding structure that supports bi-directional motion compensated prediction is commonly used for video compression efficiency. Conventional approach directly seeks the reference pixel block from each individual reference frame and use it or its linear combinations for prediction. It largely ignores the motion information between these reference frames. To fully utilize all the information from the bi-directional reference frames, this work builds a per-pixel motion field that connects the two-sided reference frames using optical flow estimation. A reference frame is then interpolated at the current frame location. This collocated reference frame effectively accounts for the true motion trajectories in the video signal including both translational and the more complex non-translational motion models, which are beyond the capability of the conventional block-based motion compensated prediction. The scheme is experimentally shown to provide substantial compression performance gains. A number of optimization designs are proposed to make the codec complexity feasible while largely maintaining the coding performance. Bohan Li 0006, Jingning Han, Yaowu Xu |
DCC | 2 |
| 2018 | Efficient AV1 Video Coding Using a Multi-layer FrameworkabstractThis paper proposes a multi-layer multi-reference prediction framework for effective video compression. Current AOM/AV1 baseline uses three reference frames for the inter prediction of each video frame. This paper first presents a new coding tool that extends the total number of reference frames in both forward and backward prediction directions. A multi-layer framework is then described, which suggests the encoder design and places different reference frames within one Golden Frame (GF) group to different layers. The multi-layer framework leverages the existing coding tools in the AV1 baseline, including the tool of "show_existing_frame" and the reference frame buffer update module of a wide flexibility. The use of extended ALTREF_FRAMEs is proposed, and multiple ALTREF_FRAME candidates are selected and widely spaced within one GF group. ALTREF_FRAME is a constructed, no-show reference obtained through temporal filtering of a look-ahead frame. In the multi-layer structure, one reference frame may serve different roles for the encoding of different frames through the virtual index manipulation. The experimental results have been collected over several video test sets of various resolutions and characteristics both texture- and motion-wise, which demonstrate that the proposed approach achieves a consistent coding gain compared to the AV1 baseline. For instance, using PSNR as the distortion metric, an average bitrate saving of 5.57+% in BDRate is obtained for the CIF-level resolution set, some of which has a gain of up to 13+%, and 4.47% on average for the VGA-level resolution set, some of which up to 18+%. Zoe Liu, Debargha Mukherjee, Jingning Han, Paul Wilkins, Yaowu Xu, Kenneth Rose |
DCC | 4 |
| 2018 | A Motion Vector Entropy Coding Scheme Based on Motion Field Referencing for Video CompressionabstractVideo codec exploits the temporal correlations in video signal through block-based motion compensated prediction. The motion vector associated with each prediction unit needs to be coded in the bit-stream. A differential coding scheme that employs the motion information from spatial neighbors and collocated blocks in the reference frames to predict the current motion vector is commonly used. Its efficacy is largely limited to track consistent or slow motion activities. A linear projection model is proposed in this work to create a motion field estimation that is capable to capture motion trajectory with high velocity. The resulting motion field motion vectors (MFMV) are fed into a dynamic motion vector referencing system as candidates in addition to those obtained from the spatial neighboring blocks. It allows the codec to closely track complex motion activities that the spatial neighbors or collocated motion vector referencing system usually fail to keep synchronized with. The MFMV system improves the prediction quality of the motion vectors and substantially reduces the energy in the difference motion vector for entropy coding, which translates into considerable compression performance improvements, especially for video sequences that contain complex motion activities. A number of design considerations to make the computation efficiency in both hardware and software platforms practical for production are discussed. Jingning Han, Yaowu Xu, Jim Bankoski |
ICIP | 1 |
| 2018 | A Hybrid Weighted Compound Motion Compensated Prediction for Video CompressionabstractCompound motion compensated prediction that combines reconstructed reference blocks to exploit the temporal correlation is a major component in the hierarchical coding scheme. A uniform combination that applies equal weights to reference blocks regardless of distances towards the current frame is widely employed in mainstream codecs. Linear distance weighted combination, while reflecting the temporal correlation, is likely to ignore the quantization noise factor and hence degrade the prediction quality. This work builds on the premise that the compound prediction mode effectively embeds two functionalities - exploiting temporal correlation in the video signal and canceling the quantization noise from reference blocks. A modified distance weighting scheme is introduced to optimize the trade-off between these two factors. It quantizes the weights to limit the minimum contribution from both reference blocks for noise cancellation. We further introduces a hybrid scheme allowing the codec to switch between the proposed distance weighted compound mode and the averaging mode to provide more flexibility for the trade-off between temporal correlation and noise cancellation. The scheme is implemented in the AV1 codec as part of the syntax definition. It is experimentally demonstrated to provide on average 1.5% compression gains across a wide range of test sets. Jingning Han, Yaowu Xu |
PCS | 2 |
| 2018 | An Overview of Core Coding Tools in the AV1 Video CodecabstractAV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz |
PCS | 3 |
| 2017 | A constrained adaptive scan order approach to transform coefficient entropy codingabstractTransform coefficient coding is a key module in modern video compression systems. Typically, a block of the quantized coefficients are processed in a pre-defined zig-zag order, starting from DC and sweeping through low frequency positions to high frequency ones. Correlation between magnitudes of adjacent coefficients is exploited via context based probability models to improve compression efficiency. Such scheme is premised on the assumption that spatial transforms compact energy towards lower frequency coefficients, and the scan pattern that follows a descending order of the likelihood of coefficients being non-zero provides more accurate probability modeling. However, a pre-defined zig-zag pattern that is agnostic to signal statistics may not be optimal. This work proposes an adaptive approach to generate scan pattern dynamically. Unlike prior attempts that directly sort a 2-D array of coefficient positions according to the appearance frequency of non-zero levels only, the proposed scheme employs a topological sort that also fully accounts for the spatial constraints due to the context dependency in entropy coding. A streamlined framework is designed for processing both intra and inter prediction residuals. This generic approach is experimentally shown to provide consistent coding performance gains across a wide range of test settings. Ching-Han Chiang, Jingning Han, Yaowu Xu |
ICASSP | 2 |
| 2017 | Adaptive interpolation filter scheme in AV1abstractVideo codecs heavily depend on sub-pixel level motion compensation to achieve superior compression performance. Interpolation filters with both anti-aliasing and denoising properties play a critical role in producing high quality prediction at sub-pixel positions. Prior research has developed many adaptive filtering schemes to improve the prediction precision for compression gains. On the other hand, such filtering operations require intense computation and may lead to scattered cache footprints, therefore, account for a major portion of the overall decoding cost in both software and hardware implementations. An adaptive interpolation filtering scheme is proposed in this work to optimize the trade off between prediction quality and decoding performance. It employs a separable model and selects filter kernels independently for horizontal and vertical directions to better capture statistical variations. In order to obtain sharper transition and reduce the ripple effect in the passband in frequency domain, a 12-tap filter is introduced in conjunction with a complimentary operation design that minimizes its impact on the decoding performance. The scheme achieves on average 1.3% coding gains across a wide range of test settings, with fairly limited additional hardware cost. Ching-Han Chiang, Jingning Han, Stan Vitvitskyy, Debargha Mukherjee, Yaowu Xu |
ICIP | 2 |
| 2017 | A level-map approach to transform coefficient codingabstractTransform coding is widely used in the video and image codec to largely remove the spatial correlation. The magnitude of transform coefficient is weakly correlated to a number of factors, including its frequency band, the neighboring coefficient magnitudes, luma/chroma planes, etc. To exploit such correlations for efficient entropy coding, one would build a probability model conditioned on the available contexts. However, the interaction of these factors creates a high dimensional space, a direct use of which would easily fall into the over-fitting problem. How to construct a compact context set which effectively captures the underlying correlations remains a major challenge in video and image compression. Prior research work primarily relies on bucketizing the previously coded coefficients into a small number of categories as the context model for next coefficient. Certain information loss is inevitable due to the classification process. To fully exploit the available context in a limited model space, a level map approach is proposed in this work. It decomposes the coding of coefficient magnitudes into consecutive runs of binary map coding, each corresponds to whether a coefficient is equal to or greater than the given level. Under the Markov assumption across the levels, nearly all the reference symbols available to each level map can be approximated as binary random variables. It hence allows the context model to account for all the surrounding coefficients information provided by the lower level maps, while retaining a reasonably compact size. Experimental evidence demonstrates that the proposed coding scheme provides considerable compression performance gains consistently over a large test settings. Jingning Han, Ching-Han Chiang, Yaowu Xu |
ICIP | 1 |
| 2016 | A staircase transform coding scheme for screen content video codingabstractDemand for screen content videos that contain computer generated text and graphics is growing. They are very different from natural videos, because they include much sharper edge transitions and very repetitive patterns. On this type of material, the efficacy of the conventional discrete cosine transform (DCT) is questionable because it relies on the assumption that a Gauss-Markov model leads to a base-band signal. However, the assumption may not hold true for screen content material. This work exploits a class of staircase transforms. Unlike the DCT whose bases are samplings of sinusoidal functions, the staircase transforms have their bases sampled from staircase functions, which better approximate the sharp transitions often encountered in the context of screen content. The staircase transform is integrated into a hybrid transform coding scheme, in conjunction with DCT. It is experimentally shown that the proposed approach provides an average of 2.9% compression performance gains in terms of BD-rate reduction. A perceptual comparison further demonstrates that the use of staircase transform achieves substantial reduction in ringing artifact due to the Gibbs phenomenon. Jingning Han, Yaowu Xu, Jim Bankoski |
ICIP | 2 |
| 2016 | A dynamic motion vector referencing scheme for video codingabstractVideo codecs exploit temporal redundancy in video signals, through the use of motion compensated prediction, to achieve superior compression performance. The coding of motion vectors takes a large portion of the total rate cost. Prior research utilizes the spatial and temporal correlation of the motion field to improve the coding efficiency of the motion information. It typically constructs a candidate pool composed of a fixed number of reference motion vectors and allows the codec to select and reuse the one that best approximates the motion of the current block. This largely disconnects the entropy coding process from the block's motion information, and throws out any information related to motion consistency, leading to sub-optimal coding performance. An alternative motion vector referencing scheme is proposed in this work to fully accommodate the dynamic nature of the motion field. It adaptively extends or shortens the candidate list according to the actual number of available reference motion vectors. The associated probability model accounts for the likelihood that an individual motion vector candidate is used. A complementary motion vector candidate ranking system is also presented here. It is experimentally shown that the proposed scheme achieves about 1.6% compression performance gains on a wide range of test clips. Jingning Han, Yaowu Xu, Jim Bankoski |
ICIP | 1 |
| 2015 | An estimation-theoretic approach to video denoiseingabstractA novel denoising scheme is proposed to fully exploit the spatio-temporal correlations of the video signal for efficient enhancement. Unlike conventional pixel domain approaches that directly connect motion compensated reference pixels and spatially neighboring pixels to build statistical models for noise filtering, this work first removes spatial correlations by applying transformations to both pixel blocks and performs estimation in the frequency domain. It is premised on the realization that the precise nature of temporal dependencies, which is entirely masked in the pixel domain by the statistics of the dominant low frequency components, emerges after signal decomposition and varies considerably across the spectrum. We derive an optimal non-linear estimator that accounts for both motion compensated reference and the noisy observations to resemble the original video signal per transform coefficient. It departs from other transform domain approaches that employ linear filters over a sizable reference set to reduce the uncertainty due to the random noise term. Instead it jointly exploits this precise statistical property appeared in the transform domain and the noise probability model in an estimation-theoretic framework that works on a compact support region. Experimental results provide evidence for substantial denoising performance improvement. Jingning Han, Timothy Kopp, Yaowu Xu |
ICIP | 1 |
| 2014 | Joint inter-intra prediction based on mode-variant and edge-directed weighting approaches in video codingabstractMost modern video compression codecs, like VP9, HEVC and H.264, encode square or rectangular blocks either by inter prediction or intra prediction. A joint inter-intra predictor that combines motion compensation and intra extrapolation by two novel weighting schemes is proposed to improve compression quality. Prior work on joint prediction employs inter-intra weights that only rely on the pixel locations. As an enhancement, we design a weighting approach by also considering the angle of intra prediction, which is the actual direction that the intra prediction errors evolve. Moreover, our second approach, inspired by prior work on geometric-partition-based motion compensation, breaks the limitation of traditional quad-tree partition by jointly using different predictors that implies soft step weighting functions for new and existing objects co-occurring around irregular motion edges. The proposed joint prediction approaches deliver consistent coding gains, as shown by extensive experiments on the experimental branch of VP9, Google's open source video compression tool. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Kenneth Rose |
ICASSP | 3 |
| 2014 | A pre-filtering approach to exploit decoupled prediction and transform block structures in video codingabstractRecent video coding techniques allow for decoupling of the transform block partition from that employed for prediction. For example, HEVC allows a transform block to overlap multiple prediction blocks. This paper is premised on the observation that in order to truly realize the potential of such enhanced flexibility, it is necessary to account for and mitigate considerable side effects due to stitching together independently predicted blocks, including the emergence of spurious high frequency components from sharp transitions across boundaries, which undermine the transform efficacy. The proposed solution involves an appropriately designed pre-filtering approach to mitigate boundary transition effects whenever a transform spans data from multiple prediction blocks. Moreover, this filtering technique enables extending the flexibility in decoupling prediction and transform structures, as various restrictions may now be eliminated. In particular, it makes it possible and beneficial to allow a transform block to span residual data from both inter and intra predicted blocks, whereas HEVC necessarily forces a single type of prediction in each coding unit. The method is further extended to include motion refinement that accounts for the pre-filtering approach. Experiments provide evidence for consistent coding gains over HEVC and VP9. Yue Chen 0040, Kenneth Rose, Jingning Han, Debargha Mukherjee |
ICIP | 3 |
| 2014 | Rate-distortion optimization and adaptation of intra prediction filter parametersabstractConventional “pixel copying” prediction used in current video standards was shown in previous work to be sub-optimal compared to 2-D non-separable Markov model based recursive extrapolation approaches. The premise of this paper is that in order to achieve the full potential of these approaches it is necessary to account for several requirements, namely, the design of prediction modes (and respective extrapolation filters) must optimize a rate-distortion cost rather than minimize the mean squared prediction error; the filters must be of sufficient complexity to cover all necessary directions; and the approach must include adaptation to available information indicative of local statistics. Hence, the proposed system employs four-tap recursive extrapolation filters that can predict from all standard directions, combined with a filter design method that accounts for the overall rate-distortion cost in conjunction with the codec decisions, along with adaptation of filter coefficients to relevant local information provided by encoder decisions on target bit rate and block size. Experimental evidence is provided for substantial coding gains over conventional intra coding. Shunyao Li, Jingning Han, Tejaswi Nanjundaswamy, Kenneth Rose |
ICIP | 3 |
| 2014 | An Estimation-Theoretic Framework for Spatially Scalable Video CodingabstractThis paper focuses on prediction optimality in spatially scalable video coding. It draws inspiration from an estimation-theoretic prediction framework for quality (SNR) scalability earlier developed by our group, which achieved optimality by fully accounting for relevant information from the current base layer (e.g., quantization intervals) and the enhancement layer, to efficiently calculate the conditional expectation that forms the optimal predictor. It was central to that approach that all layers reconstruct approximations to the same original transform coefficient. In spatial scalability, however, the layers encode different resolution versions of the signal. To approach optimality in enhancement layer prediction, this paper departs from existing spatially scalable codecs that employ pixel domain resampling to perform interlayer prediction. Instead, it incorporates a transform domain resampling technique that ensures that the base layer quantization intervals are accessible and usable at the enhancement layer despite their differing signal resolutions, which in conjunction with prior enhancement layer information, enable optimal prediction. A delayed prediction approach that complements this framework for spatial scalable video coding is then provided to further exploit future base layer frames for additional enhancement layer coding performance gains. Finally, a low-complexity variant of the proposed estimation-theoretic prediction approach is also devised, which approximates the conditional expectation by switching between three predictors depending on a simple condition involving information from both layers, and which retains significant performance gains. Simulations provide experimental evidence that the proposed approaches substantially outperform the standard scalable video codec and other leading competitors. Jingning Han, Vinay Melkote, Kenneth Rose |
IEEE Trans. Image Process. | 1 |
| 2013 | A recursive extrapolation approach to intra prediction in video codingabstractA novel intra prediction scheme, based on recursive extrapolation filters, is introduced. Standard intra prediction largely consists of copying boundary pixels (or linear combinations thereof) along certain directions, which reflects an overly simplistic model for the underlying spatial correlations. As an alternative, we view the image signal as a 2-D non-separable Markov model, whose corresponding correlation model better captures the nuanced directionality effects within blocks. This viewpoint motivates the design of a set of prediction modes represented by three-tap extrapolation filters, which replace the standard “pixel-copying” prediction modes, and provide efficient prediction at modest complexity. The Markov property is exploited by recursive predictions from nearest neighbors without recourse to simplistic separability assumptions, and while effectively accounting for correlation decay with distance from available boundary pixels. Coefficients for the set of mode filters are first trained by an efficient “k-modes” iterative technique designed to monotonically decrease the mean squared prediction error, and are then adjusted to directly optimize the overall rate-distortion objective. This prediction scheme complements the hybrid (cosine and sine) transform coding approach developed by our group, to achieve consistent coding gains, as shown for standard and commercial intra coders such as H.264/AVC and VP8. Jingning Han, Kenneth Rose |
ICASSP | 2 |
| 2013 | Approaching optimality in spatially scalable video coding: From resampling and prediction to quantization and entropy codingabstractThis paper builds on our recent work on optimal prediction in spatially scalable video coding, and is inspired by earlier work in our lab on optimal approaches for quality (or SNR) scalability. The approach we propose herein complements the optimal enhancement-layer prediction, enabled by transform domain resampling that ensures the base layer information is maximally accessible and usable at the enhancement layer despite their differing signal resolutions, with an optimal approach to quantization and entropy coding that exploits all available information, encapsulated in the appropriate conditional distribution for transform coefficients, to yield a unified coding engine for spatial scalability. For such quantizers to fully exploit base layer information, the enhancement layer transform block size must proportionally match the signal block transformed at the base layer. The overall system incorporates switching that applies the full estimation-theoretic quantizer and entropy coder at the right block size, but may optionally employ other block sizes where it defaults to optimal prediction followed by standard quantization. It is experimentally shown that the proposed scheme provides considerable performance gains over conventional codec and other leading competitors. Jingning Han, Kenneth Rose |
ICIP | 1 |
| 2013 | A joint spatio-temporal filtering approach to efficient prediction in video compressionabstractA novel filtering approach that naturally combines information from both intra-frame and motion compensated referencing for efficient prediction is proposed to fully exploit the spatio-temporal correlations of video signals, thereby achieving superior compression performance. Inspiration was drawn from our recent work on extrapolation filter based intra prediction, which views the spatial signal as a non-separable first-order Markov process and employs a 3-tap recursive filter to effectively capture the statistical characteristics. This work significantly extends the scope to further incorporate motion compensated reference in a filtering framework, whose coefficients were optimized via a “k-modes”-like iteration that accounts for various factors in the compression process including variation in statistics in the prediction loop, to minimize the rate-distortion cost. Experiments validate the efficacy of the proposed spatio-temporal approach, which translates into consistent coding performance gains. Jingning Han, Tejaswi Nanjundaswamy, Kenneth Rose |
PCS | 2 |
| 2013 | A butterfly structured design of the hybrid transform coding schemeabstractThe hybrid transform coding scheme that alternates amongst the asymmetric discrete sine transform (ADST) and the discrete cosine transform (DCT) depending on the boundary prediction conditions, is an efficient tool for video and image compression. It optimally exploits the statistical characteristics of prediction residual, thereby achieving significant coding performance gains over the conventional DCT-based approach. A practical concern lies in the intrinsic conflict between transform kernels of ADST and DCT, which prevents a butterfly structured implementation for parallel computing. Hence the hybrid transform coding scheme has to rely on matrix multiplication, which presents a speed-up barrier due to under-utilization of the hardware, especially for larger block sizes. In this work, we devise a novel ADST-like transform whose kernel is consistent with that of DCT, thereby enabling butterfly structured computation flow, while largely retaining the performance advantages of hybrid transform coding scheme in terms of compression efficiency. A prototype implementation of the proposed butterfly structured hybrid transform coding scheme is available in the VP9 codec repository. Jingning Han, Yaowu Xu, Debargha Mukherjee |
PCS | 1 |
| 2013 | The latest open-source video codec VP9 - An overview and preliminary resultsabstractGoogle has recently finalized a next generation open-source video codec called VP9, as part of the libvpx repository of the WebM project (http://www.webmproject.org/). Starting from the VP8 video codec released by Google in 2010 as the baseline, various enhancements and new tools were added, resulting in the next-generation VP9 bit-stream. This paper provides a brief technical overview of VP9 along with comparisons with other state-of-the-art video codecs H.264/AVC and HEVC on standard test sets. Results show VP9 to be quite competitive with mainstream state-of-the-art codecs. Debargha Mukherjee, Jim Bankoski, Adrian Grange, Jingning Han, John Koleszar, Paul Wilkins, Yaowu Xu, Ronald Bultje |
PCS | 4 |
| 2013 | Estimation-Theoretic Approach to Delayed Decoding of Predictively Encoded Video SequencesabstractCurrent video coders employ predictive coding with motion compensation to exploit temporal redundancies in the signal. In particular, blocks along a motion trajectory are modeled as an auto-regressive (AR) process, and it is generally assumed that the prediction errors are temporally independent and approximate the innovations of this process. Thus, zero-delay encoding and decoding is considered efficient. This paper is premised on the largely ignored fact that these prediction errors are, in fact, temporally dependent due to quantization effects in the prediction loop. It presents an estimation-theoretic delayed decoding scheme, which exploits information from future frames to improve the reconstruction quality of the current frame. In contrast to the standard decoder that reproduces every block instantaneously once the corresponding quantization indices of residues are available, the proposed delayed decoder efficiently combines all accessible (including any future) information in an appropriately derived probability density function, to obtain the optimal delayed reconstruction per transform coefficient. Experiments demonstrate significant gains over the standard decoder. Requisite information about the source AR model is estimated in a spatio-temporally adaptive manner from a bit-stream conforming to the H.264/AVC standard, i.e., no side information needs to be sent to the decoder in order to employ the proposed approach, thereby compatibility with the standard syntax and existing encoders is retained. Jingning Han, Vinay Melkote, Kenneth Rose |
IEEE Trans. Image Process. | 1 |
| 2012 | An estimation-theoretic approach to spatially scalable video codingabstractThis paper focuses on prediction optimality in spatially scalable video coding. It is inspired by the earlier estimation-theoretic prediction framework developed by our group for quality (SNR) scalability, which achieved optimality by fully accounting for relevant information from the current base layer (e.g., quantization intervals) and the enhancement layer, to efficiently calculate the conditional expectation that forms the optimal predictor. It was central to that approach that all layers reconstruct approximations to the same original transform coefficient. In spatial scalability, however, the layers encode different resolution versions of the signal. To approach optimality in enhancement layer prediction, the current work departs from existing spatially scalable codecs that employ pixel-domain resampling to perform inter-layer prediction. Instead, it incorporates a transform-domain resampling technique that ensures that the base layer quantization intervals are accessible and usable at the enhancement layer, which in conjunction with prior enhancement layer information, enable optimal prediction. Simulations provide experimental evidence that the proposed approach achieves substantial enhancement layer coding gains over the standard. Jingning Han, Vinay Melkote, Kenneth Rose |
ICASSP | 1 |
| 2012 | Towards predictor, quantizer and entropy coder optimality in scalable video codingabstractA novel coding paradigm is proposed to jointly optimize the prediction, quantization, and entropy coding modules, thereby approaching optimality in scalable video coding. It departs from conventional video coding schemes that consider prediction, transformation, quantization, and entropy coding, as largely separate sequential functional components. The method draws inspiration from an early estimation-theoretic approach, developed by our group for enhancement layer prediction, which efficiently combines all the information available to the enhancement layer coder, to produce the optimal prediction. The framework is significantly expanded here to also incorporate optimization of entropy-constrained quantization and arithmetic coding, while fully accounting for hitherto ignored relevant factors, inherent to predictive scalable coding, including information from the base layer quantization operation, and from the enhancement layer motion compensated reference. Experimental evidence is provided for substantial coding gains over conventional scalable video coding. Jingning Han, Kenneth Rose |
ICIP | 1 |
| 2012 | A Unified Estimation-Theoretic Framework for Error-Resilient Scalable Video CodingabstractA novel scalable video coding (SVC) scheme is proposed for video transmission over loss networks, which builds on an estimation-theoretic (ET) framework for optimal prediction and error concealment, given all available information from both the current base layer and prior enhancement layer frames. It incorporates a recursive end-to-end distortion estimation technique, namely, the spectral coefficient-wise optimal recursive estimate (SCORE), which accounts for all ET operations and tracks the first and second moments of decoder reconstructed transform coefficients. The overall framework enables optimization of ET-SVC systems for transmission over lossy networks, while accounting for all relevant conditions including the effects of quantization, channel loss, concealment, and error propagation. It thus resolves longstanding difficulties in combining truly optimal prediction and concealment with optimal end-to-end distortion and error-resilient SVC coding decisions. Experiments demonstrate that the proposed scheme offers substantial performance gains over existing error-resilient SVC systems, under a wide range of packet loss and bit rates. Jingning Han, Vinay Melkote, Kenneth Rose |
ICME | 1 |
| 2012 | Jointly Optimized Spatial Prediction and Block Transform for Video and Image CodingabstractThis paper proposes a novel approach to jointly optimize spatial prediction and the choice of the subsequent transform in video and image compression. Under the assumption of a separable first-order Gauss-Markov model for the image signal, it is shown that the optimal Karhunen-Loeve Transform, given available partial boundary information, is well approximated by a close relative of the discrete sine transform (DST), with basis vectors that tend to vanish at the known boundary and maximize energy at the unknown boundary. The overall intraframe coding scheme thus switches between this variant of the DST named asymmetric DST (ADST), and traditional discrete cosine transform (DCT), depending on prediction direction and boundary information. The ADST is first compared with DCT in terms of coding gain under ideal model conditions and is demonstrated to provide significantly improved compression efficiency. The proposed adaptive prediction and transform scheme is then implemented within the H.264/AVC intra-mode framework and is experimentally shown to significantly outperform the standard intra coding mode. As an added benefit, it achieves substantial reduction in blocking artifacts due to the fact that the transform now adapts to the statistics of block edges. An integer version of this ADST is also proposed. Jingning Han, Ankur Saxena, Vinay Melkote, Kenneth Rose |
IEEE Trans. Image Process. | 1 |
| 2011 | A spectral approach to recursive end-to-end distortion estimation for sub-pixel motion-compensated video codingabstractError resilient video coding critically relies on the accuracy of end to-end distortion estimation. An established solution, the recursive optimal per-pixel estimate (ROPE), is based on tracking the first and second moments of the decoder reconstructed pixels. This paper is focused on an alternative estimation approach, the spectral coefficient-wise optimal recursive estimate (SCORE), whose recursion is performed in the transform domain. The SCORE formulation is extended to derive a new technique for effective end-to-end distortion estimation, which accounts for sub-pixel motion compensation. Specifically, this technique exploits properties of the transform, such as coefficient de-correlation and energy compaction, to overcome ROPE's remaining shortcoming due to the proliferation of cross-correlation terms requiring excessive complexity or relatively crude approximations. Experiments show that the accuracy of SCORE matches ROPE in the full-pixel motion compensation setting, where ROPE is known to be optimal. More importantly, in the problematic setting of sub-pixel motion compensation, SCORE substantially outperforms ROPE and yields highly accurate distortion estimation. Jingning Han, Vinay Melkote, Kenneth Rose |
ICASSP | 1 |
| 2011 | A unified framework for spectral domain prediction and end-to-end distortion estimation in scalable video codingabstractA novel scalable coding approach is proposed for video transmission over lossy networks, which builds on two estimation-theoretic (ET) paradigms previously developed by our group: (1) an ET approach to enhancement layer prediction in scalable video coding (ET-SVC) that optimally combines all available information from both the current base layer and prior enhancement layer frames, and (2) the spectral coefficient-wise optimal recursive estimate (SCORE) of end-to-end distortion. SCORE provides the encoder with an estimate of distortion per decoder-reconstructed transform coefficient, accounting for the effects of quantization, concealment, packet loss and error propagation via the prediction loop. The current work significantly extends the scope of SCORE to encompass the setting of ET-SVC, whose prediction involves non-linear operations. This advance enables optimization of ET-SVC systems for transmission over lossy networks, thereby combining optimal prediction with optimal mode decisions at the enhancement layer. Experiments first demonstrate the estimation accuracy of SCORE in the settings of the ET-SVC coder. They then show considerable gains when SCORE is incorporated into ET-SVC to optimize encoding decisions under a wide range of packet loss and bit rates. Jingning Han, Vinay Melkote, Kenneth Rose |
ICIP | 1 |
| 2011 | Transform-domain temporal prediction in video coding with spatially adaptive spectral correlationsabstractTemporal prediction in standard video coding is performed in the spatial domain, where each pixel block is predicted from a motion-compensated pixel block in a previously reconstructed frame. Such prediction treats each pixel independently and ignores underlying spatial correlations. In contrast, this paper proposes a paradigm for motion-compensated prediction in the transform domain, that eliminates much of the spatial correlation before individual frequency components along a motion trajectory are independently predicted. The proposed scheme exploits the true temporal correlations, that emerge only after signal decomposition, and vary considerably from low to high frequency. The scheme spatially and temporally adapts to the evolving source statistics via a recursive procedure to obtain the cross-correlation between transform coefficients on the same motion trajectory. This recursion involves already reconstructed data and precludes the need for any additional side-information in the bit-stream. Experiments demonstrate substantial performance gains in comparison with the standard codec that employs conventional pixel domain motion-compensated prediction. Jingning Han, Vinay Melkote, Kenneth Rose |
MMSP | 1 |
| 2010 | Estimation-Theoretic Delayed Decoding of Predictively Encoded Video SequencesabstractCurrent video coding schemes employ motion compensation to exploit the fact that the signal forms an auto-regressive process along the motion trajectory, and remove temporal redundancies with prior reconstructed samples via prediction. However, the decoder may, in principle, also exploit correlations with received encoding information of future frames. In contrast to current decoders that reconstruct every block immediately as the corresponding quantization indices are available, we propose an estimation-theoretic delayed decoding scheme which leverages quantization and motion information of one or more future frames to refine the reconstruction of the current block. The scheme, implemented in the transform domain, efficiently combines all available (including future) information in an appropriately derived conditional pdf, to obtain the optimal delayed reconstruction of each transform coefficient in the frame. Experiments demonstrate substantial gains over the standard H.264 decoder. The scheme learns the autoregressive model from information available to the decoder, and compatibility with the standard syntax and existing encoders is retained. Jingning Han, Vinay Melkote, Kenneth Rose |
DCC | 1 |
| 2010 | Towards jointly optimal spatial prediction and adaptive transform in video/image codingabstractThis paper proposes a new approach to combined spatial (Intra) prediction and adaptive transform coding in block-based video and image compression. Context-adaptive spatial prediction from available, previously decoded boundaries of the block, is followed by optimal transform coding of the prediction residual. The derivation of both the prediction and the adaptive transform for the prediction error, assumes a separable first-order Gauss-Markov model for the image signal. The resulting optimal transform is shown to be a close relative of the sine transform with phase and frequencies such that basis vectors tend to vanish at known boundaries and maximize energy at unknown boundaries. The overall scheme switches between the above sine-like transform and discrete cosine transform (per direction, horizontal or vertical) depending on the prediction and boundary information. It is implemented within the H.264/AVC intra mode, is shown in experiments to significantly outperform the standard intra mode, and achieve significant reduction of the blocking effect. Jingning Han, Ankur Saxena, Kenneth Rose |
ICASSP | 1 |
| 2010 | Transform-domain temporal prediction in video coding: Exploiting correlation variation across coefficientsabstractTemporal prediction in standard video coding is performed in the spatial domain, where each pixel is predicted from a motion-compensated reconstructed pixel in a prior frame. This paper is premised on the realization that such standard prediction treats each pixel independently and ignores underlying spatial correlations, while transform-domain prediction would eliminate much of the spatial correlation before signal components (transform coefficients) are independently predicted. Moreover, the true temporal correlations emerge after signal decomposition, and vary considerably from low to high frequency components. This precise nature of the temporal dependencies is entirely masked in spatial domain prediction by the high temporal correlation coefficient (ρ ≈ 1) imposed on all pixels by the dominant low frequency components. We derive optimal transform-domain per-coefficient predictors for three main settings: basic inter-frame prediction; bi-directional prediction; and enhancement-layer prediction in scalable coding. Experimental results provide evidence for substantial performance gains in all settings. Jingning Han, Vinay Melkote, Kenneth Rose |
ICIP | 1 |
| 2010 | Estimation-theoretic approach to delayed prediction in scalable video codingabstractScalable video coding (SVC) employs inter-frame prediction at the base and/or the enhancement layers. Since the base layer can be encoded/decoded independent of the enhancement layers, we consider here the potential gains when prediction at the enhancement layers is delayed to accumulate and incorporate additional future information from the base layer. We build on two basic estimation-theoretic (ET) approaches developed by our group: an ET approach for enhancement layer prediction that optimally combines current base layer with prior enhancement layer information, and our recent ET approach for delayed decoding. The proposed technique fully exploits all the available information from the base layer, including any future frame information, and past enhancement layer information. It achieves considerable gains over zero-delay techniques including both standard SVC, and SVC with optimal ET prediction (but with zero encoding delay). Jingning Han, Vinay Melkote, Kenneth Rose |
ICIP | 1 |