Bohan Li 0006

dblp:123/2549-6 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0003-2285-9572ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Transform Domain Block and Reference Frame Prediction with Adaptive Predictor Coefficients
abstract
Block prediction and reference frame generation play a central role in contemporary video codecs and hold significant potential for improving compression efficiency. For motion vectors of varying reliability, directly averaging the bi-directional reference blocks introduces noise in block prediction. In advanced video codecs such as AV2, the generation of co-located reference frames (CLRFs) within the temporal interpolated prediction (TIP) module is also performed at the block level. However, the imprecise vectors in the motion field cause the associated block pairs to exhibit weak correlation in their high-frequency components, which is not addressed by existing methods. This creates an opportunity for the transformdomain temporal prediction (TDTP) approach. This work proposes an adaptive TDTP framework for block prediction. Independent linear predictors (LPs) are trained for each AC coefficient to estimate the corresponding temporal correlation. A novel adaptive backward updating scheme for linear predictors is employed, in which reconstructed blocks are used to update the statistics and refine LP parameters without introducing additional bitrate overhead. Simulation results on the TIP module in AV2 demonstrate that the proposed approach effectively adapts to frame statistics and enhances the efficiency of compound inter-prediction.
Bohan Li 0006, Dharmesh Mohanraj, Kenneth Rose
DCC2
2023 Multi-Rate Adaptive Transform Coding for Video Compression
abstract
Contemporary lossy image and video coding standards rely on transform coding, the process through which pixels are mapped to an alternative representation to facilitate efficient data compression. Despite impressive performance of end-to-end optimized compression with deep neural networks, the high computational and space demands of these models has prevented them from superseding the relatively simple transform coding found in conventional video codecs. In this study, we propose learned transforms and entropy coding that may either serve as (non)linear drop-in replacements, or enhancements for linear transforms in existing codecs. These transforms can be multi-rate, allowing a single model to operate along the entire rate-distortion curve. To demonstrate the utility of our framework, we augmented the DCT with learned quantization matrices and adaptive entropy coding to compress intra-frame AV1 block prediction residuals. We report substantial BD-rate and perceptual quality improvements over more complex nonlinear transforms at a fraction of the computational cost.
Lyndon R. Duong, Bohan Li 0006, Jingning Han
ICASSP2
2023 Learned Image Compression Guided Adaptive Quantization for Perceptual Quality
abstract
Neural network based image compression has made significant progress in recent years. The learned image codecs are commonly reported to outperform their conventional counterparts in perceptual quality. Despite the superior performance, the learned image codecs are much more complex to decode, which hinders their usage in practice. Without a significant advance in hardware capability, the conventional image codec will likely remain a primary component for large scale image services. It is therefore desirable to improve the quality of conventional image codecs. In this paper, we present an adaptive quantization approach to the conventional image codec with the help of learned image codecs to improve its perceptual quality. It exploits the bit allocation of the neural network based image codec to adapt the quantizers on a block basis. It is experimentally shown that the proposed method provides considerable perceptual quality improvements over other leading contenders.
Ruiqi Geng, Bohan Li 0006, Maryla Ustarroz-Calonge, Frank Galligan, Jingning Han, Yaowu Xu
ICIP3
2022 An Efficient Scheme of Multi-Hypothesis Motion Compensated Prediction for Video Coding Applications
abstract
Prior research has demonstrated that the multi-hypothesis motion compensated prediction (MCP) can theoretically provide a better prediction quality than single-reference MCP, thereby improving the compression efficiency in video coding. However, the existing multi-hypothesis MCP methods typically require either additional rate cost to transmit the motion vectors, or significant decoding complexity to conduct the motion search at the decoder end, which is usually expensive. In this work, we propose a novel scheme to materialize the multi-hypothesis MCP that requires no additional rate cost, nor extra motion search on either the encoder or decoder side. Various approaches to synthesize these available multiple references to form the inter prediction are presented. We experimentally demonstrate that the proposed scheme provides considerable and consistent coding gains across a wide range of operating points.
Bohan Li 0006, Jingning Han, Yaowu Xu
ICIP1
2021 Adaptive GOP Size Decision for Multi-Pass Video Coding Based on Hidden Markov Model
abstract
Multi-pass coding is a widely utilized technique to improve the compression efficiency in video coding, where frame statistics are collected from the previous passes and then analyzed to provide better encoder decisions, such as rate control parameters, prediction mode selection, motion estimation, etc. In this paper, a novel method to determine the size of each group of picture (GOP) using the multi-pass information is presented. In particular, we propose to categorize frames into regions with different natures, including stationary, high-variance, blending, and scene cut, through analyzing the frame statistics generated from the previous passes using a hidden Markov model. The GOP size is then determined based on the region types and the inter frame correlations. It is experimentally shown that the proposed adaptive GOP size decision provides considerable coding performance improvements over conventional fixed GOP length.
Bohan Li 0006, Jingning Han, Yaowu Xu
ICASSP1
2021 A Temporal Filtering Approach Based on Optical Flow Estimation for Video Coding
abstract
Video coding uses motion compensated prediction to exploit temporal correlations for compression efficiency. Prior works have demonstrated that substantial coding gains can be achieved by decomposing a long-term reference frame into a synthetic reference-only (non-displayable) frame and an overlay displayable frame that resembles the original frame. The source of the reference-only frame is typically generated by temporal filtering along the motion trajectories across nearby frames, where the motion trajectories are built using block matching algorithms (BMAs), thereby reducing the noise level within this synthetic frame. Noting that the efficacy of the conventional BMAs are limited to capturing translational motion activities, this paper proposes a novel approach that uses a per-pixel motion field generated by an optical flow estimation to form the motion trajectory for more efficient temporal filtering. It is experimentally shown that the proposed method better captures non-translational motion activities, which translates into considerable coding gains for video signals with such complicate motion patterns.
Bohan Li 0006, Lauren Partin, Jingning Han, Yaowu Xu
MMSP1
2021 A Technical Overview of AV1
abstract
The AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than a 30% reduction in bit rate compared to its predecessor VP9 for the same decoded video quality. This article provides a technical overview of the AV1 codec design that enables the compression performance gains with considerations for hardware feasibility.
Jingning Han, Bohan Li 0006, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, Yue Chen 0040, Yunqing Wang, Paul Wilkins, Yaowu Xu, Jim Bankoski
Proc. IEEE2
2020 An Adaptive Linear Estimator Based Approach to Bi-Directional Motion Compensated Prediction
abstract
Bi-directional motion compensated prediction is widely utilized in video coding. Conventionally, the encoder searches for two motion vectors pointing to reference frames in both directions, and transmits these motion vectors to the decoder. Recognizing that the two reference frames are already available to the decoder, prior work proposed decoder-side motion estimation to extract motion information or optical flow, at the cost of dramatic increase in decoder complexity. This paper proposes a novel bi-directional motion compensation mode that efficiently utilizes the motion information that is already available to the decoder, without recourse to extensive search. An estimation theory based approach is proposed and utilized to provide a high quality prediction, which adaptively combines contributions from multiple motion-compensated references. Experimental results show that the proposed method, while yielding a greatly reduced decoder side complexity, introduces a significant coding gain for a diverse set of video sequences.
Bohan Li 0006, Jingning Han, Kenneth Rose
ICASSP1
2020 Optical Flow Based Co-Located Reference Frame for Video Compression
abstract
This paper proposes a novel bi-directional motion compensation framework that extracts existing motion information associated with the reference frames and interpolates an additional reference frame candidate that is co-located with the current frame. The approach generates a dense motion field by performing optical flow estimation, so as to capture complex motion between the reference frames without recourse to additional side information. The estimated optical flow is then complemented by transmission of offset motion vectors to correct for possible deviation from the linearity assumption in the interpolation. Various optimization schemes specifically tailored to the video coding framework are presented to further improve the performance. To accommodate applications where decoder complexity is a cardinal concern, a block-constrained speed-up algorithm is also proposed. Experimental results show that the main approach and optimization methods yield significant coding gains across a diverse set of video sequences. Further experiments focus on the trade-off between performance and complexity, and demonstrate that the proposed speed-up algorithm offers complexity reduction by a large factor while maintaining most of the performance gains.
Bohan Li 0006, Jingning Han, Yaowu Xu, Kenneth Rose
IEEE Trans. Image Process.1
2018 Co-located Reference Frame Interpolation Using Optical Flow Estimation for Video Compression
abstract
The hierarchical coding structure that supports bi-directional motion compensated prediction is commonly used for video compression efficiency. Conventional approach directly seeks the reference pixel block from each individual reference frame and use it or its linear combinations for prediction. It largely ignores the motion information between these reference frames. To fully utilize all the information from the bi-directional reference frames, this work builds a per-pixel motion field that connects the two-sided reference frames using optical flow estimation. A reference frame is then interpolated at the current frame location. This collocated reference frame effectively accounts for the true motion trajectories in the video signal including both translational and the more complex non-translational motion models, which are beyond the capability of the conventional block-based motion compensated prediction. The scheme is experimentally shown to provide substantial compression performance gains. A number of optimization designs are proposed to make the codec complexity feasible while largely maintaining the coding performance.
Bohan Li 0006, Jingning Han, Yaowu Xu
DCC1
2017 On generalizing the estimation-theoretic framework to scalable video coding with quadtree structured block partitions
abstract
Scalable video coding suffers from the under-utilization of base layer information, where usually only the reconstruction in the base layer is used for enhancement layer prediction. Prior work from our lab proposed an optimal estimation-theoretic (ET) approach for quality scalable coding, wherein the estimates are obtained by utilizing all the available information from base layer quantization interval and enhancement layer distribution for transform coefficients. While this approach was proposed for fixed block size encoding, modern codecs employ variable block size quadtree structured partitioning, which results in different partitions at base layer and enhancement layer based on the rate-distortion trade-off, thus makes the base layer information not directly usable in the enhancement layer. Other new tools such as hybrid transform and the rate-distortion optimized quantizer (RDOQ) also have an impact on the information available for optimal estimation. In this paper, we generalize the ET framework for quality scalable video coding to account for the quadtree structured partitioning, hybrid transform and the RDOQ adjustment. Experimental evidence is provided for consistent coding gains over standard SHVC.
Shunyao Li, Tejaswi Nanjundaswamy, Bohan Li 0006, Kenneth Rose
ICIP3
2017 An error-resilient video coding framework with soft reset and end-to-end distortion optimization
abstract
Temporal prediction plays a crucial role in most video coding applications. However, due to error propagation via the prediction loop, it also increases the vulnerability to channel loss. The standard counter measure to mitigate error propagation is the `intra refresh' mode, which in effect resets temporal prediction to block error propagation, but at a significant rate overhead. This on/off switch for temporal prediction is overly crude to optimize the compression-resilience tradeoff. In this paper, we propose a novel framework that significantly expands the options available to counter error propagation by introducing optimally controlled soft resets, wherein intra and inter predictions are combined with adjustable weights to control the dependency on previous frames while accounting for the overall rate and distortion. Since the optimal control of such soft resets can only be achieved if the encoder can effectively estimate its impact on the end-to-end distortion (EED), we propose to extend the well known recursive optimal per-pixel estimation (ROPE) approach to accurately account for the soft reset mode, then optimize encoder mode decisions to minimize the estimated EED for the given rate. Experimental results show that the proposed framework achieves significant performance gains for video streaming over unreliable networks.
Bohan Li 0006, Tejaswi Nanjundaswamy, Kenneth Rose
ICIP1
2016 Block-size adaptive transform domain estimation of end-to-end distortion for error-resilient video coding
abstract
The accuracy of end-to-end distortion (EED) estimation is crucial to achieving effective error resilient video coding. An established solution, the recursive optimal per-pixel estimate (ROPE), does so by tracking the first and second moments of decoder-reconstructed pixels. An alternative estimation approach, the spectral coefficient-wise optimal recursive estimate (SCORE), tracks instead moments of decoder-reconstructed transform coefficients, which enables accounting for transform domain operations. However, the SCORE formulation relies on a fixed transform block size, which is incompatible with recent standards. This paper proposes a non-trivial generalization of the SCORE framework which, in particular, accounts for arbitrary block size combinations involving the current and reference block partitions. This seemingly intractable objective is achieved by a two-step approach: i) Given the fixed block size moments of a reference frame, estimate moments of transform coefficients for the codec-selected current block partition; ii) Convert the current results to transform coefficient moments corresponding to a regular fixed block size grid, to facilitate EED estimation for the next frame. Experimental results first demonstrate the accuracy of the proposed estimate in conjunction with transform domain temporal prediction. Then the estimate is leveraged to optimize the coding mode and yields considerable gains in rate-distortion performance.
Bohan Li 0006, Tejaswi Nanjundaswamy, Kenneth Rose
ICIP1