Xin Zhao 0003

dblp:68/2766-3 · DBLP profile ↗
← Back
44ranked-venue papers
16as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 16 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSystems, architecture and hardware · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Video Coding With Cross-Component Sample Offset
abstract
Beyond the exploration of traditional spatial, temporal and subjective visual signal redundancy in image and video compression, recent research has focused on leveraging cross-color component redundancy to enhance coding efficiency. Cross-component coding approaches are motivated by the statistical correlations among different color components, such as those in the Y'CbCr color space, where luma (Y) color component typically exhibits finer details than chroma (Cb/Cr) color components. Inspired by previous cross-component coding algorithms, this paper introduces a novel in-loop filtering approach named Cross-Component Sample Offset (CCSO). CCSO utilizes co-located and neighboring luma samples to generate correction signals for both luma and chroma reconstructed samples. It is a multiplication-free, non-linear mapping process implemented using a look-up-table. The input to the mapping is a group of reconstructed luma samples, and the output is an offset value applied on the center luma or co-located chroma sample. Experimental results demonstrate that the proposed CCSO can be applied to both image and video coding, resulting in improved coding efficiency and visual quality. The method has been adopted into an experimental next-generation video codec beyond AV1 developed by the Alliance for Open Media (AOMedia), demonstrating average -0.81% and -0.69% coding gain on PSNR and VMAF quality metric, respectively, under random access configuration. Additionally, CCSO notably improves the subjective visual quality.
Han Gao 0001, Xin Zhao 0003, Shan Liu 0001
IEEE Trans. Image Process.2
2024 Adaptive Online Learning of Separable Path Graph Transforms for Intra-Prediction
abstract
Current video coding standards, including H.264/AVC, HEVC, and VVC, employ discrete cosine transform (DCT), discrete sine transform (DST), and secondary Karhunen-Loéve transforms (KLTs) to decorrelate the intra-prediction residuals. However, the efficiency of these transforms in decorrelation can be limited when the signal has a non-smooth and non-periodic structure, such as those occurring in textures with intricate patterns. This paper introduces a novel adaptive separable path graph-based transform (GBT) that can provide better decorrelation than the DCT for intra-predicted texture data. The proposed GBT is learned in an online scenario with sequential$K$-means clustering, which groups similar blocks during encoding and decoding to adaptively learn the GBT for the current block from previously reconstructed areas with similar characteristics. A signaling overhead is added to the bitstream of each coding block to indicate the usage of the proposed graph-based transform. We assess the performance of this method combined with H.264/AVC intra-coding tools and demonstrate that it can significantly outperform H.264/AVC DCT for intra-predicted texture data.
Wen-Yang Lu, Eduardo Pavez, Antonio Ortega, Xin Zhao 0003, Shan Liu 0001
PCS4
2024 Improvements of the BD-Rate Metrics Using Monotonic Curve-Fitting Methods
abstract
The Bj⊘ntegaard Delta rate (BD-rate) measurements have been used as the primary metrics to evaluate performance of video codecs. However, current BD-rate calculation methods are only applicable under the condition that the rate-distortion (R-D) values maintain a monotonic relationship, as this prerequisite is essential for computing integral along the distortion axis. To address this limitation, we propose a curve-fitting based BD-rate solution that guarantees the reconstructed R-D curve to be monotonic. Considering different use cases, we provide a four parameters logistic curve and a constraint cubic curve to approximate the underlying R-D curve. Computation of BD-rate and BD-metric using fitted R-D curve are elaborated in detail. Experimental results indicate that the proposed solutions work well on non-monotonic data. Furthermore, we verified through quantitative analysis that curve-fitting solutions provide more precise measurements of coding efficiency compared to interpolation methods. This improved accuracy contributed by the proposed methods is attributed to the higher resilience to the inherent randomness present in observed data. The proposed method has been adopted by the MPEG WG4 VCM study group for standardization activities. The source code was released at https://multimedia.tencent.com/resources/tvd.
Haiqiang Wang, Xin Zhao 0003, Ding Ding 0004, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001
PCS2
2023 Adaptive Probability Estimation Techniques for Context Adaptive Arithmetic Coding
abstract
Context-adaptive arithmetic coding is an essential entropy coding scheme used in all modern video codecs. An arithmetic coder with binary symbol size is used by video codecs like H.264/AVC, HEVC and VVC while AV1 utilizes an arithmetic coder with syntax adaptive M-ary symbols. Recently, the Alliance for Open Media (AOMedia) has initiated exploration activities towards next-generation video coding tools beyond AV1. In this regard, improvements on probability estimation techniques for the context-adaptive M-ary arithmetic coder in AV1, are explored in this paper. The proposed improvements are applied and tested on top of the reference implementation of the exploratory codec beyond AV1, known as AVM (AOMedia Video Model). Experimental results show that, compared to AVM, the proposed method achieves an average 0.27%, 0.34% and 0.32% overall BD-rate coding gains for All Intra (AI), Random Access (RA) and Low Delay (LD) coding configurations for a wide range of video content.
Madhu Peringassery Krishnan, Xin Zhao 0003, Shan Liu 0001
ISCAS2
2023 Improved Chroma From Luma Intra Prediction Mode Beyond AV1
abstract
In AV1, Chroma from Luma (CfL) intra prediction mode is adopted to predict chroma samples by exploiting the linear correlation between the co-located samples of luma and chroma components, wherein the scaling factor of the linear model is transmitted to the decoder side and the offset factor is derived as the average of neighboring chroma pixels. In this paper, the CfL prediction mode is improved in following three aspects. Firstly, the calculation of DC contribution between luma and chroma is aligned to improve the accuracy of CfL prediction. Secondly, to further enhance the CfL prediction mode, a new cross-component intra prediction mode without signaling of scaling factor is employed. Thirdly, the down-sampling filter for CfL mode is adaptively selected at the encoder side for each video sequence. Simulation results show that, on top of AOMedia Video Model (AVM) v3.0.0, an average coding gain of 0.9%, 0.5%, and 0.4% in terms of YUV-PNSR is achieved for all intra, random access, and low delay configurations, respectively.
Xin Zhao 0003, Shan Liu 0001
ISCAS3
2023 Content Adaptive Weighted Prediction for Video Coding Beyond AV1
abstract
In AV1, the compound prediction supports three methods for averaging two prediction blocks. The weighting factors for two prediction blocks can be based on either the wedge mask, the difference between prediction samples, or predefined weighting factors. However, for single prediction, only one predefined weighting factor is used, which is suboptimal when there are illumination changes between reference frames and the current frame. To address this limitation, this paper proposes a content adaptive weighted prediction method. This approach aims to enhance the prediction accuracy for single prediction. It involves the use of multiple predefined weighting factor look-up tables, and the selection among these different look-up tables are implicitly determined based on coded information of the current block, and the index of scaling factor in the look-up table is signaled in bitstream and parsed at the decoder side. Experimental results show that the proposed method can achieve an average 0.31% and 0.24% coding gain in terms of BD-rate with random access and low delay configurations, respectively.
Xin Zhao 0003, Han Gao 0001, Shan Liu 0001
VCIP2
2022 Advanced Motion Vector Difference Coding Beyond AV1
abstract
In AV1, for inter coded blocks with compound reference mode, motion vector differences (MVDs) are signaled for reference frame list 0 or list 1 separately with the same MVD precision regardless of the motion vector magnitude. In this paper, two advanced MVD coding methods are proposed. Firstly, to reduce the overhead for signaling MVD, the precision of the MVD is implicitly determined based on the associated MV class and MVD magnitude. Secondly, a new inter prediction mode is added to explore the correlation of MVDs between two reference frames, wherein one joint MVD is signaled for two reference frames. Experimental results demonstrate that, in the random-access common test condition luma coding gains of around 1.1% in terms of BD-rate can be achieved on top of a recent release of AOMedia Video Model (AVM).
Xin Zhao 0003, Shan Liu 0001
ICIP2
2022 Unified Fast Partitioning Algorithm for Intra and Inter Predictions in Versatile Video Coding
abstract
The Versatile Video Coding (VVC) standard adopts a more flexible partitioning structure beyond the High Efficiency Video Coding (HEVC) standard. The quadtree with nested multi-type tree (QTMT) partitioning structure greatly improves coding efficiency. Nevertheless, the typical recursive coding unit (CU) partitioning scheme causes a substantial increase in computational complexity at the encoder. In this paper, a unified fast algorithm is proposed for both intra and inter predictions, which makes use of various historical information of previously checked partitions. The proposed algorithm is implemented on top of the VVC reference software VTM-14.0. Experimental results show that the proposed algorithm achieves an average 42.47%, 40.46%, and 40.38% encoder runtime reduction with only 0.45%, 1.08%, and 1.18% Bjøntegaard delta bitrate loss (BDBR) in All Intra (AI), Random Access (RA), and Low Delay P (LDP) configurations, respectively.
Wei Kuang, Xiang Li 0003, Xin Zhao 0003, Shan Liu 0001
PCS3
2022 An Open Video Dataset For Screen Content Coding
abstract
In recent years, screen content video is becoming increasingly popular in several major video applications, such as video recording and video conferencing. Due to the unique features of screen content videos that are not captured by camera sensors but produced artificially, dedicated coding tools have been developed for achieving significant compression efficiency gain. In recognition of the popularity of screen content applications, an open video dataset for screen content is proposed in this paper for the development of screen content coding technologies. The proposed video dataset consists of 12 typical screen content type video clips that are publicly available. In addition, to better understand the characteristics of the proposed video dataset, several major screen content coding tools in AOMedia Video 1 (AV1) have been evaluated on this dataset and analyzed in this paper.
Yingbin Wang, Xin Zhao 0003, Xiaozhong Xu, Shan Liu 0001, Zhijun Lei, Mariana Afonso, Andrey Norkin, Thomas Daede
PCS2
2021 Improved Intra Mode Coding Beyond Av1
abstract
In AOMedia Video 1 (AV1), directional intra prediction modes are applied to model local texture patterns that present certain directionality. Each intra prediction direction is represented with a nominal mode index and a delta angle. The delta angle is entropy coded using shared context between luma and chroma, and the context is derived using the associated nominal mode. In this paper, two methods are proposed to further reduce the signaling cost of delta angles: cross-component delta angle coding, and context-adaptive delta angle coding, whereby the cross-component and spatial correlation of the delta angles are explored, respectively. The proposed methods were implemented on top of a recent version of libaom. Experimental results show that the proposed cross-component delta angle coding achieved average 0.4% BD-rate reduction with 4% encoding time saving over all intra configurations. By combining both methods, an average 1.2% BD-rate reduction is achieved.
Yize Jin, Xin Zhao 0003, Shan Liu 0001, Alan C. Bovik
ICASSP3
2021 Context-Adaptive Secondary Transform For Video Coding
abstract
It is well-known that non-separable transforms can efficiently decorrelate arbitrarily directed textures that are often present in image and video content. Due to the computational complexity involved, it is usually applied as a secondary transform operating on low frequency primary transform coefficients. In order to represent a variety of arbitrary directional textures in natural images /videos, it is ideal to have sufficient coverage of secondary transform kernels for the codec to choose from. However, this may lead to increased signaling cost and encoder complexity. This paper proposes a context-adaptive secondary transform (CAST) kernel selection approach to enable the usage of more secondary transform kernels with no signaling cost increase and minimal encoder and decoder complexity increase. The proposed approach uses the variance of the top row and left column of reconstructed pixels adjacent to the transform block, if available, as a context for selecting the set of transform kernels. Experimental results show that, compared to libaom, the proposed algorithm achieves a luma BD-rate reduction of 2.17% and 3.11% for All Intra coding using PSNR and SSIM quality metrics, respectively.
Samruddhi Kahu, Madhu Peringassery Krishnan, Xin Zhao 0003, Shan Liu 0001
ICIP3
2021 Semi-Decoupled Partitioning for Video Coding Beyond AV1
abstract
Recently, the Alliance for Open Media (AOMedia) has initiated activities on exploring new coding tools with capabilities beyond AOMedia Video 1 (AV1). Among various coding modules within the conventional hybrid video coding structure, the block partitioning scheme builds the foundation of the codec and the related design needs to be sought out from the beginning. In this paper, a Semi-Decoupled Partitioning (SDP) method is proposed for coding block partitioning. With SDP, luma and chroma share the same coding block partitioning toward a specified partitioning depth. After this specified depth, the partitioning patterns of luma and chroma components can be optimized and signaled independently. The benefit of SDP is the additional flexibility of switching between dependent and independent partitioning patterns for luma and chroma since the characteristics of these color components can largely differ. The proposed method has been integrated on top of libaom research branch, and experimental results show that, significant coding gain can be achieved comparing to the libaom research anchor.
Xin Zhao 0003, Shan Liu 0001
ICIP2
2021 Study On Coding Tools Beyond AV1
abstract
The Alliance for Open Media has recently initiated coding tool exploration activities towards the next-generation video coding beyond AV1. With this regard, this paper presents a package of coding tools that have been investigated, implemented and tested on top of the codebase, known as libaom, which is used for the exploration of next-generation video compression tools. The proposed tools cover several technical areas based on a traditional hybrid video coding structure, including block partitioning, prediction, transform and loop filtering. The proposed coding tools are integrated as a package, and a combined coding gain over AV1 is demonstrated in this paper. Furthermore, to better understand the behavior of each tool, besides the combined coding gain, the tool-on and tool-off tests are also simulated and reported for each individual coding tool. Experimental results show that, compared to libaom, the proposed methods achieve an average 8.0% (up to 22.0%) overall BD-rate reduction for All Intra coding configuration a wide range of image and video content.
Xin Zhao 0003, Madhu Peringassery Krishnan, Yixin Du, Shan Liu 0001, Debargha Mukherjee, Yaowu Xu, Adrian Grange
ICME1
2021 Video Coding Tool Analysis and Dataset for Gaming Content
abstract
The gaming market has kept growing significantly in recent years. Driven by multiple technology advances, such as cloud computing and video technologies, new gaming applications, e.g. AR, VR and cloud gaming, are becoming more and more practical and popular. Among different types of gaming, the emergence of cloud gaming is driving the market with enhanced gamer experience as well as new challenges to the services. One of the key technological challenges of gaming applications is the video coding, which is the foundation of several popular gaming applications, including cloud gaming and game live streaming. Comparing to the typical camera captured content and screen content, gaming content presents unique features that directly lead to different preferences on the selection of coding tool sets. To better understand the behaviors of known video coding tools and provide test materials for research and development on future coding tools that benefits more on gaming content, in this paper, a dataset consists of a set of gaming video is proposed together with analysis of the performances of existing coding tools on these materials. It is observed that, several known coding tools are exceptionally beneficial for gaming content and the rational is analyzed in this paper.
Xin Zhao 0003, Shan Liu 0001, Xiang Li 0003, Guichun Li, Xiaozhong Xu
PCS1
2021 Cross-Component Sample Offset for Image and Video Coding
abstract
Existing cross-component video coding technologies have shown great potential on improving coding efficiency. The fundamental insight of cross-component coding technology is respecting the statistical correlations among different color components. In this paper, a Cross-Component Sample Offset (CCSO) approach for image and video coding is proposed inspired by the observation that, luma component tends to contain more texture, while chroma component is relatively smoother. The key component of CCSO is a non-linear offset mapping mechanism implemented as a look-up-table (LUT). The input of the mapping is the co-located reconstructed samples of luma component, and the output is offset values applied on chroma component. The proposed method has been implemented on top of a recent version of libaom. Experimental results show that the proposed approach brings 1.16% Random Access (RA) BD-rate saving on top of AV1 with marginal encoding/decoding time increase.
Yixin Du, Xin Zhao 0003, Shan Liu 0001
VCIP2
2021 Multicomponent Secondary Transform
abstract
The Alliance for Open Media has recently initiated coding tool exploration activities towards the next-generation video coding beyond AV1. In this regard, a frequency-domain coding tool, which is designed to leverage the cross-component correlation existing between collocated chroma blocks, is explored in this paper. The tool, henceforth known as multi-component secondary transform (MCST), is implemented as a low complexity secondary transform with primary transform coefficients of multiple color components as input. The proposed tool is implemented and tested on top of libaom. Experimental results show that, compared to libaom, the proposed method achieves an average 0.34% to 0.44% overall coding efficiency for All Intra (AI) coding configuration for a wide range of video content.
Madhu Peringassery Krishnan, Xin Zhao 0003, Shan Liu 0001
VCIP2
2021 Intra Prediction and Mode Coding in VVC
abstract
This paper presents the intra prediction and mode coding of the Versatile Video Coding (VVC) standard. This standard was collaboratively developed by the Joint Video Experts Team (JVET). It follows the traditional architecture of a hybrid block-based codec that was also the basis of previous standards. Almost all intra prediction features of VVC either contain substantial modifications in comparison with its predecessor H.265/HEVC or were newly added. The key aspects of these tools are the following: 65 angular intra prediction modes with block shape-adaptive directions and 4-tap interpolation filters are supported as well as the DC and Planar mode, Position Dependent Prediction Combination is applied for most of these modes, Multiple Reference Line Prediction can be used, an intra block can be further subdivided by the Intra Subpartition mode, Matrix-based Intra Prediction is supported, and the chroma prediction signal can be generated by the Cross Component Linear Model method. Finally, the intra prediction mode in VVC is coded separately for luma and chroma. Here, a Most Probable Mode list containing six modes is applied for luma. The individual compression performance of tools is reported in this paper. For the full VVC intra codec, a bitrate saving of 25% on average is reported over H.265/HEVC using an objective metric. Significant subjective benefits are illustrated with specific examples.
Jonathan Pfaff, Alexey Filippov, Shan Liu 0001, Xin Zhao 0003, Jianle Chen, Santiago De-Luxán-Hernández, Thomas Wiegand 0001, Vasily Rufitskiy, Adarsh K. Ramasubramonian, Geert Van der Auwera
IEEE Trans. Circuits Syst. Video Technol.4
2021 Fast DST-VII/DCT-VIII With Dual Implementation Support for Versatile Video Coding
abstract
The Joint Video Exploration Team (JVET) recently launched the standardization of the next-generation video coding named Versatile Video Coding (VVC) with the inherited technical framework from its predecessor High-Efficiency Video Coding (HEVC). The simplified Enhanced Multiple Transform (EMT) has been adopted as the primary residual coding transform solution, termed Multiple Transform Selection (MTS). In MTS, only the transform set consisting of DST-VII and DCT-VIII remains, excluding the other transform sets and the dependency on intra prediction modes. Significant coding gains are achieved by introducing new DST/DCT transforms, but the full matrix implementation is relatively costly compared to partial butterfly in terms of both software run-time and operation counts. In this work, we exploit the inherent features existing in DST-VII and DCT-VIII. Instead of repeating the element-wise additions and multiplications in full matrix operation, these features can be leveraged to achieve more efficient implementations which only use partial elements to derive the identical results. Existing transform matrices are further tuned to utilize these (anti-)symmetric features. A partial butterfly-type fast algorithm with dual-implementation support is proposed for DST-VII/DCT-VIII transform in VVC. Complexity analysis including operation counts and software run-time are conducted to validate the effectiveness. In addition, we prove the features are perfectly supported by theory. The proposed fast methods achieve noticeable software run-time savings without compromising on coding performance by comparing with the VVC Test Model VTM-3.0. It is shown that under Common Test Condition (CTC) with inter MTS enabled, an average of 9%, 0%, and 3% decoding time savings are achieved for All Intra (AI), Random Access (RA) and Low Delay B (LDB), respectively. Under low QP test condition with inter MTS enabled, the proposed fast methods achieve 1%, 2% and 4% decoding time savings on average for AI, RA, and LDB, respectively.
Zhaobin Zhang, Xin Zhao 0003, Xiang Li 0003, Li Li 0040, Shan Liu 0001, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Transform Coding in the VVC Standard
abstract
In the past decade, the development of transform coding techniques has achieved significant progress and several advanced transform tools have been adopted in the new generation Versatile Video Coding (VVC) standard. In this paper, a brief history of transform coding development during VVC standardization is presented, and the transform coding tools in the VVC standard are described in detail together with their initial design, incremental improvements and implementation aspects. To improve coding efficiency, four new transform coding techniques are introduced in VVC, which are namely Multiple Transform Selection (MTS), Low-Frequency Non-separable Secondary Transform (LFNST) and Sub-Block Transform (SBT), as well as a large (64-point) type-2 DCT. The experimental results on VVC reference software (VTM-9.0) show that average 4.5% and 3.6% overall coding gain can be achieved by the VVC transform coding tools for All Intra and Random Access configurations, respectively.
Xin Zhao 0003, Seung-Hwan Kim 0001, Yin Zhao, Hilmi E. Egilmez, Moonmo Koo, Shan Liu 0001, Jani Lainema, Marta Karczewicz
IEEE Trans. Circuits Syst. Video Technol.1
2020 Unified Secondary Transform for Intra Coding Beyond Av1
abstract
In AV1, only separable transforms are applied as primary transform. Separable transform can efficiently capture the statistical correlation of residual samples along horizontal and vertical directions. However, for typical natural image/video contents which usually present arbitrary directionality in the residual samples, the efficiency of separable transform is rather limited comparing to the high-complexity non-separable primary transform. To capture the directionality of residual samples with relatively lower complexity, in this paper, a non-separable unified secondary transform scheme, which is customized for AV1 intra coding scheme, is proposed. The proposed method categorizes the secondary transform kernels according to the nominal intra prediction angles defined in AV1 and share the bases among the recursive filtering modes with DC prediction mode. Experimental results show that, based on the libaom implementation of AV1, the proposed method provides average 2.5% luma BD-rate for intra coding using a set of well-known video content as test material, and up to 14% luma BD-rate saving for still image coding using the DIV2K data set as test material.
Xin Zhao 0003, Shan Liu 0001
ICIP1
2020 Improved Intra Coding Beyond AV1 Using Adaptive Prediction Angles and Reference Lines
abstract
A fixed set of intra prediction angles by using the reconstructed samples in adjacent reference line are employed in AV1 to remove the spatial redundancy of video signals. Two methods are proposed in this paper to further improve the intra coding performance of AV1. Firsly, to better signal the intra prediction modes, only a subset of the intra prediction modes (IPMs) are allowed and signaled for each block, which is adaptively selected according to the IPMs of neighboring blocks. Secondly, to reduce the prediction errors when there is a strong discontinuity between the samples in current block and its adjacent reference samples, an adaptive reference line selection method is proposed by enabling farther reference lines for intra prediction. Experimental results show that, the proposed methods achieve 2.2% luma BD-rate savings with around 150% encoding time for intra coding on top of the libaom implementation of AV1.
Xin Zhao 0003, Shan Liu 0001
ICIP2
2019 Multiple Reference Line Coding for Most Probable Modes in Intra Prediction
abstract
Intra-picture prediction as in HEVC exploits the nearest reference line adjacent to the current coding unit (CU) for prediction of samples. If this reference line represents a discontinuity, the reference samples in this reference line can differ to a large extent from the original samples and may lead to a large prediction error. We propose a multiple reference lines (MRLs) coding to allow not only the nearest reference line 0 but also reference lines 1 and 3 to be candidates for angular intra prediction as shown in Fig. 1. To reduce the complexity arising from additional lines to be checked at encoder side, we further propose to restrict the MRL to angular most probable modes (MPMs) only. The MRL coding signals the reference line index before the intra prediction mode. This allows to not signal the MPM flag of the current CU and implicitly derive it as true when a non-zero reference line index is signaled. Experimental results are provided to evaluate the performance of the proposed MRL coding on top of the VVC test model VTM-2.0.1. 26 test sequences in different categories, including 4k, 1080p, 720p, WVGA, WQVGA resolutions and screen contents are tested. Two coding structures are evaluated, all intra (AI) and random access (RA). The objective coding efficiency is measured in terms of Bjøntegaard Delta (BD) rate (%) computed using four rate/PSNR points that were generated by using quantization parameters 22, 27, 32 and 37. Lower (negative) BD-rate implies better compression rate. Table 1 shows that the presented MRL provides 0.46% bitrate savings for an all-intra and 0.2% for a random-access configuration on average. Furthermore, it provides 1.45% bitrate reduction for screen content test sequences, which are representing an increasingly important video application. Because of a fairly good trade-off between coding efficiency and complexity, the proposed MRL coding mode with MPM restriction was adopted into the current VVC draft standard.
Yao-Jen Chang, Hong-Jheng Jhu, Hui-Yu Jiang, Xin Zhao 0003, Xiang Li 0003, Shan Liu 0001, Benjamin Bross, Paul Keydel, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
DCC5
2019 Fast Adaptive Multiple Transform for Versatile Video Coding
abstract
The Joint Video Exploration Team (JVET) recently launched the standardization of next-generation video coding named Versatile Video Coding (VVC) in which the Adaptive Multiple Transforms (AMT) is adopted as the primary residual coding transform solution. AMT introduces multiple transforms selected from the DST/DCT families and achieves noticeable coding gains. However, the set of transforms are calculated using direct matrix multiplication which induces higher run-time complexity and limits the application for practical video codec. In this paper, a fast DST-VII/DCT-VIII algorithm based on partial butterfly with dual implementation support is proposed, which aims at achieving reduced operation counts and run-time cost meanwhile yield almost the same coding performance. The proposed method has been implemented on top of the VTM-1.1 and experiments have been conducted using Common Test Conditions (CTC) to validate the efficacy. The experimental results show that the proposed methods, in the state-of-the-art codec, can provide an average of 7%, 5% and 8% overall decoding time savings under All Intra (AI), Random Access (RA) and Low Delay B (LDB) configuration, respectively yet still maintains coding performance.
Zhaobin Zhang, Xin Zhao 0003, Xiang Li 0003, Zhu Li 0001, Shan Liu 0001
DCC2
2019 Wide Angular Intra Prediction for Versatile Video Coding
abstract
This paper presents a technical overview of Wide Angular Intra Prediction (WAIP) that was adopted into the test model of Versatile Video Coding (VVC) standard. Due to the adoption of flexible block partitioning using binary and ternary splits, a Coding Unit (CU) can have either a square or a rectangular block shape. However, the conventional angular intra prediction directions, ranging from 45 degrees to -135 degrees in clockwise direction, were designed for square CUs. To better optimize the intra prediction for rectangular blocks, WAIP modes were proposed to enable intra prediction directions beyond the range of conventional intra prediction directions. For different aspect ratios of rectangular block shapes, different number of conventional angular intra prediction modes were replaced by WAIP modes. The replaced intra prediction modes are signaled using the original signaling method. Simulation results reportedly show that, with almost no impact on the run-time, on average 0.31% BD-rate reduction is achieved for intra coding using VVC test model (VTM).
Xin Zhao 0003, Shan Liu 0001, Xiang Li 0003, Jani Lainema, Gagan Rath, Fabrice Urban, Fabien Racapé
DCC2
2018 Low-Complexity Intra Prediction Refinements for Video Coding
abstract
In existing video coding standards such as H.264/AVC and HEVC, the intra prediction is typically derived using fixed, symmetric prediction filters along the prediction direction, e.g., in planar mode, top-right and bottom-left samples are predicted using symmetric prediction filters. However, in case of asymmetric availability of neighboring reference samples, the performance of intra prediction filters designed in HEVC may not be optimal. To further refine the intra prediction and achieve higher accuracy of prediction samples, this paper proposes low-complexity refinements over HEVC intra prediction, which are applied on frequently used planar, DC, horizontal and vertical modes. The proposed method only requires simple addition and bit-shift operations on top of HEVC's intra prediction implementation. Experimental results show that, an average of 0.7% coding gain is achieved for intra coding with no increase in run-time complexity.
Xin Zhao 0003, Vadim Seregin, Amir Said, Kai Zhang 0007, Hilmi E. Egilmez, Marta Karczewicz
PCS1
2018 Coupled Primary and Secondary Transform for Next Generation Video Coding
abstract
The discrete cosine transform type II can efficiently approximate the Karhunen-Loeve transform under the first-order stationary Markov condition. However, the highly dynamic characteristics of natural images will not always follow the first-order stationary Markov condition. It is well known that multi-core transforms and non-separable transforms capture diversified and directional texture patterns more efficiently. And a combination of enhanced multiple transform (EMT) and non-separable secondary transform (NSST) are provided in the reference software of the next generation video coding standard to solve this problem. However, the current method of combining the EMT and NSST may lead to quite significant encoder complexity increase, which makes the video codec rather impractical for real applications. Therefore, in this paper, we investigate the interactions between EMT and NSST, and propose a coupled primary and secondary transform to simplify the combination to obtain a better trade-off between the performance and the encoder complexity. With the proposed method, the transform for the Luma and Chroma components is also unified for a consistent design as an additional benefit. We implement the proposed transform on top of the Next software, which has been proposed for the next generation video coding standard. The experimental results demonstrate that the proposed algorithm can provide significant time reduction while keeping the majority of the performance.
Xin Zhao 0003, Li Li 0040, Zhu Li 0001, Xiang Li 0003, Shan Liu 0001
VCIP1
2018 Joint Separable and Non-Separable Transforms for Next-Generation Video Coding
abstract
Throughout the past few decades, the separable Discrete Cosine Transform (DCT), particularly the DCT type II, has been widely used in image and video compression. It is well known that, under first-order stationary Markov conditions, DCT is an efficient approximation of the optimal Karhunen-Loève transform. However, for natural image and video sources, the adaptivity of a single separable transform with fixed core is rather limited for the highly dynamic image statistics, e.g., textures and arbitrarily directed edges. It is also known that non-separable transforms can achieve better compression efficiency for images with directional texture patterns, yet they are computationally complex, especially when the transform size is large. In order to achieve higher transform coding gains with relatively low-complexity implementations, we propose a joint separable and non-separable transform. The proposed separable primary transform, named Enhanced Multiple Transform (EMT), applies multiple transform cores from a pre-defined subset of sinusoidal transforms, and the transform selection is signaled in a joint block level manner. Moreover, a Non-Separable Secondary Transform (NSST) method is proposed to operate in conjunction with EMT. Unlike the existing non-separable transform schemes which require excessive amounts of memory and computation, the proposed NSST efficiently improves coding gain with much lower complexity. Extensive experimental results show that the proposed methods, in a state-of-the-art video codec, such as HEVC, can provide significant coding gains (average 6.9% and 4.5% bitrate reductions for intra and random-access coding, respectively).
Xin Zhao 0003, Jianle Chen, Marta Karczewicz, Amir Said, Vadim Seregin
IEEE Trans. Image Process.1
2017 Multiple direct mode for intra coding
abstract
In this paper, a multiple direct mode (MDM) method is presented for chroma intra coding. The main contributions of the proposed MDM method include two aspects: selection of multiple luma intra prediction modes from co-located luma blocks, and the derivation of chroma intra prediction modes from spatial neighbouring blocks. With the proposed method, both the cross-component correlation and spatial correlation of intra prediction modes can be better utilized for more efficient chroma intra coding. Simulation results have validated the efficiency of MDM especially under the decoupled luma-chroma partition trees. The proposed method has been adopted in the Joint Exploration Model (JEM) which is the test platform for future video coding technology exploration in Joint Video Exploration Team (JVET).
Li Zhang 0006, Wei-Jung Chien, Jianle Chen, Xin Zhao 0003, Marta Karczewicz
VCIP4
2016 Enhanced Multiple Transform for Video Coding
abstract
The Discrete Cosine Transform (DCT), and in particular the DCT type II, has been widely used for image and video compression. Although DCT efficiently approximates the optimal Karhunen–Loève transform under first-order Markov conditions with low complexity, the energy packing efficiency is still limited since a fixed transform cannot always capture the highly dynamic statistics of natural video content. In this paper, to further improve the transform efficiency, an Enhanced Multiple Transform (EMT) scheme is proposed. In the proposed EMT, a few sinusoidal transforms, other than DCT, have also been utilized for coding both Intra and Inter prediction residuals. The best transform, as selected from a pre-defined transform subset specified by prediction mode, is explicitly signaled in a joint coding block level manner. Moreover, to accelerate encoding process, fast methods have also been proposed by skipping unnecessary transform rate-distortion evaluations using previously encoding statistics. The proposed method has been implemented on top of High-Efficiency Video Coding (HEVC) reference software, and significant coding gain has been verified.
Xin Zhao 0003, Jianle Chen, Marta Karczewicz, Li Zhang 0006, Xiang Li 0003, Wei-Jung Chien
DCC1
2016 Position dependent prediction combination for intra-frame video coding
abstract
Intra-frame prediction in the High Efficiency Video Coding (HEVC) standard can be empirically improved by applying sets of recursive two-dimensional filters to the predicted values. However, this approach does not allow (or complicates significantly) the parallel computation of pixel predictions. In this work we analyze why the recursive filters are effective, and use the results to derive sets of non-recursive predictors that have superior performance. We present an extension to HEVC intra prediction that combines values predicted using non-filtered and filtered (smoothed) reference samples, depending on the prediction mode, and block size. Simulations using the HEVC common test conditions show that a 2.0% bit rate average reduction can be achieved compared to HEVC, for All Intra (AI) configurations.
Amir Said, Xin Zhao 0003, Marta Karczewicz, Jianle Chen
ICIP2
2016 Highly efficient non-separable transforms for next generation video coding
abstract
For the last few decades, the application of signal-adaptive transform coding to video compression has been stymied by the large computational complexity of matrix-based solutions. In this paper, we propose a novel parametric approach to greatly reduce the complexity without degrading the compression performance. In our approach, instead of following the conventional technique of identifying full transform matrices that yield best compression efficiency, we look for the best transform parameters defining a new class of transforms, called HyGTs, which have low complexity implementations that are easy to parallelize. The proposed HyGTs are implemented as an extension of High Efficiency Video Coding (HEVC), and our comprehensive experimental results demonstrate that proposed HyGTs improve average coding gain by 6% bit rate reduction, while using 6.8 times less memory than KLT matrices.
Amir Said, Xin Zhao 0003, Marta Karczewicz, Hilmi E. Egilmez, Vadim Seregin, Jianle Chen
PCS2
2016 NSST: Non-separable secondary transforms for next generation video coding
abstract
In traditional image and video coding schemes, separable transforms are typically employed due to their low-complexity implementations. However, the compression efficiency of separable transforms is limited for most natural image/video blocks which generally have arbitrarily directed edge and texture patterns. It is well known that non-separable transforms can achieve better compression efficiency for directional texture patterns, yet they are computationally complex, especially for larger block sizes. In order to achieve higher transform coding gains with relatively low-complexity implementations, in this paper, we propose non-separable secondary transforms (NSSTs). The proposed approach applies a secondary non-separable transform on a sub-block of low frequency coefficients generated using a primary separable transform, such as discrete cosine transform (DCT). Since the proposed NSST is a non-separable transform applied on low frequency coefficients in a much smaller block size, which typically captures most of the signal energy, better coding gains can be achieved with at a relatively low-computational cost. Experimental results show that, compared to the latest HEVC reference software (HM16.6), the proposed method achieves up to a significant 12% coding gain for Intra coding.
Xin Zhao 0003, Jianle Chen, Amir Said, Vadim Seregin, Hilmi E. Egilmez, Marta Karczewicz
PCS1
2016 Multiview and 3D Video Compression Using Neighboring Block Based Disparity Vectors
abstract
Compression of the statistical redundancy among different viewpoints, i.e., inter-view redundancy, is a fundamental and critical problem in multiview and three-dimensional (3D) video coding. To exploit the inter-view redundancy, disparity vectors are required to identify pixels of the same objects within two different views; in this way, the enhancement coding tools can be efficiently employed as new modes in block-based video codecs to achieve higher compression efficiency. Although disparity can be converted from depth, it is not possible in multiview video coding since depth information is not considered. Even when depth information is coded, it breaks the so-called multiview compatibility wherein texture views can be decoded without depth information. To resolve this problem, in this paper, a neighboring block-based disparity vector derivation (NBDV) method is proposed. The basic concept of NBDV is to derive a disparity vector (DV) of a current block by utilizing the motion information of spatially and temporally neighboring blocks predicted from another view. Through extensive experiments and analysis, it is shown that the proposed NBDV method achieves efficient DV derivation in the state-of-art video codecs, and it keeps the multiview compatibility with a relatively lower complexity. The proposed method has become an essential part of the 3D video standard extensions of H.264/AVC and HEVC.
Ying Chen 0011, Xin Zhao 0003, Li Zhang 0006, Je-Won Kang
IEEE Trans. Multim.2
2015 Texture based sub-PU motion inheritance for depth coding
abstract
The 3D extension of High Efficiency Video Coding (HEVC) has been developed as a standard to code multiview data with depth information. In addition to the existing coding tools in HEVC, in the 3D extension, depth pictures can be coded with supplemental techniques typically designed to utilize the special characteristic of depth pictures. To better utilize the correlation between texture and depth pictures, in this paper, we present an enhanced motion prediction method for depth coding using the motion from the associated texture picture. Such a method inherits the motion information for sub-blocks of depth prediction units. The proposed method provides in average 3.2% bit rate reduction for a 3D video system, wherein the quality is evaluated by the synthesized views.
Ying Chen 0011, Xin Zhao 0003
ICIP3
2014 Unified wedgelet genenration for depth coding in 3D-HEVC
abstract
In 3D-HEVC, bi-partition based modes, i.e., depth modeling modes (DMM), are applied for depth intra coding. With DMM, a depth prediction unit (PU) is partitioned into two parts using a Wedgelet pattern selected from a predefined large Wedgelet set, and the Wedgelet set is generated during both encoder and decoder initialization for each block size ranging from 4 ×4 to 32×32. The generation of Wedgelet sets involves relatively high complexity in terms of both storage and computation, especially for large block sizes. To simplify the Wedgelet generation, in this paper, we propose a unified Wedgelet generation method which derives large Wedgelet patterns by extending the primitive 4×4 Wedgelet pattern along its partition boundary line, and the storage of large Wedgelet patterns can be saved. Experimental results show almost no coding efficiency degradation using the proposed method, and the complexity of Wedgelet generation process is largely reduced in a unified manner.
Xin Zhao 0003, Li Zhang 0006, Ying Chen 0011
ICIP1
2014 Derived disparity vector based NBDV for 3D-AVC
abstract
In the 3D video extension of H.264/AVC, namely 3D-AVC, Neighboring Based Disparity Vector (NBDV) derivation has been proposed to support multiview/stereo compatibility, therefore texture views can be decoded independently to depth views. NBDV generates a disparity vector for the current macroblock (MB) using the motion information of neighboring blocks, especially those coded with motion vectors pointing to inter-view reference pictures. In 3D-AVC, NBDV has been utilized to access minimum number of spatial and temporal neighboring blocks, therefore there is a high probability that NBDV does not derive an efficient disparity vector. This paper introduces a derived disparity vector scheme, wherein only one disparity vector derived from NBDV is maintained for the whole slice and it is used as the disparity vector of the current MB if NBDV does not derive one from neighboring blocks. Simulation results show that the proposed method provides 3.6% bit rate reduction for multiview coding.
Xin Zhao 0003, Ying Chen 0011, Li Zhang 0006
VCIP1
2013 Texture mode dependent depth coding in 3D-HEVC
abstract
In 3D-HEVC, depth modeling modes (DMM) are applied for efficient intra depth coding. With DMM modes, a prediction unit is partitioned into two parts, and each part is predicted by a single value. In one DMM mode, the partition pattern is implicitly derived at the decoder by searching all pre-defined Wedgelet patterns on a Co-located Texture Luma Block (CTLB), which increases the decoding complexity drastically. To simplify the design of this DMM mode, in this paper, we propose to utilize the Intra Prediction Mode (IPM) of CTLB to largely skip unnecessary Wedgelet searches at the decoder. To achieve this goal, for each IPM, only a limited number of Wedgelet patterns are selected as candidates in this DMM mode. Experimental results demonstrate that, with almost no coding performance degradation, the proposed method significantly reduces the decoding complexity by skipping 90% of Wedgelet searches. The proposed method has been adopted by 3D-HEVC.
Xin Zhao 0003, Ying Chen 0011, Li Zhang 0006, Marta Karczewicz
ICIP1
2013 Neighboring block based disparity vector derivation for 3D-AVC
abstract
3D-AVC, being developed under Joint Collaborative Team on 3D Video Coding (JCT-3V), significantly outperforms the Multiview Video Coding plus Depth (MVC+D) which has no new macroblock level coding tools compared to Multiview video coding extension of H.264/AVC (MVC). However, for multiview compatible configuration, i.e., when texture views are decoded without accessing depth information, the performance of the current 3D-AVC is only marginally better than MVC+D. The problem is caused by the lack of disparity vectors which can be obtained only from the coded depth views in 3D-AVC. In this paper, a disparity vector derivation method is proposed by using the motion information of neighboring blocks and applied along with existing coding tools in 3D-AVC. The proposed method improves 3D-AVC in the multiview compatible mode substantially, resulting in about 20% bitrate reduction for texture coding. When enabling the so-called view synthesis prediction to further refine the disparity vectors, the performance of the proposed method is 31% better than MVC+D and even better than 3D-AVC under the best performing 3D-AVC configuration.
Li Zhang 0006, Je-Won Kang, Xin Zhao 0003, Ying Chen 0011, Rajan L. Joshi
VCIP3
2012 Video Coding With Rate-Distortion Optimized Transform
abstract
Block-based discrete cosine transform (DCT) has been successfully adopted into several international image/video coding standards, e.g., MPEG-2, H.264/AVC, as it can achieve a good tradeoff between performance and complexity. Although DCT theoretically approximates the optimum Karhunen–Loève transform under first-order Markov conditions, one fixed set of transform basis functions (TBF) cannot handle all the cases efficiently due to the non-stationary nature of video contents. To further improve the performance of block-based transform coding, in this paper, we present the design of rate-distortion optimized transform (RDOT) which contributes to both intraframe and interframe coding. The most important property which makes a difference between RDOT and the conventional DCT is that, in the proposed method, transform is implemented with multiple TBF candidates which are obtained from off-line training. With this feature, for coding each residual block, the encoder is capable to select the optimal set of TBF in terms of rate-distortion performance, and better energy compaction is achieved in the transform domain. To obtain an optimum group of candidate TBF, we have developed a two-step iterative optimization technique for the off-line training, with which the TBF candidates are refined at each iteration until the training process becomes converged. Moreover, analysis on the optimal group of candidate TBF is also presented in this paper, with a detailed description of a practical implementation for the proposed algorithm on the latest VCEG key technical area software platform. Extensive experimental results show that, compared with the conventional DCT-based transform scheme adopted into the state-of-the-art H.264/AVC video coding standard, significant improvement of coding performance has been achieved for both intraframe and interframe coding with our proposed method.
Xin Zhao 0003, Li Zhang 0006, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2011 Enhanced line-based intra prediction with fixed interpolation filtering
abstract
In this paper, a novel intra prediction algorithm, named enhanced line-based intra prediction (ELIP), is proposed to improve the traditional intra prediction methods for image/intra-frame coding. Different from the existing schemes where the interpolation filtering only depends on the intra prediction mode, in the proposed ELIP, the linear filtering coefficients are further refined by imposing both mode and position dependencies. For each intra prediction mode and each position within one block, a set of linear filtering coefficients are obtained from off-line training using the least square method. The proposed algorithm has been implemented in the latest ITU-T VCEG KTA software. Experimental results demonstrate that, compared with the intra coding scheme of KTA with new intra coding tool enabled, up to 0.44dB additional coding gain is achieved by the proposed method, while keeping applicable computational complexity for practical video codecs.
Li Zhang 0006, Siwei Ma 0001, Wen Gao 0001, Xin Zhao 0003
ISCAS4
2011 Novel intra prediction via position-dependent filtering
Li Zhang 0006, Xin Zhao 0003, Siwei Ma 0001, Qiang Wang 0011, Wen Gao 0001
J. Vis. Commun. Image Represent.2
2010 Rate-distortion optimized transform for intra-frame coding
abstract
In this paper, a novel algorithm is proposed for intra-frame coding, named as rate-distortion optimized transform (RDOT). Unlike existing intra-frame coding schemes where the transform matrices are either fixed or mode dependent, in the proposed algorithm, transform is implemented with multiple candidate transform matrices. With this flexibility, for coding each residual block, the encoder is endowed with the power to select the optimal transform matrix in terms of rate-distortion tradeoff. The proposed algorithm has been implemented in the latest ITU-T VCEG-KTA software. Experimental results show that, over a wide range of test set, the proposed method achieves average 0.43dB coding gain compared with the recent Mode-Dependent Directional Transform (MDDT). The improvement is more significant at high bit-rates, and up to 1dB coding gain can be achieved.
Xin Zhao 0003, Li Zhang 0006, Siwei Ma 0001, Wen Gao 0001
ICASSP1
2010 Fast rate-distortion optimized transform for Intra coding
abstract
In our previous work, the rate-distortion optimized transform (RDOT) is introduced for Intra coding, which is featured by the usage of multiple offline-trained transform matrix candidates. The proposed RDOT achieves remarkable coding gain for KTA Intra coding, while maintaining almost the same computational complexity at the decoder. However, at the encoder, the computational complexity is increased drastically by the expensive ratedistortion (R-D) optimized selection of transform matrix. To resolve this problem, in this paper, we propose a fast RDOT scheme using macroblock- and block-level R-D cost thresholding. With the proposed method, unnecessary mode trials and R-D evaluations of transform matrices can be efficiently skipped from the mode decision process. Extensive experimental results show that, with negligible coding performance degradation, about 88.9% of the total encoding time is saved by the proposed method.
Xin Zhao 0003, Li Zhang 0006, Siwei Ma 0001, Wen Gao 0001
PCS1
2010 Novel Statistical Modeling, Analysis and Implementation of Rate-Distortion Estimation for H.264/AVC Coders
abstract
In H.264/advanced video coding, the encoder employs the rate-distortion optimization (RDO) to select the optimal coding mode of each block. Although it is effective to employ the RDO technique for mode decision, the computation load increases drastically. To reduce the computation complexity of the RDO technique, in this paper, we propose efficient algorithms for the estimation of block-level rate and distortion. For rate estimation, we model the transform coefficients with accurate generalized Gaussian distributions, and the weighted sum of absolute quantized transform coefficients is proposed as an efficient rate estimator, where the weights provide an implicit mechanism for evaluating different contributions of different frequency components to the coding bits. For distortion estimation, we first analyze the origins of distortion thoroughly. Then a direct relationship between the discarded bits in quantization and the distortion is explored. According to this investigation, a simple and efficient algorithm is proposed for the distortion estimation. With above proposed algorithms, the RDO technique can be efficiently implemented in a low-complexity way. Extensive experimental results demonstrate that, compared with the original RDO implementation, the proposed algorithms achieve about 32% reduced total encoding time with ignorable coding performance degradation.
Xin Zhao 0003, Jun Sun 0012, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1