EDBT 2026 Demo / reviewers in the wild / expert
Xiang Li 0003
dblp:40/1491-3
· DBLP profile ↗
45ranked-venue papers
20as first author
9since 2021 · last 2025
0000-0002-0575-2143ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 18 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorSystems, architecture and hardware · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Loop Filters and Edge Enhancement for Variable Resolution Video Coding
Tim Claßen, Xiang Li 0003, Priyanka Das 0005, Mathias Wien |
PCS | 2 |
| 2023 | Panoramic Video Salient Object Detection with Ambisonic Audio GuidanceabstractVideo salient object detection (VSOD), as a fundamental computer vision problem, has been extensively discussed in the last decade. However, all existing works focus on addressing the VSOD problem in 2D scenarios. With the rapid development of VR devices, panoramic videos have been a promising alternative to 2D videos to provide immersive feelings of the real world. In this paper, we aim to tackle the video salient object detection problem for panoramic videos, with their corresponding ambisonic audios. A multimodal fusion module equipped with two pseudo-siamese audio-visual context fusion (ACF) blocks is proposed to effectively conduct audio-visual interaction. The ACF block equipped with spherical positional encoding enables the fusion in the 3D context to capture the spatial correspondence between pixels and sound sources from the equirectangular frames and ambisonic audios. Experimental results verify the effectiveness of our proposed components and demonstrate that our method achieves state-of-the-art performance on the ASOD60K dataset. Xiang Li 0003, Haoyuan Cao, Shijie Zhao 0001, Li Zhang 0006, Bhiksha Raj |
AAAI | 1 |
| 2022 | Towards Joint Frame-Level and MOS Quality Predictions with Low-Complexity Objective ModelsabstractThe evaluation of the quality of gaming content, with low-complexity and low-delay approaches is a major challenge raised by the emerging gaming video streaming and cloud-gaming services. Considering two existing and a newly created gaming databases this paper confirms that some low-complexity metrics match well with subjective scores when considering usual correlation indicators. It is however argued such a result is insufficient: gaming content suffers from sudden large quality drops that these indicators do not capture. In addition to proposing three new low-complexity models based on various machine learning techniques, this paper introduces a new indicator to capture sudden quality variations and reports poor results for most of the models when applying this indicator. Consequently, an original way to train the models, using jointly the subjective scores and the frame level scores of a full-reference metric, is proposed. The high correlation through traditional indicators is preserved, while the efficiency on the new indicator is drastically improved. Joël Jung, Alexandre Giraud, Meijia Song, Songnan Li, Xiang Li 0003, Shan Liu 0001 |
ICASSP | 5 |
| 2022 | Unified Fast Partitioning Algorithm for Intra and Inter Predictions in Versatile Video CodingabstractThe Versatile Video Coding (VVC) standard adopts a more flexible partitioning structure beyond the High Efficiency Video Coding (HEVC) standard. The quadtree with nested multi-type tree (QTMT) partitioning structure greatly improves coding efficiency. Nevertheless, the typical recursive coding unit (CU) partitioning scheme causes a substantial increase in computational complexity at the encoder. In this paper, a unified fast algorithm is proposed for both intra and inter predictions, which makes use of various historical information of previously checked partitions. The proposed algorithm is implemented on top of the VVC reference software VTM-14.0. Experimental results show that the proposed algorithm achieves an average 42.47%, 40.46%, and 40.38% encoder runtime reduction with only 0.45%, 1.08%, and 1.18% Bjøntegaard delta bitrate loss (BDBR) in All Intra (AI), Random Access (RA), and Low Delay P (LDP) configurations, respectively. Wei Kuang, Xiang Li 0003, Xin Zhao 0003, Shan Liu 0001 |
PCS | 2 |
| 2021 | Video Coding Tool Analysis and Dataset for Gaming ContentabstractThe gaming market has kept growing significantly in recent years. Driven by multiple technology advances, such as cloud computing and video technologies, new gaming applications, e.g. AR, VR and cloud gaming, are becoming more and more practical and popular. Among different types of gaming, the emergence of cloud gaming is driving the market with enhanced gamer experience as well as new challenges to the services. One of the key technological challenges of gaming applications is the video coding, which is the foundation of several popular gaming applications, including cloud gaming and game live streaming. Comparing to the typical camera captured content and screen content, gaming content presents unique features that directly lead to different preferences on the selection of coding tool sets. To better understand the behaviors of known video coding tools and provide test materials for research and development on future coding tools that benefits more on gaming content, in this paper, a dataset consists of a set of gaming video is proposed together with analysis of the performances of existing coding tools on these materials. It is observed that, several known coding tools are exceptionally beneficial for gaming content and the rational is analyzed in this paper. Xin Zhao 0003, Shan Liu 0001, Xiang Li 0003, Guichun Li, Xiaozhong Xu |
PCS | 3 |
| 2021 | A Novel Video Coding Strategy in HEVC for Object DetectionabstractOccupying the most significant portion of global data traffic, video is being generated in almost every aspect of our life. Because of its huge volume, we are depending much more heavily on machine intelligence based analysis. In the meantime, video coding technology has been continuously improved for better compression efficiency. However, the state-of-the-art video coding standards, such as H.265/HEVC and versatile video coding (VVC), are still designed assuming that the compressed video will be watched by a human later. Such a design is not optimal when the compressed video will be used by computer vision applications. While the human visual system (HVS) is consistently sensitive to the content with high contrast, the impact of pixels on computer vision algorithms is task driven. For example, because of the different categories of objects used to train detection algorithms, the influence of the same image content on those detectors also varies. Therefore, human oriented video coding strategies may not be optimal when the compressed signal is further processed by algorithms, as the encoder is unaware of the task specific information. In this article, taking object detection as an example, we propose a novel video coding strategy for computer vision. By protecting the information according to its importance for an object detector rather than for the human visual system, our proposed method has the potential to achieve a better object detection performance with the same bandwidth. The main contributions of our paper are: 1) the modeling of the relationship between object detection accuracy and bit rate; 2) a back propagation based method to analyze the influence of each pixel on the detection of target objects; 3) an object detection oriented bit allocation and codec control parameter determination scheme; 4) an evaluation metric to compare the impact of video coding strategies on a given object detector over a predefined range of bit rate. Experimental results demonstrate that our proposed algorithm can better preserve the video content vital for object detection than state-of-the-art video coding schemes. Dapeng Oliver Wu, Shan Liu 0001, Xiang Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Motion Vector Coding and Block Merging in the Versatile Video Coding StandardabstractThis paper overviews the motion vector coding and block merging techniques in the Versatile Video Coding (VVC) standard developed by the Joint Video Experts Team (JVET). In general, inter-prediction techniques in VVC can be classified into two major groups: “whole block-based inter prediction” and “subblock-based inter prediction”. In this paper, we focus on techniques for whole block-based inter prediction. As in its predecessor, High Efficiency Video Coding (HEVC), whole block-based inter prediction in VVC is represented by adaptive motion vector prediction (AMVP) mode or merge mode. Newly introduced features purely for AMVP mode include symmetric motion vector difference and adaptive motion vector resolution. The features purely for merge mode include pairwise average merge, merge with motion vector difference, combined inter-intra prediction and geometric partitioning mode. Coding tools such as history-based motion vector prediction and bidirectional prediction with coding unit weights can be applied on both AMVP mode and merge mode. This paper discusses the design principles and the implementation of the new inter-prediction methods. Using objective metrics, simulation results show that the methods overviewed in the paper can jointly achieve 6.2% and 4.7% BD-rate savings on average with the random access and low-delay configurations, respectively. Significant subjective picture quality improvements of some tools are also reported when comparing the resulting pictures at same bitrates. Wei-Jung Chien, Li Zhang 0006, Martin Winken, Xiang Li 0003, Ru-Ling Liao, Han Gao 0001, Hongbin Liu 0004, Chun-Chi Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Block Partitioning Structure in the VVC StandardabstractVersatile Video Coding (VVC) is the latest video coding standard jointly developed by ITU-T VCEG and ISO/IEC MPEG. In this paper, technical details and experimental results for the VVC block partitioning structure are provided. Among all the new technical aspects of VVC, the block partitioning structure is identified as one of the most substantial changes relative to the previous video coding standards and provides the most significant coding gains. The new partitioning structure is designed using a more flexible scheme. Each coding tree unit (CTU) is either treated as one coding unit or split into multiple coding units by one or more recursive quaternary tree partitions followed by one or more recursive multi-type tree splits. The latter can be horizontal binary tree split, vertical binary tree split, horizontal ternary tree split, or vertical ternary tree split. A CTU dual tree for intra-coded slices is described on top of the new block partitioning structure, allowing separate coding trees for luma and chroma. Also, a new way of handling picture boundaries is presented. Additionally, to reduce hardware decoder complexity, virtual pipeline data unit constraints are introduced, which forbid certain multi-type tree splits. Finally, a local dual tree is described, which reduces the number of small chroma intra blocks. Yu-Wen Huang, Jicheng An, Han Huang 0001, Xiang Li 0003, Shih-Ta Hsiang, Kai Zhang 0007, Han Gao 0001, Jackie Ma, Olena Chubach |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Fast DST-VII/DCT-VIII With Dual Implementation Support for Versatile Video CodingabstractThe Joint Video Exploration Team (JVET) recently launched the standardization of the next-generation video coding named Versatile Video Coding (VVC) with the inherited technical framework from its predecessor High-Efficiency Video Coding (HEVC). The simplified Enhanced Multiple Transform (EMT) has been adopted as the primary residual coding transform solution, termed Multiple Transform Selection (MTS). In MTS, only the transform set consisting of DST-VII and DCT-VIII remains, excluding the other transform sets and the dependency on intra prediction modes. Significant coding gains are achieved by introducing new DST/DCT transforms, but the full matrix implementation is relatively costly compared to partial butterfly in terms of both software run-time and operation counts. In this work, we exploit the inherent features existing in DST-VII and DCT-VIII. Instead of repeating the element-wise additions and multiplications in full matrix operation, these features can be leveraged to achieve more efficient implementations which only use partial elements to derive the identical results. Existing transform matrices are further tuned to utilize these (anti-)symmetric features. A partial butterfly-type fast algorithm with dual-implementation support is proposed for DST-VII/DCT-VIII transform in VVC. Complexity analysis including operation counts and software run-time are conducted to validate the effectiveness. In addition, we prove the features are perfectly supported by theory. The proposed fast methods achieve noticeable software run-time savings without compromising on coding performance by comparing with the VVC Test Model VTM-3.0. It is shown that under Common Test Condition (CTC) with inter MTS enabled, an average of 9%, 0%, and 3% decoding time savings are achieved for All Intra (AI), Random Access (RA) and Low Delay B (LDB), respectively. Under low QP test condition with inter MTS enabled, the proposed fast methods achieve 1%, 2% and 4% decoding time savings on average for AI, RA, and LDB, respectively. Zhaobin Zhang, Xin Zhao 0003, Xiang Li 0003, Li Li 0040, Shan Liu 0001, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Multiple Reference Line Coding for Most Probable Modes in Intra PredictionabstractIntra-picture prediction as in HEVC exploits the nearest reference line adjacent to the current coding unit (CU) for prediction of samples. If this reference line represents a discontinuity, the reference samples in this reference line can differ to a large extent from the original samples and may lead to a large prediction error. We propose a multiple reference lines (MRLs) coding to allow not only the nearest reference line 0 but also reference lines 1 and 3 to be candidates for angular intra prediction as shown in Fig. 1. To reduce the complexity arising from additional lines to be checked at encoder side, we further propose to restrict the MRL to angular most probable modes (MPMs) only. The MRL coding signals the reference line index before the intra prediction mode. This allows to not signal the MPM flag of the current CU and implicitly derive it as true when a non-zero reference line index is signaled. Experimental results are provided to evaluate the performance of the proposed MRL coding on top of the VVC test model VTM-2.0.1. 26 test sequences in different categories, including 4k, 1080p, 720p, WVGA, WQVGA resolutions and screen contents are tested. Two coding structures are evaluated, all intra (AI) and random access (RA). The objective coding efficiency is measured in terms of Bjøntegaard Delta (BD) rate (%) computed using four rate/PSNR points that were generated by using quantization parameters 22, 27, 32 and 37. Lower (negative) BD-rate implies better compression rate. Table 1 shows that the presented MRL provides 0.46% bitrate savings for an all-intra and 0.2% for a random-access configuration on average. Furthermore, it provides 1.45% bitrate reduction for screen content test sequences, which are representing an increasingly important video application. Because of a fairly good trade-off between coding efficiency and complexity, the proposed MRL coding mode with MPM restriction was adopted into the current VVC draft standard. Yao-Jen Chang, Hong-Jheng Jhu, Hui-Yu Jiang, Xin Zhao 0003, Xiang Li 0003, Shan Liu 0001, Benjamin Bross, Paul Keydel, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
DCC | 6 |
| 2019 | Fast Adaptive Multiple Transform for Versatile Video CodingabstractThe Joint Video Exploration Team (JVET) recently launched the standardization of next-generation video coding named Versatile Video Coding (VVC) in which the Adaptive Multiple Transforms (AMT) is adopted as the primary residual coding transform solution. AMT introduces multiple transforms selected from the DST/DCT families and achieves noticeable coding gains. However, the set of transforms are calculated using direct matrix multiplication which induces higher run-time complexity and limits the application for practical video codec. In this paper, a fast DST-VII/DCT-VIII algorithm based on partial butterfly with dual implementation support is proposed, which aims at achieving reduced operation counts and run-time cost meanwhile yield almost the same coding performance. The proposed method has been implemented on top of the VTM-1.1 and experiments have been conducted using Common Test Conditions (CTC) to validate the efficacy. The experimental results show that the proposed methods, in the state-of-the-art codec, can provide an average of 7%, 5% and 8% overall decoding time savings under All Intra (AI), Random Access (RA) and Low Delay B (LDB) configuration, respectively yet still maintains coding performance. Zhaobin Zhang, Xin Zhao 0003, Xiang Li 0003, Zhu Li 0001, Shan Liu 0001 |
DCC | 3 |
| 2019 | Wide Angular Intra Prediction for Versatile Video CodingabstractThis paper presents a technical overview of Wide Angular Intra Prediction (WAIP) that was adopted into the test model of Versatile Video Coding (VVC) standard. Due to the adoption of flexible block partitioning using binary and ternary splits, a Coding Unit (CU) can have either a square or a rectangular block shape. However, the conventional angular intra prediction directions, ranging from 45 degrees to -135 degrees in clockwise direction, were designed for square CUs. To better optimize the intra prediction for rectangular blocks, WAIP modes were proposed to enable intra prediction directions beyond the range of conventional intra prediction directions. For different aspect ratios of rectangular block shapes, different number of conventional angular intra prediction modes were replaced by WAIP modes. The replaced intra prediction modes are signaled using the original signaling method. Simulation results reportedly show that, with almost no impact on the run-time, on average 0.31% BD-rate reduction is achieved for intra coding using VVC test model (VTM). Xin Zhao 0003, Shan Liu 0001, Xiang Li 0003, Jani Lainema, Gagan Rath, Fabrice Urban, Fabien Racapé |
DCC | 4 |
| 2019 | Intra block copy in Versatile Video Coding with Reference Sample Memory ReuseabstractScreen contents such as online gaming streaming, remote desktop and WIFI display, become popular in current mainstream video applications. In versatile video coding (VVC), the most recent international video coding standard development, coding tools have been evaluated for optimizing screen content materials. Intra block copy (IBC) has shown its effectiveness in coding of computer-generated contents such as texts and graphics. Therefore, it has been previously included into the HEVC standard version 4, extensions for screen content coding (SCC). A constrained version of IBC mode has also been adopted in the VVC standard where the compensation range is limited within the current coding-tree unit (CTU), assuming a 1-CTU size of memory is allocated for storing IBC’s reference samples. In this paper, methods are proposed to efficiently utilize this reference sample memory for IBC mode such that effectively the search range for IBC mode can be increased without requiring more memory to store the reference samples. As a result, significant coding efficiency improvement over the traditional 1-CTU search range setting can be achieved. One of the proposed memory reuse strategies is considered practical for implementation and therefore has been included in the VVC standard. Xiaozhong Xu, Xiang Li 0003, Shan Liu 0001 |
PCS | 2 |
| 2018 | Coupled Primary and Secondary Transform for Next Generation Video CodingabstractThe discrete cosine transform type II can efficiently approximate the Karhunen-Loeve transform under the first-order stationary Markov condition. However, the highly dynamic characteristics of natural images will not always follow the first-order stationary Markov condition. It is well known that multi-core transforms and non-separable transforms capture diversified and directional texture patterns more efficiently. And a combination of enhanced multiple transform (EMT) and non-separable secondary transform (NSST) are provided in the reference software of the next generation video coding standard to solve this problem. However, the current method of combining the EMT and NSST may lead to quite significant encoder complexity increase, which makes the video codec rather impractical for real applications. Therefore, in this paper, we investigate the interactions between EMT and NSST, and propose a coupled primary and secondary transform to simplify the combination to obtain a better trade-off between the performance and the encoder complexity. With the proposed method, the transform for the Luma and Chroma components is also unified for a consistent design as an additional benefit. We implement the proposed transform on top of the Next software, which has been proposed for the next generation video coding standard. The experimental results demonstrate that the proposed algorithm can provide significant time reduction while keeping the majority of the performance. Xin Zhao 0003, Li Li 0040, Zhu Li 0001, Xiang Li 0003, Shan Liu 0001 |
VCIP | 4 |
| 2018 | Enhanced Cross-Component Linear Model for Chroma Intra-Prediction in Video CodingabstractCross-Component Linear Model (CCLM) for chroma intra-prediction is a promising coding tool in Joint Exploration Model (JEM) developed by the Joint Video Exploration Team (JVET). CCLM assumes a linear correlation between the luma and chroma components in a coding block. With this assumption, the chroma components can be predicted by the Linear Model (LM) mode, which utilizes the reconstructed neighbouring samples to derive parameters of a linear model by linear regression. This paper presents three new methods to further improve the coding efficiency of CCLM. First, we introduce a multi-model CCLM (MM-CCLM) approach, which applies more than one linear models to a coding block. With MM-CCLM, reconstructed neighbouring luma and chroma samples of the current block are classified into several groups, and a particular set of linear model parameters is derived for each group. The reconstructed luma samples of the current block are also classified to predict the associated chroma samples with the corresponding linear model. Second, we propose a multi-filter CCLM (MF-CCLM) technique, which allows the encoder to select the optimal down-sampling filter for the luma component with the 4:2:0 colour format. Third, we present a LM-angular prediction (LAP) method, which synthesizes the angular intra-prediction and the MM-CCLM intra-prediction into a new chroma intra coding mode. Simulation results show that 0.55%, 4.66% and 5.08% BD rate savings in average on Y, Cb and Cr components respectively, are achieved for All Intra (AI) configurations with the proposed three methods. MM-CCLM and MF-CCLM have been adopted into the JEM by JVET. Kai Zhang 0007, Jianle Chen, Li Zhang 0006, Xiang Li 0003, Marta Karczewicz |
IEEE Trans. Image Process. | 4 |
| 2017 | Frame Rate Up-Conversion Based Motion Vector Derivation for Hybrid Video CodingabstractIn this paper, a MV derivation method based on the idea of frame rate up-conversion (FRUC) is proposed. When a block is signaled as FRUC mode, the motion information of the block is derived without signaling. Moreover, derived MVs are refined at sub-block level for more accurate motion field. In addition, two matching methods, i.e., bilateral matching and template matching are supported to obtain good performance in both bi-directional and uni-directional prediction. Simulations under HEVC common test conditions show that over 4.2% average BD-rate reduction was achieved over HEVC reference software HM-16.6 in the case of random access configuration. The method has been adopted into the Joint Exploration Model (JEM) developed by the joint video exploration team (JVET) of MPEG and ITU-T VCEG for the study of next generation video coding standard. Xiang Li 0003, Jianle Chen, Marta Karczewicz |
DCC | 1 |
| 2017 | Multi-model based cross-component linear model chroma intra-prediction for video codingabstractCross-component Linear Model (CCLM) chroma intra prediction assumes a linear correlation between the luma and chroma components in a coding block. With this assumption, the chroma components can be predicted by LM mode, which utilizes the reconstructed neighbouring samples to derive parameters of the linear model by linear regression. This paper presents a multi-model CCLM (MM-CCLM) approach, which applies more than one linear models in a coding block. With MM-CCLM, reconstructed neighbouring luma and chroma samples of the current block are classified into several groups and each group is used as a training set to derive its own linear model. The reconstructed luma samples of the current block are also classified to use corresponding linear model to predict the associated chroma samples. Simulation results show that 0.26%, 1.89% and 1.96% BD rate savings on Y, Cb and Cr components are achieved for All Intra (AI) configurations in average. The proposed method has been adopted in the Joint Exploration Model (JEM) by Joint Video Exploration Team (JVET). Kai Zhang 0007, Jianle Chen, Li Zhang 0006, Xiang Li 0003, Marta Karczewicz |
VCIP | 4 |
| 2017 | Overview of Color Gamut ScalabilityabstractDisplays' new rendering capabilities combined with the ever-growing number of video applications have fueled the emergence of new video formats addressing wider color gamut and larger frame size. Thus, the need in scalable compression technology to provide backward compatibility with legacy devices and capitalize on the superior compression performance of High Efficiency Video Coding (HEVC) has increased significantly. This paper gives an overview of the work carried out in the Joint Collaborative Team on Video Coding of ITU-T Study Group 16 (VCEG) and ISO/IEC JTC1/SC29/WG11 motion picture experts group (MPEG) to define scalable extensions of HEVC (SHVC) targeting these market requirements. The color gamut scalability (CGS) tool of SHVC is specially designed to support efficient scalable coding with multiple layers in different color spaces. The genesis of the SHVC-CGS tool is presented and the performance of the various proposals is compared. Finally, the design of the recently adopted SHVC-CGS inter-layer prediction is detailed. The experimental results validate its efficiency in coding video with extended color gamut and high dynamic range. Philippe Bordes, Pierre Andrivon, Xiang Li 0003, Yan Ye 0003, Yuwen He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Enhanced Multiple Transform for Video CodingabstractThe Discrete Cosine Transform (DCT), and in particular the DCT type II, has been widely used for image and video compression. Although DCT efficiently approximates the optimal Karhunen–Loève transform under first-order Markov conditions with low complexity, the energy packing efficiency is still limited since a fixed transform cannot always capture the highly dynamic statistics of natural video content. In this paper, to further improve the transform efficiency, an Enhanced Multiple Transform (EMT) scheme is proposed. In the proposed EMT, a few sinusoidal transforms, other than DCT, have also been utilized for coding both Intra and Inter prediction residuals. The best transform, as selected from a pre-defined transform subset specified by prediction mode, is explicitly signaled in a joint coding block level manner. Moreover, to accelerate encoding process, fast methods have also been proposed by skipping unnecessary transform rate-distortion evaluations using previously encoding statistics. The proposed method has been implemented on top of High-Efficiency Video Coding (HEVC) reference software, and significant coding gain has been verified. Xin Zhao 0003, Jianle Chen, Marta Karczewicz, Li Zhang 0006, Xiang Li 0003, Wei-Jung Chien |
DCC | 5 |
| 2016 | Geometry transformation-based adaptive in-loop filterabstractRecently, adaptive in-loop filter (ALF) for image/video coding has attracted increasing attention by its proven capability in improving coding performance. ALF is aiming to minimize the mean square error between original samples and decoded samples by using Wiener-based adaptive filter. Samples in a picture are classified into multiple categories and the samples in each category are then filtered with their associated adaptive filter. The filter coefficients may be signaled or inherited to optimize the tradeoff between the mean square error and the overhead. In this paper, a Geometry transformation-based ALF (GALF) scheme is proposed to further improve the performance of ALF, which introduces geometric transformations, such as rotation, diagonal and vertical flip, to be applied to the samples in filter support region depending on the orientation of the gradient of the reconstructed samples before ALF. With the introduction of geometric transformations, more spatial adaptation is supported without excessive signaling of filter coefficients. The experimental results show that GALF outperforms the existing ALF techniques and it has been adopted by the JEM reference software used as the test platform for future video coding technology exploration in JVET. Marta Karczewicz, Li Zhang 0006, Wei-Jung Chien, Xiang Li 0003 |
PCS | 4 |
| 2016 | High Dynamic Range and Wide Color Gamut Video Coding in HEVC: Status and Potential Future EnhancementsabstractAs the video industry begins deployment of ultrahigh-definition TV in both professional and consumer markets, including support for higher dynamic range and wider color gamut services is considered essential within the industry. Higher dynamic range and wider color gamut offer end users a significantly enhanced viewing experience by supporting intensity ranges and colors unattainable in existing distribution ecosystems. In response to this trend, several standardization organizations have launched efforts to better enable these features in both short term and midterm. In this paper, we provide a survey of these standardization activities, with the specific goal of providing a summary of the underlying technologies. Our emphasis is on both existing and potential extensions to the High Efficiency Video Coding standard. Edouard François, Chad Fogg, Yuwen He, Xiang Li 0003, Ajay Luthra, C. Andrew Segall |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Resampling Process of the Scalable High Efficiency Video CodingabstractSHVC is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC) and spatial resampling process is inevitable module to support spatial scalability. This paper describes in details the resampling process, including both texture and motion data resampling in SHVC, and using experimental evidence, demonstrate their benefits in terms of coding efficiency. Jianle Chen, Elena Alshina, Xiang Li 0003, Marta Karczewicz, Alexander Alshin |
DCC | 3 |
| 2015 | Asymmetric 3D Lookup Table Based Color Gamut Scalability in SHVCabstractSHVC is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC). Color Gamut Scalability (CGS) refers to a scalable use case in which base layer and enhancement layer have different color gamuts. In this case, special inter-layer prediction is needed to improve coding efficiency in SHVC. In this paper, a solution based on asymmetric 3D lookup table is presented for color gamut scalability. Compared to SHVC without CGS coding tool, the proposed solution provides 9.6% - 16.1% overall luma BD-rate reduction in different test cases. Xiang Li 0003, Jianle Chen, Marta Karczewicz, Yuwen He, Yan Ye 0003, Cheung Auyeung |
DCC | 1 |
| 2014 | Region based inter-layer cross-color filtering for scalable extension of HEVCabstractInter-layer filtering is a key module of the emerging Scalable Extension of High Efficiency Video Coding Standard (SHVC). In SHVC, up-sampled based layer reconstructed pictures are used as inter-layer references to predict enhancement layer frames such that inter-layer redundancy is reduced. To improve the coding performance of inter-layer filtering, luma plane based chroma plane enhancement was proposed at picture level. However, the efficiency of the picture level adaptation is not very promising when picture resolution is high. To address this issue, region based inter-layer cross-color filtering is proposed in this paper. Simulations under the common test conditions defined by Joint Collaborative Team on Video Coding (JCT-VC) showed that significant chroma coding gain and moderate luma improvement were achieved by the proposed method. When compared to the luma plane based chroma plane enhancement method, the coding gain over SHVC reference software SHM-2.0 is about doubled while the decoding complexity is kept even lower. Moreover, the proposed method outperforms other tools studied in SHVC core experiment on inter-layer filtering. Xiang Li 0003, Jianle Chen, Marta Karczewicz, Elena Alshina, Alexander Alshin, Yongjin Cho |
ICIP | 1 |
| 2014 | Low-complexity advanced residual prediction design in 3D-HEVCabstractAdvanced residual prediction (ARP) is an efficient tool for 3D video coding by exploiting the residual correlation between views. In ARP, the residual predictor could be efficiently produced by aligning the motion information at the current view for motion compensation in the reference view. On the other hand, such on-the-fly residual predictor derivation process increases the complexity significantly due to the increased motion compensation steps. In this paper, a low-complexity ARP scheme is proposed. Experimental results demonstrate that the proposed scheme significantly reduces the decoding complexity of the original design in terms of both memory access and computational complexity while keeping comparable coding performance. Li Zhang 0006, Ying Chen 0011, Xiang Li 0003, Shanhua Xue |
ISCAS | 3 |
| 2013 | Scalable Video Coding Extension for HEVCabstractThis paper describes a scalable video codec that was submitted as a response to the joint call for proposals issued by ISO/IEC MPEG and ITU-T VCEG on HEVC scalable extension. The proposed codec uses a multi-loop decoding structure. Several inter-layer texture prediction methods are employed to remove the inter-layer redundancy. Inter-layer prediction is also used when coding enhancement layer syntax elements such as motion parameter and intra prediction mode, to further reduce bit overhead. Additionally, alternative transforms as well as adaptive coefficients scanning are used to code the prediction residues more efficiently. Experimental results are presented to demonstrate the effectiveness of the proposed scheme. When compared to HEVC single-layer coding, the additional rate overhead for the proposed scalable extension is 1.2% to 6.4% to achieve two layers of SNR and spatial scalability. Jianle Chen, Krishnakanth Rapaka, Xiang Li 0003, Vadim Seregin, Marta Karczewicz, Geert Van der Auwera, Joel Sole, Xianglin Wang, Chengjie Tu, Ying Chen 0011, Rajan L. Joshi |
DCC | 3 |
| 2013 | Generalized inter-layer residual prediction for scalable extension of HEVCabstractScalable video coding extension of HEVC (SHVC) is being developed by Joint Collaborative Team on Video Coding (JCT-VC) of ISO/IEC MPEG and ITU-T VCEG. Different from scalable video coding extension of H.264/AVC (SVC), SHVC employs a multi-loop decoding framework so that the inter-layer residual prediction in SVC does not perform well in SHVC. In this paper, a method called generalized inter-layer residual prediction (GILRP) is proposed. To improve prediction accuracy, the residual predictor is derived with the information from both base and enhancement layers. Moreover, three additional weighting types are introduced on top of inter coding modes to further compensate errors caused by base layer quantization. Simulations under SHVC common test conditions defined by JCT-VC show that 2.9%, 5.1% and 4.9% overall luma BD-rate reduction on average were obtained over SHVC reference software for configurations of random access, low delay with P slices, and low delay with B slices, respectively. Xiang Li 0003, Jianle Chen, Krishnakanth Rapaka, Marta Karczewicz |
ICIP | 1 |
| 2013 | Advanced residual predction in 3D-HEVabstractInter-view residual prediction (IVRP) is employed t o efficiently code non-base texture views by exploiting the correlation of residues between two views in 3D video coding extension of HEVC(3D-HEVC). To further improve the performance of texture coding, advanced residual prediction (ARP) is proposed in this paper. In IVRP, residues in a non-base view are predicted from decoded residues in base view. In contrast, ARP makes the residual predictor for the non-base view block based on newly generated base-view residues by applying motion vector at the non-base view to the baseview. Moreover, an adaptive weighting factor is applied to reduce prediction error. Comprehensive simulations show that up to 6.0% luma BD-rate reduction was obtained over IVRP in the reference software of 3D-HEVC. Xiang Li 0003, Li Zhang 0006, Ying Chen 0011 |
ICIP | 1 |
| 2013 | Inter-layer filtering for scalable extension of HEVCabstractThis paper introduces inter-layer filters for the scalable extension of High Efficiency Video Coding (SHVC) standard, which is being developed by the Joint Collaborative Team on Video Coding (JCT-VC). The major new coding tool in SHVC is inter-layer texture prediction. It provides about 18% average BD-rate reduction compared with HEVC two-layer simulcast. In the case of spatial scalability, base layer reconstructed pictures are up-sampled to the enhancement layer resolution to generate inter-layer texture prediction. A set of 2D separable 8 taps (luma) and 4 taps (chroma) DCT based interpolation filters, which follow the design principles of HEVC motion compensation interpolation filter, are used in the up-sampling process. In the case of SNR scalability, the up-sampling process is not needed since the reference layer has the same spatial resolution as the current layer but encoded with lower quality. This paper proposes a novel inter-layer filter with denoising effect for SNR scalability to improve enhancement layer coding efficiency and equalize the number of stages in inter-layer processing between SNR and spatial scalabilities. Experimental results show that the usage of inter-layer de-noising filter in SNR scalability provides up to 7.5% BD-rate reduction and has observable improvement on subjective visual quality. Elena Alshina, Alexander Alshin, Yongjin Cho, Jeong-Hoon Park, Jianle Chen, Xiang Li 0003, Vadim Seregin, Marta Karczewicz |
PCS | 7 |
| 2013 | High Frequency SAO for scalable extension of HEVCabstractScalable extension of HEVC, a.k.a. SHVC, is being standardized by the Joint Collaborative Team on Video Coding (JCT-VC). SHVC employs one of the most important coding tools called interlayer prediction, in which. Reconstructed base layer pictures can be used as reference pictures to predict enhancement layer pictures. Therefore, how to efficiently generate interlayer reference pictures to improve coding efficiency is one of the core research topics for the new international standard. In this paper, we presented High Frequency SAO filter (HF-SAO), which extends Sample Adaptive Offset filter (SAO) in HEVC to SHVC. Experimental results based on SHVC reference software version 1.0 show that HFSAO achieves 1.2% (Luma), 1.4% (Cb), 1.4% (Cr) average BD-rate reduction for the enhancement layer coding, which makes itself one of the most promising candidate interlayer filters to the new generation of scalable video coding standard. Jianle Chen, Krishnakanth Rapaka, Xiang Li 0003, Marta Karczewicz |
PCS | 4 |
| 2011 | Rate-Complexity-Distortion Optimization for Hybrid Video CodingabstractIn recent years, video applications on handheld devices became more and more popular. Due to limited computational capability and power supply in handheld devices, rate-complexity-distortion optimization (RCDO) algorithms at encoder side draw increasing attention. The target of RCDO is to obtain the best rate-distortion (R-D) performance under a constraint of complexity. Generally, there are three essential problems in RCDO. First, complexity needs to be properly mapped to a target in terms of coding parameters such that the control over complexity can be achieved. Second, the complexity budget should be efficiently distributed among frames or other coding units. Third, the allocated budget for each coding unit has to be effectively used to obtain good R-D performance. In this paper, these problems are well addressed. To obtain a large dynamic range in complexity control, medium-granularity control methods are presented. Then, a frame level complexity allocation algorithm is developed based on dependent rate-distortion function. Finally, an adaptive mode and reference searching method is proposed for motion compensation process. Comprehensive simulations verify the proposed algorithms. In the environment of the H.264/AVC reference software, an average gain of over 0.5 dB and 0.7 dB in BD-PSNR was achieved for nine sequences at low complexity when compared to two RCDO methods from literature. Moreover, experiments on x264 (a practical implementation of H.264/AVC) show that the proposed algorithms outperform predefined complexity levels by x264 in terms of both coding efficiency and computational scalability. Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Optimized channel rate allocation for H.264/AVC scalable video multicast streaming over heterogeneous networksabstractWe present an algorithm to optimize the allocation of channel bitrate to different network abstraction layer (NAL) units of the H.264/AVC scalable video bitstreams for real-time multicast streaming over heterogeneous networks. We focus on the problem of achieving a high robustness of video streaming under varying channel conditions in terms of the reconstructed video qualities at different users. As an extension of our previous work for unicast streaming, the proposed algorithm can achieve an optimized allocation of channel bitrate for multicast streaming with any user distribution. Our simulations show that a good performance on the video qualities among the multicast users can be achieved for different user distributions. A gain in terms of the overall multicast PSNR can be achieved against the protection strategies targeting at users with medium channel qualities in our experiments. Bin Zhang 0018, Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
ICIP | 2 |
| 2010 | Rate-complexity-distortion evaluation for hybrid video codingabstractTo objectively evaluate the coding efficiency of video codecs, Bj⊘ntegaard Delta PSNR (BD-PSNR) was proposed. Based on the rate-distortion (R-D) curve fitting, BD-PSNR is able to provide a good evaluation of the R-D performance. However, BD-PSNR has a critical drawback: It doesn't take the coding complexity into account. Clearly for practical video applications, especially for those on handheld devices, coding complexity has to be considered when evaluating the overall coding performance. Therefore in this paper, a new coding efficiency measurement is developed by generalizing BD-PSNR from R-D curve fitting to rate-complexity-distortion (R-C-D) surface fitting. Simulations show that a comprehensive performance evaluation can easily be obtained with the proposed method. Moreover, the idea can be used for rate-distortion optimization for complexity-constrained video coding. Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
ICME | 1 |
| 2010 | Adaptive quantization parameter cascading for hierarchical video codingabstractQuantization parameter (QP) cascaded hierarchical prediction structures have been proved as efficient techniques in hybrid video coding. However, the current QP cascading method is empirical and not adaptive. The reason for the higher coding efficiency of this method has not been fully explored so far. In this paper, the rate-distortion performance of QP cascaded hierarchical video coding is first analyzed with dependent rate-distortion function. Then the optimal offset in linear QP cascading scheme is derived theoretically. It is shown that the widely accepted empirical QP cascading method is actually an approximation to the theoretical solution in fast movement environment. For slow sequences, an average gain of 0.43 dB can be achieved by the combination of two proposed adaptive algorithms. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
ISCAS | 1 |
| 2010 | Medium-granularity computational complexity control for H.264/AVCabstractToday, video applications on handheld devices become more and more popular. Due to limited computational capability of handheld devices, complexity constrained video coding draws much attention. In this paper, a medium-granularity computational complexity control (MGCC) is proposed for H.264/AVC. First, a large dynamic range in complexity is achieved by taking 16×16 motion estimation in a single reference frame as the basic computational unit. Then a high coding efficiency is obtained by an adaptive computation allocation at MB level. Simulations show that coarse-granularity methods cannot work when the normalized complexity is below 15%. In contrast, the proposed MGCC performs well even when the complexity is reduced to 8.8%. Moreover, an average gain of 0.3 dB over coarse-granularity methods in BD-PSNR is obtained for 11 sequences when the complexity is around 20%. Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
PCS | 1 |
| 2009 | One-pass multi-layer rate-distortion optimization for quality scalable video codingabstractIn this paper, a one-pass multi-layer rate-distortion optimization algorithm is proposed for quality scalable video coding. To improve the overall coding efficiency, the MB mode in the base layer is selected not only based on its rate-distortion performance relative to this layer but also according to its impact on the enhancement layer. Moreover, the optimization module for residues is also improved to benefit inter-layer prediction. Simulations show that the proposed algorithm outperforms the most recent SVC reference software. For eight test sequences, a gain of 0.35 dB on average and 0.75 dB at maximum is achieved at a cost of less than 8% increase of the total coding time. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
ICASSP | 1 |
| 2009 | Model based analysis for quantization parameter cascading in hierarchical video codingabstractOriginally, hierarchical prediction structures were proposed to achieve temporal scalability. Soon after, it was realized that with a proper quantization parameter cascading (QPC) scheme the general performance can be significantly improved by hierarchical coding. However, the theory behind the gain has not been explored so far. In this paper, the QPC in hierarchical coding is investigated by model based emulations. From the analysis, it is noticed that a parameter β which represents the error propagation in a group of pictures greatly affects the performance of QPC in hierarchical coding. Based on β, a simple adaptive QPC algorithm is designed. Simulations verify the efficiency of this algorithm: a gain up to 0.89 dB is obtained over the most recent SVC reference software. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 1 |
| 2009 | One-pass frame level budget allocation in video coding using inter-frame dependencyabstractIn this paper, a one-pass budget allocation algorithm is proposed for hybrid video coding. Taking the percentage of skipped MBs as the measure of inter-frame dependency, the optimal budget allocation is first modeled for a two-frame case. Then this model is extended to a practical method in slow movement scenario, where the information of inter-frame dependency is predicted based on the previously coded frames. Simulations show that the proposed algorithm outperforms the recommended MB and frame level rate control algorithms in H.264/AVC reference software JM 15.1. For eight CIF sequences with slow movement, significant gains of 0.91 dB and 0.43 dB on average (1.66 dB and 1.49 dB at maximum) were obtained over the reference MB and frame level rate control, respectively. Considering that the computational complexity by the proposed algorithm is quite low (less than 1% of the total coding time when fast motion estimation algorithm is enabled), it is quite appealing for real-time video applications. Xiang Li 0003, Andreas Hutter, André Kaup |
MMSP | 1 |
| 2009 | Lagrange multiplier selection for rate-distortion optimization in SVCabstractThe Lagrangian multiplier based rate-distortion optimization (RDO) has been widely employed in single layer video coding. During the development of scalable video coding (SVC) extension of H.264/AVC, it was directly applied in a multilayer scenario. However, such an application is not very efficient since the correlation between layers is not considered in the Lagrange multiplier selection. To improve the overall performance, in this paper a new selection algorithm is presented for RDO in SVC. Simulations show that the proposed method outperforms the recent SVC reference software. With a tiny computational cost, average gains of 0.22 dB and 0.35 dB were achieved in the tests of four-layer quality scalability and three-layer spatial scalability, respectively. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
PCS | 1 |
| 2009 | Efficient one-pass frame level rate control for H.264/AVC
Xiang Li 0003, Andreas Hutter, André Kaup |
J. Vis. Commun. Image Represent. | 1 |
| 2009 | Laplace Distribution Based Lagrangian Rate Distortion Optimization for Hybrid Video CodingabstractIn today's hybrid video coding, Rate-Distortion Optimization (RDO) plays a critical role. It aims at minimizing the distortion under a constraint on the rate. Currently, the most popular RDO algorithm for one-pass coding is the one recommended in the H.264/AVC reference software. It, or HR-$\lambda $for convenience, is actually a kind of universal method which performs the optimization only according to the quantization process while ignoring the properties of input sequences. Intuitively, it is not efficient all the time and an adaptive scheme should be better. Therefore, a new algorithm Lap-$\lambda $is presented in this paper. Based on the Laplace distribution of transformed residuals, the proposed Lap-$\lambda $is able to adaptively optimize the input sequences so that the overall coding efficiency is improved. Cases which cannot be well captured by the proposed models are considered via escape methods. Comprehensive simulations verify that compared with HR-$\lambda $, Lap-$\lambda $shows a much better or similar performance in all scenarios. Particularly, significant gains of 1.79 dB and 1.60 dB in PSNR are obtained for slow sequences and B-frames, respectively. Xiang Li 0003, Norbert Oertel, Andreas Hutter, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Advanced Lagrange Multiplier Selection for Hybrid Video CodingabstractThe Lagrangian multiplier based rate-distortion optimization has been proved to be an effective way in hybrid video coding. In this paper, an advanced Lagrange multiplier selection method is presented. Based on Laplace distribution, the variance of transformed residuals is introduced into the rate and distortion models. Moreover, inspired by the ρ-domain method, the percentage of non-zeros among quantized residuals is also considered to refine the rate model. Thanks to the more accurate models, the proposed method is able to adaptively optimize the encoding process for videos with different properties. It outperforms the algorithm used in the reference software of H.264/AVC. According to the simulation, a gain up to 1.1dB was achieved. Xiang Li 0003, Norbert Oertel, Andreas Hutter, André Kaup |
ICME | 1 |
| 2007 | Adaptive Lagrange Multiplier Selection for Intra-Frame Video CodingabstractThe Lagrangian technique proves to be an effective way in Rate-Distortion optimization for hybrid video coding. In this paper, an new Lagrange multiplier selection method for Intra-Frame coding is presented. Based on Laplacian distribution, the variance of transformed residual coefficients is introduced into the rate and distortion model, so that the proposed method is able to adaptively optimize the encoding process for different types of videos. It outperforms the algorithm used in the reference software of H.264/AVC, especially for computer animations and videos with small movements. According to the simulations, a gain up to 0.3 dB was achieved. Xiang Li 0003, Norbert Oertel, André Kaup |
ISCAS | 1 |
| 2006 | Gradient Intra Prediction for Coding of Computer Animated VideosabstractThe increasing interest in computer animations initiated the need of efficient coding for such applications. This paper proposes two gradient based approaches to improve the efficiency of Intra Prediction. Taking advantage of a special property of computer animations, namely the gradient distribution of intensity on object surfaces, the devised methods provided a better performance over H.264/AVC. According to the simulations, up to 0.3 dB gain in PSNR has been achieved. Xiang Li 0003, Norbert Oertel, André Kaup |
MMSP | 1 |
| 2004 | Fast multi-frame motion estimation algorithm with adaptive search strategies in H.264abstractIn the new H.264/AVC video coding standard, motion estimation takes up a significant encoding time, especially when using the straightforward full search algorithm (FS). A fast flexible multi-frame motion estimation algorithm with adaptive search strategies (FMASS) is presented. With special considerations on the multiple reference frames and block modes, several techniques, i.e., adaptive search strategies for single frame and flexible multi-frame selection, have been utilized to improve significantly the speed-up performance in H.264. Extensive simulations show that it can minimize the matching points by more than 1190 times compared with FS. In addition, the output quality of the encoded sequences loses only 0.053 dB in terms of PSNR on average. Fast speed-up performance and unnoticeable quality losses make the proposed algorithm outperform most of the other well-known algorithms proposed recently, such as ARPS3, MVFAST and UMHexagonS, of which the latter two have been already accepted by MPEG-4 and JVT respectively. Xiang Li 0003, Eric Q. Li, Yen-Kuang Chen |
ICASSP (3) | 1 |