EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyu Xiu
dblp:64/8116
· DBLP profile ↗
8ranked-venue papers in the field
2as first author
2since 2021 · last 2023
—ORCID · none
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Gradient Linear Model for Chroma Intra PredictionabstractIn Versatile Video Coding (VVC), Cross-component Linear Model (CCLM) predicts chroma samples by assuming a linear relationship between luma and chroma components. In performing CCLM for video in YUV 4:2:0 chroma format, collocated luma samples are firstly downsampled by a low-pass filter to match luma resolution with chroma, and one linear model of luma-chroma sample pairs is applied on the reconstructed luma samples to generate the predicted chroma samples. However, the low-pass downsampling procedure ignores relative spatial variations among luma samples in proximity, such as edge and gradient information. To solve this issue, a new coding technique, namely gradient linear model (GLM), is proposed for further compression efficiency exploration beyond VVC. Instead of using a low-pass filter in CCLM, the GLM utilizes high-pass gradient filters to generate the downsampled luma values. In this paper, two GLM schemes are provided with different trade-offs between coding gain and complexity, including: 1) a 2-parameter scheme that shares the CCLM module framework but replaces the downsampling filter with high-pass gradient filters; 2) a 3-parameter scheme that further combines the luma gradients with the low-pass downsampled luma values. Based on the enhanced compression model (ECM-5.0) software from the joint video experts team (JVET), simulation results show that the 2-parameter GLM achieves average Bjontegaard delta-rate (BD-rate) savings of {1.01%, 1.66%, 1.81%} and {0.69%, 0.95%, 1.12%} for {Y, U, V} components under the All Intra and Random Access configurations, respectively, and the 3-parameter GLM provides {1.28%, 3.23%, 3.28%} and {0.92%, 2.19%, 2.26%} BD-rate savings for {Y, U, V} components under the All Intra and Random Access configurations, respectively. Both of the proposed GLM schemes have been adopted to the ECM software platform. Che-Wei Kuo, Xiaoyu Xiu, Hong-Jheng Jhu, Xianglin Wang, Yan Ye 0003, Jie Chen 0006, Ru-Ling Liao |
DCC | 3 |
| 2022 | Cross-component Sample Adaptive OffsetabstractThis paper proposes one new In-loop filtering technique cross-component sample adaptive offset (CCSAO) for further coding efficiency improvement beyond Versatile Video Coding (VVC). The CCSAO reduces the sample distortion by 1) utilizing the strong correlation between luma and chroma components to classify the reconstructed samples into different categories and 2) deriving one offset for each category and adding the offset to the samples in the category. The offset of each category is properly derived at encoder and signaled to decoder. To keep the design at low complexity, only band information of reconstructed samples is considered for the sample classification of the CCSAO. To verify the performance, the proposed CCSAO is implemented on top of the enhanced compression model (ECM) for the joint video exploration team (JVET)'s exploratory work of future video coding technologies beyond VVC. Simulation results show that the CCSAO achieves average {0.20%, 2.83%, 2.98%} and {0.41%, 7.36%, 7.36%} Bjentegaard delta (BD)-rate savings for {Y, U, V} components under the Random Access and Low Delay B configuration, with negligible complexity impacts on encoding and decoding complexity. The proposed CCSAO scheme has been adopted to the ECM-2.0 software platform. Che-Wei Kuo, Xiaoyu Xiu, Yi-Wen Chen, Hong-Jheng Jhu, Xianglin Wang |
DCC | 2 |
| 2019 | Improved Video Coding Techniques for Next Generation Video Coding StandardabstractThis paper describes a video coding scheme submitted in response to the joint call for proposals (CfP) on video compression for capability beyond HEVC issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11(MPEG) in October 2017. It includes video coding techniques for the standard dynamic range (SDR) and high dynamic range (HDR) categories. Design of the core SDR codec in the response is based on the joint exploration model (JEM) reference software developed by the joint video exploration team (JVET). Some of key coding tools in the JEM are significantly simplified to reduce both average and worst-case complexity for hardware design with negligible coding performance loss. Furthermore, two additional coding technologies, namely multi-type tree (MTT) and decoder-side intra mode derivation (DIMD), are used to further improve coding efficiency. For the HDR category, besides the tools used in SDR category, two additional coding tools: an in-loop reshaper and a luma-based QP prediction method are used to further improve HDR coding efficiency. Simulation results demonstrate the high coding efficiency achieved by the proposed video codec at the expense of moderate coding complexity over HEVC. For random access configuration, it achieves average bit rate savings of 35.7% and 4.00% over the HM and JEM anchors with decoding time of 263% and 33%, respectively, for the SDR sequences. For the HDR sequences, the proposed in-loop reshaper is configured to maximize HDR objective metrics, it achieves average bit rate savings of 31.3% and 4.6% over the HM and JEM for wPSNRY metrics for the HDR PQ content. Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Rahul Vanam, Philippe Hanhart, Taoran Lu, Fangjun Pu, Peng Yin 0002, Walt Husak, Tao Chen 0044 |
DCC | 1 |
| 2018 | Hybrid Cubemap Projection Format for 360-Degree Video Codingabstract360-degree video has become popular in recent years with the advances in virtual reality (VR) and augmented reality (AR) technologies and has been rapidly commercialized. To provide viewers with an immersive experience, 360-degree video requires higher resolution and much higher bandwidth compared with conventional 2D video. In a typical 360-degree video compression and delivery framework, the stitched input 360-degree videos, represented in a native projection format, e.g., equirectangular (ERP), are converted into another projection format, e.g., cubemap (CMP), octahedron (OHP), etc. and frame packed before being fed into existing video codecs. The intermediate projection format is important and would potentially improve the representation efficiency and coding performance. Among all the projection solutions, CMP is very popular and has been widely used in the computer graphics community. The intrinsic rectilinear properties of the CMP format are advantageous for the translational motion model in the modern codec architecture. However, in the CMP representation, the samples on the sphere are not evenly distributed within the faces, resulting in a higher density near the face boundaries and a lower density near the face center. Such non-uniform sampling scheme penalizes the video representation efficiency and degrades the coding performance. Adjusted cubemap projection (ACP) was proposed to address such non-uniform sampling by introducing transform functions to improve the sampling uniformity. However, the transform function parameters in ACP are fixed regardless of the content inside each cube face. In this paper, a generalized hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency beyond ACP. HCP is defined by a pair of forward transform and inverse transform functions with a pair of horizontal and vertical transform parameters per cube face. The encoder can choose the optimal sampling for each face by adjusting the parameters in the horizontal and vertical directions based on the 360-degree video content characteristics inside each cube face. In order to maintain the boundary continuities between two neighboring faces, in a 3x2 packing layout, vertical parameter constraints are imposed such that faces in each face-row have the same vertical parameters. The HCP parameters are chosen to minimize the end-to-end weighted conversion error and determined using iterative search between the horizontal and the vertical directions. Significant changes in HCP parameter values can cause drastic change in sampling distribution, and may affect the inter-picture coding efficiency. Therefore, an efficient HCP parameter estimation algorithm is proposed to achieve a better trade-off between the temporal sampling adaptation and the inter-picture prediction efficiency by reducing the temporal variation of HCP parameters. The proposed HCP parameter search algorithm reduces the computational complexity by 5x compared to the exhaustive search method. The HCP parameters are selected by the encoder using the first picture of each Intra Random-Access Point (IRAP) and signalled once per IRAP. In SPS, projection format, frame packing parameters including number of faces in horizontal and vertical directions and each face's position and orientation are signalled. In PPS, the horizontal and vertical HCP parameters in 6-bit precision are encapsulated. The proposed HCP solution is implemented upon JEM-6.0 and 360Lib-3.0 software. Simulation results are reported using the test conditions specified in the JVET Call-for-Evidence (CfE) document. Compared with the CMP and ACP formats, the proposed HCP format demonstrates average 3.0 dB (up to 3.6 dB) and 0.2dB (up to 0.4 dB) End-to-End WS-PSNR improvement for the luma (Y) component, respectively, and average luma (Y) BD-rate reductions of 11.5% (up to 23.0%) and 0.5% (up to 1.0%), respectively. Fanyi Duanmu, Yuwen He, Xiaoyu Xiu, Philippe Hanhart, Yan Ye 0003, Yao Wang 0001 |
DCC | 3 |
| 2017 | Geometry Padding for Motion Compensated Prediction in 360 Video Codingabstract360 Video has become popular in recent years, as commercial interests in deploying Virtual Reality (VR) applications rise. This type of video is usually captured using multi-camera arrays, such as the GoPro Omni camera rig. After separate video streams are captured from multiple cameras, image stitching is applied to obtain a spherical representation of the scene, which spans 360 degrees horizontally and 180 degrees vertically, hence the name 360 video. In the existing workflow of 360 spherical video coding, the 360 video is projected onto the 2D plane with a projection format, such as equirectangular (ERP), cubemap (CMP), equal-area (EAP), octahedron (OHP), etc. Most, if not all, of the currently available 360 video content are provided in ERP format defined in longitude and latitude. Projection format conversion may be performed to convert the native ERP format to another format before coding is applied. Some projection formats contain more than one face, for example, CMP projects the sphere onto a cube of six faces or OHP projects the sphere onto an octahedron of eight faces. For these multi-face projection formats, the faces are packed onto a 2D rectangular picture with a frame packing method. For example, the six faces of CMP can be packed with 4x3 configuration, or 3x2 configuration. Finally, the frame packed picture is coded as a 2D conventional video. Existing video codecs are designed only considering conventional 2D video captured on a plane. When motion compensated prediction uses any samples outside of a reference picture's padding will be performed by simply copying the sample values from the picture boundaries. This repetitive padding method is referred as conventional 2D padding method, which is widely used in video coding standards such as H.264, High Efficiency Video Coding (HEVC). However, a 360 video encompasses video information on the whole sphere, and thus intrinsically has a cyclic property. When considering this cyclic property, the reference pictures of a 360 video no longer have boundaries, as the information they contain is all wrapped around a sphere. This cyclic property holds regardless of which projection format or which frame packing is used to represent the 360 video on a 2D plane. The paper presents a new geometry padding method for motion compensated prediction in 360 video coding. Unlike the conventional padding method for 2D video coding, the proposed geometry padding method extends samples outside of a 2D picture's boundaries using neighboring samples on the sphere. The geometry projection format is considered when performing padding. The corresponding sample outside of a face's boundary (which may come from another side in the same face or from another face), is derived with rectilinear projection. Each face is extended with geometry padding separately. When visualized, the extended faces using geometry padding show continuous texture representing natural extension of the texture inside the face. The proposed geometry padding method is implemented in the HEVC reference software HM-16.12 for the ERP and CMP projection formats. In the simulation, a total of sixteen 4K ERP video and eight 8K ERP video are used. For 8K ground truth 8K video, they are converted to 4K video in ERP and CMP projection formats, coded, and converted back to reconstructed 8K video in ERP format. For 4K ground truth video, they are directly coded in 4K ERP, or converted to CMP consisting of 75% of effective samples, coded, and converted back to reconstructed 4K video in ERP format. Then, the end-to-end spherical PSNR (S-PSNR) is calculated between the original 8K or 4K and the reconstructed 8K or 4K ERP video. BD-rate is calculated between the reference unmodified HEVC, which uses the conventional padding method, and HEVC modified with the proposed geometry padding method. Simulation results showed that geometry padding performs better. For 8K sequences, the proposed geometry padding gives on average luma (Y) BD-rate reduction of 0.3% for ERP and 0.8% for CMP, for 4K sequences, the proposed geometry padding gives on average Y BD-rate reduction of 0.2% for ERP and 1.0% for CMP. Comparing the gains in ERP format with the gains in CMP format, the improvement for CMP is larger. This is because CMP has six faces, therefore the improvement from geometry padding method affects more out-of-boundary samples. The proposed geometry padding method is also especially effective for sequences with fast motion. For example, it achieves BD rate reductions of 4.3%, 2.7%, 1.9%, and 2.5% for Glacier, Chairlift, Sb_in_lot, and Driving, respectively. These four sequences are all captured using moving cameras and have fast moving objects. As a result, the sequences contain a lot of across-the-face-boundary motion which can benefit from improved padding method. Detailed simulation results can be found in JVET contribution JVET-D0075 available at http://phenix.int-evry.fr/jvet/doc_end_user/documents/4_Chengdu/wg11/JVET-D0075-v3.zip. Yuwen He, Yan Ye 0003, Philippe Hanhart, Xiaoyu Xiu |
DCC | 4 |
| 2015 | Palette-Based Coding in the Screen Content Coding Extension of the HEVC StandardabstractThis paper provides a technical overview of palette-based coding that was adopted into the test model for the screen content coding (SCC) extension of High Efficiency Video Coding (HEVC) standard at the 18th JCT-VC meeting. Key techniques that enable the palette mode to deliver significant coding gains for screen contents are highlighted, including palette table generation, palette table coding, and the coding methods for palette indices and escape colors. Proposed and adopted techniques up to the first version of the working draft of HEVC SCC extension and test model SCM-2.0 are presented. Experimental results are provided to evaluate the performance of the palette mode in the SCC extension of HEVC. Xiaoyu Xiu, Yuwen He, Rajan L. Joshi, Marta Karczewicz, Patrice Onno, Christophe Gisquet, Guillaume Laroche |
DCC | 1 |
| 2015 | Adaptive Color-Space Transform for HEVC Screen Content CodingabstractThis paper presents an in-loop adaptive color-space transform for the HEVC Screen Content Coding extension. In the proposed method, the prediction residual is adaptively converted into a different color space to reduce the cross-component redundancy. After the ACT, the signal is coded following the existing HEVC framework. To keep the complexity as low as possible, fixed color-space transforms that are easily implemented with shift and add operations are utilized. Significant coding gains are achieved by this method in the current HEVC Screen Content Coding reference software with no increase of decoding runtime. The proposed method has been adopted to the HEVC Screen Content Coding extension. Li Zhang 0006, Jianle Chen, Joel Sole, Marta Karczewicz, Xiaoyu Xiu, Ji-Zheng Xu |
DCC | 5 |
| 2014 | Improved Inter-Layer Prediction for the Scalable Extensions of HEVCabstractSummary form only given. Upon the completion of the single-layer H.265/HEVC, scalable extensions of the H.265/HEVC standard, called Scalable High Efficiency Video Coding (SHVC), are currently under development. Compared to the simulcast solution that simply compresses each layer separately, SHVC offers higher coding efficiency by means of inter-layer prediction which is implemented by inserting inter-layer reference (ILR) pictures generated from reconstructed base layer (BL) pictures into the enhancement layer (EL) decoded picture buffer (DPB) for motion-compensated prediction of the collocated pictures in the EL. If the EL has a higher resolution than that of the BL, the reconstructed BL pictures need to be up-sampled to form the ILR pictures. Given that the ILR picture is generated based on the reconstructed BL picture, its suitability for an efficient inter-layer prediction may be limited due to the following reasons. Firstly, quantization is usually applied when coding the BL pictures. Quantization causes the BL reconstructed texture to contain undesired coding artifacts, such as blocking artifacts, ringing artifacts, and color artifacts. Secondly, in case of spatial scalability, a down-sampling process is used to create the BL pictures. To reduce aliasing, the high frequency information in the video signal is typically removed by the down-sampling process. As a result, the texture information in the ILR picture lacks certain high frequency information. In contrast to the ILR picture, the EL temporal reference pictures contain plentiful high frequency information, which could be extracted to enhance the quality of the ILR picture. To further improve the efficiency of inter-layer prediction, a low pass filter may be applied to the ILR picture to alleviate the quantization noise introduced by the BL coding process. In this paper, an ILR enhancement method is proposed to improve the quality of the ILR picture by combining the high frequency information extracted from the EL temporal reference pictures together with the low frequency information extracted from the ILR picture. Experimental results show that the proposed method can significantly increase the ILR efficiency for EL coding, under the Common Test Condition of SHVC, which defines a number of temporal prediction structures called Random Access (RA), Low-delay B (LD-B) and Low-delay P (LD-P), on average the proposed method provides {Y, U, V} BD-rate (BL+EL) gains of {2.0%, 7.1%, 8.2%}, {2.2%, 6.7%, 7.6%} and {4.0%, 7.4%, 8.4%} for RA, LD-B, and LD-P, respectively, in comparison to the performance of the SHVC reference software SHM-2.0. Thorsten Laude, Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Jörn Ostermann |
DCC | 2 |