Jianle Chen

dblp:73/2298 · DBLP profile ↗
← Back
47ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-4196-3754ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Extension of Semi-Decoupled Partitioning in Inter Frames
abstract
The Alliance for Open Media (AOMedia) has been exploring new coding tools to enhance AV1 capabilities. Semi-Decoupled Partitioning (SDP), originally designed for intra frames in research-v2.0.0, improves coding by decoupling luma and chroma block partitioning. This study extends SDP to inter frames by introducing intra region coding, where the root node is explicitly signaled in the bitstream. Within the intra region, luma components of the intra-coded blocks can be further split, while chroma components remain unsplit. The experiments are implemented on the 8thanchor, research-v8.0.0, of AVM reference software with CTCv7, and experimental results show that the proposed method can achieve 0.12%, 2.39%, 2.58% coding gain for Y, U, and V component separately with random access configuration and 5% encoding time increase and almost no decoding time increase.
Madhu Peringassery Krishnan, Shan Liu 0001, Jayasingam Adhuran, Minhao Tang, Jianle Chen, Urvang Joshi, Mohammed Golam Sarwer, Debargha Mukerjee
ICIP6
2025 Google Industry Seminar: Video Processing in the New Age of AI
abstract
Video processing and compression are being re-envisioned in the age of AI. Traditional video codecs, which rely on rigid, pre-defined rules, are being augmented and, in some cases, replaced by AI-driven approaches. These new methods leverage machine learning to intelligently analyze video content, allowing for more adaptive and efficient compression. We are going to discuss AOM's new codec AV2, its low- and high-level features that enable significantly smaller file sizes with no perceptible loss in quality, a crucial development for streaming and storage. The shift to AI has also transformed how we evaluate video quality. Traditional metrics, while directionally useful, don't always align with human perception, especially for user generated content (UGC). They fail to capture what's most important for machine vision tasks. We will talk about new AI-based quality metrics that are being developed. They correlate better with a human's subjective experience and a machine's ability to perform tasks like object recognition. Along the way, we'll cover large scale industrial infrastructure challenges and the ways to achieve high reliability and accuracy.
Balu Adsumilli, Jianle Chen, In Suk Chong, Yilin Wang 0001
ACM Multimedia2
2022 Dynamic Point Cloud Interpolation
abstract
Dense photorealistic point clouds can depict real-world dynamic objects in high resolution and with a high frame rate. Frame interpolation of such dynamic point clouds would enable the distribution, processing, and compression of such content. In this work, we propose a first point cloud interpolation framework for photorealistic dynamic point clouds. Given two consecutive dynamic point cloud frames, our framework aims to generate intermediate frame(s) between them. The proposed deep learning framework has three major components: the encoder module, the fusion network, and the multi-scale point cloud synthesis module. The encoder module extracts multi-scale features from two consecutive frames. The fusion network employs a novel 4D feature learning technique to merge the multi-scale features from consecutive frames. Finally, the multi-scale point cloud synthesis module hierarchically reconstructs the interpolated point cloud intermediate frame at different resolutions. We evaluate our framework on high-resolution point cloud datasets used in MPEG, JPEG Pleno, and AVS standards. The quantitative and qualitative results demonstrate the effectiveness of the proposed method.
Anique Akhtar, Zhu Li 0001, Geert Van der Auwera, Jianle Chen
ICASSP4
2022 PU-Dense: Sparse Tensor-Based Point Cloud Geometry Upsampling
abstract
Due to the increased popularity of augmented and virtual reality experiences, the interest in capturing high-resolution real-world point clouds has never been higher. Loss of details and irregularities in point cloud geometry can occur during the capturing, processing, and compression pipeline. It is essential to address these challenges by being able to upsample a low Level-of-Detail (LoD) point cloud into a high LoD point cloud. Current upsampling methods suffer from several weaknesses in handling point cloud upsampling, especially in dense real-world photo-realistic point clouds. In this paper, we present a novel geometry upsampling technique, PU-Dense, which can process a diverse set of point clouds including synthetic mesh-based point clouds, real-world high-resolution point clouds, real-world indoor LiDAR scanned objects, as well as outdoor dynamically acquired LiDAR-based point clouds. PU-Dense employs a 3D multiscale architecture using sparse convolutional networks that hierarchically reconstruct an upsampled point cloud geometry via progressive rescaling and multiscale feature extraction. The framework employs a UNet type architecture that downscales the point cloud to a bottleneck and then upscales it to a higher level-of-detail (LoD) point cloud. PU-Dense introduces a novel Feature Extraction Unit that incorporates multiscale spatial learning by employing filters at multiple sampling rates and receptive fields. The architecture is memory efficient and is driven by a binary voxel occupancy classification loss that allows it to process high-resolution dense point clouds with millions of points during inference time. Qualitative and quantitative experimental results show that our method significantly outperforms the state-of-the-art approaches by a large margin while having much lower inference time complexity. We further test our dataset on high-resolution photo-realistic datasets. In addition, our method can handle noisy data well. We further show that our approach is memory efficient compared to the state-of-the-art methods.
Anique Akhtar, Zhu Li 0001, Geert Van der Auwera, Li Li 0040, Jianle Chen
IEEE Trans. Image Process.5
2021 Gradient Compression with a Variational Coding Scheme for Federated Learning
abstract
Federated Learning (FL), a distributed machine learning architecture, emerged to solve the intelligent data analysis on massive data generated at network edge-devices. With this paradigm, a model is jointly learned in parallel at edge-devices without needing to send voluminous data to a central FL server. This not only allows a model to learn in a feasible duration by reducing network latency but also preserves data privacy. Nonetheless, when thousands of edge-devices are attached to an FL framework, limited network resources inevitably impose intolerable training latency. In this work, we propose model-update compression to solve this issue in a very novel way. The proposed method learns multiple Gaussian distributions that best describe the high dimensional gradient parameters. In the FL server, high dimensional gradients are repopulated from Gaussian distributions utilizing likelihood function parameters which are communicated to the server. Since the distribution information parameters constitute a very small percentage of values compared to the high dimensional gradients themselves, our proposed method is able to save significant uplink band-width while preserving the model accuracy. Experimental results validated our claim.
Birendra Kathariya, Zhu Li 0001, Jianle Chen, Geert Van der Auwera
VCIP3
2021 Developments in International Video Coding Standardization After AVC, With an Overview of Versatile Video Coding (VVC)
abstract
In the last 17 years, since the finalization of the first version of the now-dominant H.264/Moving Picture Experts Group-4 (MPEG-4) Advanced Video Coding (AVC) standard in 2003, two major new generations of video coding standards have been developed. These include the standards known as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). HEVC was finalized in 2013, repeating the ten-year cycle time set by its predecessor and providing about 50% bit-rate reduction over AVC. The cycle was shortened by three years for the VVC project, which was finalized in July 2020, yet again achieving about a 50% bit-rate reduction over its predecessor (HEVC). This article summarizes these developments in video coding standardization after AVC. It especially focuses on providing an overview of the first version of VVC, including comparisons against HEVC. Besides further advances in hybrid video compression, as in previous development cycles, the broad versatility of the application domain that is highlighted in the title of VVC is explained. Included in VVC is the support for a wide range of applications beyond the typical standard- and high-definition camera-captured content codings, including features to support computer-generated/screen content, high dynamic range content, multilayer and multiview coding, and support for immersive media such as 360° video.
Benjamin Bross, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Ye-Kui Wang
Proc. IEEE2
2021 Guest Editorial Introduction to the Special Section on the VVC Standard
abstract
In this Special Section of the IEEE Transactions on Circuits and Systems for Video Technology, it is our honor to introduce the Versatile Video Coding (VVC) standard, the latest of the historic partnership collaborations between the International Telecommunication Union Telecommunication Standardization Sector (ITU-T), the International Organization for Standardization (ISO), and the International Electrotechnical Commission (IEC) in the field of video coding standardization.
Jill M. Boyce, Jianle Chen, Shan Liu 0001, Jens-Rainer Ohm, Gary J. Sullivan, Thomas Wiegand 0001, Yan Ye 0003, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Overview of the Versatile Video Coding (VVC) Standard and its Applications
abstract
Versatile Video Coding (VVC) was finalized in July 2020 as the most recent international video coding standard. It was developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG) to serve an ever-growing need for improved video compression as well as to support a wider variety of today’s media content and emerging applications. This paper provides an overview of the novel technical features for new applications and the core compression technologies for achieving significant bit rate reductions in the neighborhood of 50% over its predecessor for equal video quality, the High Efficiency Video Coding (HEVC) standard, and 75% over the currently most-used format, the Advanced Video Coding (AVC) standard. It is explained how these new features in VVC provide greater versatility for applications. Highlighted applications include video with resolutions beyond standard- and high-definition, video with high dynamic range and wide color gamut, adaptive streaming with resolution changes, computer-generated and screen-captured video, ultralow-delay streaming, 360° immersive video, and multilayer coding e.g., for scalability. Furthermore, early implementations are presented to show that the new VVC standard is implementable and ready for real-world deployment.
Benjamin Bross, Ye-Kui Wang, Yan Ye 0003, Shan Liu 0001, Jianle Chen, Gary J. Sullivan, Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.5
2021 Decoder-Side Motion Vector Refinement in VVC: Algorithm and Hardware Implementation Considerations
abstract
This paper presents an overview of the decoder-side motion vector refinement (DMVR) algorithm in the Versatile Video Coding (VVC) standard. The proposed DMVR algorithm aims to increase the prediction accuracy of the blocks coded in merge mode using the bilateral matching-based refinement method. Compared with previous decoder-side motion vector derivation approaches, the proposed method significantly increases the coding efficiency without signaling additional side information. Furthermore, the hardware implementation considerations of the DMVR design are particularly focused in this study. This paper details and analyzes the novel features of DMVR contributing to the increase in coding efficiency and the reduction in computational complexity and implementation difficulty. Experimental results based on the VVC test model version 8.0 demonstrate that average Bjøntegaard Delta rate savings of 0.80 % and 2.81 % are achieved for the “tool-off” and “tool-on” test configurations, respectively. Moreover, 4 % additional decoding time and negligible additional external memory bandwidth requirements of DMVR based on the common test conditions for VVC are reported.
Han Gao 0001, Semih Esenlik, Jianle Chen, Eckehard G. Steinbach
IEEE Trans. Circuits Syst. Video Technol.4
2021 Intra Prediction and Mode Coding in VVC
abstract
This paper presents the intra prediction and mode coding of the Versatile Video Coding (VVC) standard. This standard was collaboratively developed by the Joint Video Experts Team (JVET). It follows the traditional architecture of a hybrid block-based codec that was also the basis of previous standards. Almost all intra prediction features of VVC either contain substantial modifications in comparison with its predecessor H.265/HEVC or were newly added. The key aspects of these tools are the following: 65 angular intra prediction modes with block shape-adaptive directions and 4-tap interpolation filters are supported as well as the DC and Planar mode, Position Dependent Prediction Combination is applied for most of these modes, Multiple Reference Line Prediction can be used, an intra block can be further subdivided by the Intra Subpartition mode, Matrix-based Intra Prediction is supported, and the chroma prediction signal can be generated by the Cross Component Linear Model method. Finally, the intra prediction mode in VVC is coded separately for luma and chroma. Here, a Most Probable Mode list containing six modes is applied for luma. The individual compression performance of tools is reported in this paper. For the full VVC intra codec, a bitrate saving of 25% on average is reported over H.265/HEVC using an objective metric. Significant subjective benefits are illustrated with specific examples.
Jonathan Pfaff, Alexey Filippov, Shan Liu 0001, Xin Zhao 0003, Jianle Chen, Santiago De-Luxán-Hernández, Thomas Wiegand 0001, Vasily Rufitskiy, Adarsh K. Ramasubramonian, Geert Van der Auwera
IEEE Trans. Circuits Syst. Video Technol.5
2021 Subblock-Based Motion Derivation and Inter Prediction Refinement in the Versatile Video Coding Standard
abstract
Efficient representation and coding of fine-granular motion information is one of the key research areas for exploiting inter-frame correlation in video coding. Representative techniques towards this direction are affine motion compensation (AMC), decoder-side motion vector refinement (DMVR), and subblock-based temporal motion vector prediction (SbTMVP). Fine-granular motion information is derived at subblock level for all the three coding tools. In addition, the obtained inter prediction can be further refined by two optical flow-based coding tools, the bi-directional optical flow (BDOF) for bi-directional inter prediction and the prediction refinement with optical flow (PROF) exclusively used in combination with AMC. The aforementioned five coding tools have been extensively studied and finally adopted in the Versatile Video Coding (VVC) standard. This paper presents technical details of each tool and highlights the design elements with the consideration of typical hardware implementations. Following the common test conditions defined by Joint Video Experts Team (JVET) for the development of VVC, 5.7% bitrate reduction on average is achieved by the five tools. For test sequences characterized by large and complex motion, up to 13.4% bitrate reduction is observed. Additionally, visual quality improvement is demonstrated and analyzed.
Haitao Yang 0001, Huanbang Chen, Jianle Chen, Semih Esenlik, Sriram Sethuraman, Xiaoyu Xiu, Elena Alshina, Jiancong Luo
IEEE Trans. Circuits Syst. Video Technol.3
2020 Intra Prediction in the Emerging VVC Video Coding Standard
abstract
The focus of this work is on intra-prediction tools that distinguish Versatile Video Coding (VVC) [1] from its predecessors and do not add RD-checks on the encoder side, i.e. their coding efficiency is achieved due to their inherent properties but not by additional encoder-side complexity increase.
Alexey Filippov, Vasily Rufitskiy, Jianle Chen, Elena Alshina
DCC3
2020 Guest Editorial Introduction to the Special Section on the Joint Call for Proposals on Video Compression With Capability Beyond HEVC
abstract
Standardization for digital video compression has shown significant evolution over the last three decades. Starting in 1988 with ITU-T H.261 as the first such standard that was practical for consumer use, ISO/IEC MPEG-1 and H.262/MPEG-2 video (the latter jointly standardized by ITU-T and ISO/IEC) were developed very soon thereafter, creating the first wave of broad usage of digital technology in consumer video, such as broadcast and disc player applications. Later, the H.264/MPEG-4 Advanced Video Coding (AVC) standard was again developed jointly by ITU-T and ISO/IEC experts, with its High Profile becoming dominant from 2004 in HD broadcast and storage, as well as network-based streaming services and private capture of video. With ever-increasing demands for higher quality and the advent of flat-panel displays, the H.265/MPEG-H High Efficiency Video Coding (HEVC) standard became the next generation of video compression standard; with its first version defined in 2013, HEVC has been especially instrumental for the recent deployment of Ultra High Definition (UHD, a.k.a. 4K video). As time has moved forward, video content has continued to become an increasing presence in our lives, with an ever-growing diversification of usage models and continuing demands for higher quality. For example, flat panels evolved towards support of high dynamic range (HDR) video with a wider color gamut, and new modalities for consuming video have appeared, such as head-mounted displays (HMDs).
Jill M. Boyce, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Thomas Wiegand 0001, Yan Ye 0003
IEEE Trans. Circuits Syst. Video Technol.2
2020 The Joint Exploration Model (JEM) for Video Compression With Capability Beyond HEVC
abstract
This paper provides an overview of the coding algorithms of the Joint Exploration Model (JEM) for video compression with capability beyond HEVC, which was developed by the Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). The goal of the JEM development and experimentation was to provide evidence that sufficient coding efficiency improvement over the High Efficiency Video Coding (HEVC) standard can be achieved, which would justify the need for a new video coding standard with a compression capability significantly exceeding that of HEVC. The development of the JEM provided an ability to conduct studies toward that goal in a verifiable and collaborative manner and led to the launching of the project to develop the new Versatile Video Coding (VVC) standard. Objective metric gains exceeding 30% were measured for most of the tested high-resolution video content that represents current demanding new applications, and subjective testing using human observers showed even more benefit.
Jianle Chen, Marta Karczewicz, Yu-Wen Huang, Kiho Choi, Jens-Rainer Ohm, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.1
2020 Video Codec Using Flexible Block Partitioning and Advanced Prediction, Transform and Loop Filtering Technologies
abstract
This paper describes a joint response to the Call for Proposals by Samsung, Huawei, GoPro, and HiSilicon on Video Compression with Capability beyond HEVC/H.265, jointly issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG). In the proposed codec, the coding framework supports hierarchical splitting with binary and ternary trees and flexible coding order representations. Additionally, novel compression tools on inter/intra prediction, in-loop filtering, and entropy coding have been proposed. The proposed compression scheme provides significantly higher compression capability than the state-of-the-art HEVC/H.265 standard for SDR (Standard Dynamic Range) category while maintaining complexity acceptable for emerging applications. When all the proposed algorithmic tools are used, the proposed video codec achieves approximately 40% bit-saving for the SDR cetegory on average compared to HEVC/H.265 anchor.
Kiho Choi, Jianle Chen, Haitao Yang 0001, Woongil Choi, Sergey Ikonin, Yinji Piao, Semih Esenlik, Minsoo Park, Ye-Kui Wang, Narae Choi, Yin Zhao, Seungsoo Jeong, Anish Tamse, Alexey Filippov, Heechul Yang, Junghye Min, Roman Chernyak, Bora Jin, Anand Meher Kotra, Sunil Lee, Han Gao 0001, Chanyul Kim, Timofey Solovyev, Kwangpyo Choi, Vasily Rufitskiy, Maxim Sychev, Jeonghoon Park
IEEE Trans. Circuits Syst. Video Technol.2
2020 Cross-Component Prediction in HEVC
abstract
Video coding in the YCbCr color space has been widely used, since it is efficient for compression, but it can result in color distortion due to conversion error. Meanwhile, coding in the RGB color space maintains high color fidelity, having the drawback of a substantial bitrate increase with respect to YCbCr coding. Cross-component prediction (CCP) efficiently compresses video content by decorrelating color components while keeping high color fidelity. In this scheme, the chroma residual signal is predicted from the luma residual signal inside the coding loop. This paper gives a description of the CCP scheme from several points of view, from theoretical background to practical implementation. The proposed CCP scheme has been evaluated in standardization communities and adopted into H.265/High Efficiency Video Coding (HEVC) Range Extensions. The experimental results show significant coding performance improvements for both natural and screen content video, while the quality of all color components is maintained. The average coding gains for natural video are 17% and 5% bitrate reduction in the case of intra coding and 11% and 4% in the case of inter coding for RGB and YCbCr coding, respectively, while the average increment of encoding and decoding times in the HEVC reference software implementation are 10% and 4%, respectively.
Woo-Shik Kim, Ali Khairat, Mischa Siekmann, Joel Sole, Jianle Chen, Marta Karczewicz, Tung Nguyen 0001, Detlev Marpe
IEEE Trans. Circuits Syst. Video Technol.6
2020 Advanced 3D Motion Prediction for Video-Based Dynamic Point Cloud Compression
abstract
Point cloud based immersive media representation format has provided many opportunities for extended reality applications and has become widely used in volumetric content capturing scenarios. The high data rate of the point cloud is one of the key problems preventing the adoption of this media format. MPEG Immersive media working group (MPEG-I) aims to create a point cloud compression methodology relying on the existing video coding hardware implementations to solve this problem. However, in the scope of the state-of-the-art video-based dynamic point cloud compression (V-PCC) standard under MPEG-I, the intrinsic 3D object's motion continuity is destroyed by the 2D projections resulting in a significant loss of inter prediction coding efficiency. In this paper, we first propose a general model utilizing the 3D motion and 3D to 2D correspondence to calculate the 2D motion vector (MV). Then under the V-PCC, we propose a geometry-based method using the accurate 3D reconstructed geometry from the 2D geometry video to estimate the 2D MV in the 2D attribute video. In addition, we propose an auxiliary-information-based method using the coarse 3D reconstructed geometry provided by the auxiliary information to estimate the 2D MV in both the 2D geometry and attribute videos. Furthermore, we provide the following two ways to use the estimated 2D MV to improve the coding efficiency. The first one is normative. We propose adding the estimated MV into the advanced motion vector candidate list and find a better motion vector predictor for each prediction unit (PU). The second one is non-normative. We propose applying the estimated MV as an additional candidate of the centers for motion estimation. We implement the proposed algorithms in the V-PCC reference software. The experimental results show that the proposed methods present significant coding gains compared with the current state-of-the-art motion prediction algorithm.
Li Li 0040, Zhu Li 0001, Vladyslav Zakharchenko, Jianle Chen, Houqiang Li
IEEE Trans. Image Process.4
2019 Advanced 3D Motion Prediction for Video Based Point Cloud Attributes Compression
abstract
Point cloud media representation format has provided various opportunities for extended reality applications and had become widely used in volumetric content capturing scenarios. At the same time ambiguous storage format representations and network throughput are key problems for wide adoption of this media format. Compression algorithms in corresponding standard activities are aimed to solve this problem. MPEG-I standard has an aim of creating the point cloud compression methodology relying on existing video coding hardware implementations. In scope of the state-of-the-art video-based dynamic point cloud (DPC) compression method, similar 3D patches may be projected in totally different 2D positions in different frames. In this way, the motion vector predictors especially those in the patch boundary may be very inaccurate which may lead to significant bitrate increase. In this paper, we propose to use the reconstructed geometry information to help predict the motion vector more accurately and improve the coding efficiency of the attribute video. First, we propose to use the motion vector of the co-located blocks in the geometry frame as a merge candidate of the current block in the attribute frame. Second, we perform a motion estimation between the current reconstructed point cloud with only the geometry information and the reference point cloud to find the corresponding block. The motion information derived is used as motion vector predictor of the current block in the attribute frame. As far as we can see, this is the first work using the geometry information to compress the attribute in the DPC compression scenario. Significant compression efficiency is achieved with this new 3D point cloud geometry derived motion prediction scheme when compared with the state-of-the-art DPC compression method.
Li Li 0040, Zhu Li 0001, Vladyslav Zakharchenko, Jianle Chen
DCC4
2019 New Video Codec for High-Quality Video Service and Emerging Applications
abstract
This paper proposes a novel video compression scheme for high-quality video service and emerging applications such as 360-degree omnidirectional and high dynamic range video coding. The coding framework supports hierarchical splitting of blocks with binary and ternary-split trees and flexible coding order representations. Moreover, minimal tool set to obtain high precision prediction and compression enhancement has been proposed. Compared to HEVC, bit-rate reduction of around 40% based on objective measures has been shown. This was one of the responses to the Call for Proposals (CfP) for VVC standardization.
Kiho Choi, Jianle Chen, Anish Tamse, Haitao Yang 0001, Sergey Ikonin, Woongil Choi, Semih Esenlik
DCC2
2019 Level-of-Detail Generation Using Binary-Tree for Lifting Scheme in LiDAR Point Cloud Attributes Coding
abstract
Point clouds are one of the emerging 3D visual representations of real word and plenty of useful applications has already been demonstrated. However, a huge amount of data associated with it has added challenges in both transmission and storage. This requires an efficient coding solution and brought a great attention among compression community. MPEG and JPEG standardization group has already started developing coding solution and proposed two test-models namely V-PCC, video-based coding solution, for dynamic point cloud and G-PCC, a native geometry-based coding solution, for static and LiDAR point cloud. In G-PCC, octree (lossless) and tri-soup(lossy) for geometry coding, similarly regional adaptive hierarchical transform (RAHT) and lifting-scheme for attributes coding are currently being explored. Lifting-scheme relies on level-of-details(LOD) structure for attributes prediction where LOD is generated with distance based subsampling approach. In this work we proposed a new LOD generation scheme using binary-tree and showed it provides better coding solution for sparse point cloud such as LiDAR. The experimental results demonstrated 12% bitrate reduction for reflectance and 8%, 6% and 7% bitrate reduction for luma, chroma Cb and chroma Cr respectively as well as up to 4 times computational complexity reduction compared to current G-PCC lifting-scheme.
Birendra Kathariya, Vladyslav Zakharchenko, Zhu Li 0001, Jianle Chen
DCC4
2019 Decoder Side Motion Vector Refinement for Versatile Video Coding
abstract
Inter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the bitrate required for motion vector (MV) signaling, the High Efficiency Video Coding (HEVC) standard and the latest Versatile Video Coding (VVC) draft utilize a merge mode to signal the MV. While the merge mode saves the bits for MV indication, it generates inaccurate MVs, which lead to imprecise prediction. To improve the coding performance, novel decoder side motion vector refinement (DMVR) schemes are currently being proposed. The DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. We present two variants of block matching-based DMVR, namely template matching and bilateral matching. Experimental results obtained with the VTM-2.0 reference software, after integrating our approaches, demonstrate that our proposed methods provide an average luma BD-rate reduction of 4.71% for the template matching-based DMVR and 4.92% for the bilateral matching-based DMVR when using the random access configuration.
Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
MMSP5
2019 Low Complexity Decoder Side Motion Vector Refinement for VVC
abstract
Inter picture prediction is an essential component of today's hybrid video codecs. In order to reduce the motion vector signaling overhead, a merge mode with subsequent decoder side motion vector refinement (DMVR) is current under investigation for the first working draft of the standardization activity on Versatile Video Coding (VVC). While the DMVR method searches the refined MVs at the decoder side, it heavily increases the decoding complexity and the memory bandwidth requirements. To address these issues, a novel low complexity DMVR scheme is proposed in this paper. The low complexity DMVR approach refines the initial MV from the merge mode by searching the block with the smallest matching cost in the previous decoded reference pictures. The proposed low complexity improvements are added to a previously proposed bilateral matching-based DMVR approach. Experimental results obtained with the VTM 2.0 reference software, after integrating our approaches, show that the previously proposed DMVR provides an average luma BD-rate reduction of 4.59% with 32% additional decoding time and the proposed low complexity DMVR provides an average luma BD-rate reduction of 1.67% with only 6% additional decoding time when using the random access configuration.
Han Gao 0001, Semih Esenlik, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
PCS5
2018 Scalable Point Cloud Geometry Coding with Binary Tree Embedded Quadtree
abstract
Many applications of point cloud have recently been identified in automobile navigation system, visual communication, and so on. However, the huge data size of point cloud has been a bottleneck for the practical implementations. In this paper, we present a compression scheme that utilizes variable-rate coding of a same point cloud data at different quality. Point cloud is encoded at fixed-rate for highest representation. Encoder, however, can present variable-rate encoded data for any lowest to highest representation to decoder which is then decoded to reconstruct point cloud at different quality. Variable-rate encoding is achieved through the so-called binary tree quadtree (BTQT) scheme. The BTQT scheme made the compression more effective by dividing point cloud frame into blocks using binary-tree and encoding flat surfaces in the blocks by quadtree and non-flat surfaces by octree. Simulation results show that scalable coding solution can efficiently compress point cloud data at variable rate compensating the quality.
Birendra Kathariya, Li Li 0040, Zhu Li 0001, Jose R. Alvarez, Jianle Chen
ICME5
2018 Improving Picture Boundary Handling for Video Coding Beyond HEVC
abstract
Block partitioning is an essential component of today's hybrid video codecs. In order to support video sequences with arbitrary spatial resolution, the HEVC standard and the JEM reference software utilize a forced quadtree (QT) method to process CTUs/CUs located at the picture boundaries. While the forced QT method is simple to implement, it is unsatisfactory from a coding efficiency perspective. To improve the compression performance, novel block partitioning methods are currently being proposed. Our approach for picture boundary handling uses a binary tree (BT) based splitting scheme to deal with the boundary located CTUs/CUs. We present two variants with the first one combining forced BT with forced QT partitioning (forced QTBT) and the second one adaptively choosing between BT and QT (adaptive QTBT). Experimental results obtained with the JEM-6.0 reference software after integrating our approaches demonstrate that our proposed methods provide an average luma BD-rate reduction of 0.85% for the forced QTBT scheme and 0.96% for the adaptive QTBT scheme when using the random access configuration.
Han Gao 0001, Zhijie Zhao, Eckehard G. Steinbach, Jianle Chen
VCIP4
2018 Enhanced Cross-Component Linear Model for Chroma Intra-Prediction in Video Coding
abstract
Cross-Component Linear Model (CCLM) for chroma intra-prediction is a promising coding tool in Joint Exploration Model (JEM) developed by the Joint Video Exploration Team (JVET). CCLM assumes a linear correlation between the luma and chroma components in a coding block. With this assumption, the chroma components can be predicted by the Linear Model (LM) mode, which utilizes the reconstructed neighbouring samples to derive parameters of a linear model by linear regression. This paper presents three new methods to further improve the coding efficiency of CCLM. First, we introduce a multi-model CCLM (MM-CCLM) approach, which applies more than one linear models to a coding block. With MM-CCLM, reconstructed neighbouring luma and chroma samples of the current block are classified into several groups, and a particular set of linear model parameters is derived for each group. The reconstructed luma samples of the current block are also classified to predict the associated chroma samples with the corresponding linear model. Second, we propose a multi-filter CCLM (MF-CCLM) technique, which allows the encoder to select the optimal down-sampling filter for the luma component with the 4:2:0 colour format. Third, we present a LM-angular prediction (LAP) method, which synthesizes the angular intra-prediction and the MM-CCLM intra-prediction into a new chroma intra coding mode. Simulation results show that 0.55%, 4.66% and 5.08% BD rate savings in average on Y, Cb and Cr components respectively, are achieved for All Intra (AI) configurations with the proposed three methods. MM-CCLM and MF-CCLM have been adopted into the JEM by JVET.
Kai Zhang 0007, Jianle Chen, Li Zhang 0006, Xiang Li 0003, Marta Karczewicz
IEEE Trans. Image Process.2
2018 Joint Separable and Non-Separable Transforms for Next-Generation Video Coding
abstract
Throughout the past few decades, the separable Discrete Cosine Transform (DCT), particularly the DCT type II, has been widely used in image and video compression. It is well known that, under first-order stationary Markov conditions, DCT is an efficient approximation of the optimal Karhunen-Loève transform. However, for natural image and video sources, the adaptivity of a single separable transform with fixed core is rather limited for the highly dynamic image statistics, e.g., textures and arbitrarily directed edges. It is also known that non-separable transforms can achieve better compression efficiency for images with directional texture patterns, yet they are computationally complex, especially when the transform size is large. In order to achieve higher transform coding gains with relatively low-complexity implementations, we propose a joint separable and non-separable transform. The proposed separable primary transform, named Enhanced Multiple Transform (EMT), applies multiple transform cores from a pre-defined subset of sinusoidal transforms, and the transform selection is signaled in a joint block level manner. Moreover, a Non-Separable Secondary Transform (NSST) method is proposed to operate in conjunction with EMT. Unlike the existing non-separable transform schemes which require excessive amounts of memory and computation, the proposed NSST efficiently improves coding gain with much lower complexity. Extensive experimental results show that the proposed methods, in a state-of-the-art video codec, such as HEVC, can provide significant coding gains (average 6.9% and 4.5% bitrate reductions for intra and random-access coding, respectively).
Xin Zhao 0003, Jianle Chen, Marta Karczewicz, Amir Said, Vadim Seregin
IEEE Trans. Image Process.2
2017 Frame Rate Up-Conversion Based Motion Vector Derivation for Hybrid Video Coding
abstract
In this paper, a MV derivation method based on the idea of frame rate up-conversion (FRUC) is proposed. When a block is signaled as FRUC mode, the motion information of the block is derived without signaling. Moreover, derived MVs are refined at sub-block level for more accurate motion field. In addition, two matching methods, i.e., bilateral matching and template matching are supported to obtain good performance in both bi-directional and uni-directional prediction. Simulations under HEVC common test conditions show that over 4.2% average BD-rate reduction was achieved over HEVC reference software HM-16.6 in the case of random access configuration. The method has been adopted into the Joint Exploration Model (JEM) developed by the joint video exploration team (JVET) of MPEG and ITU-T VCEG for the study of next generation video coding standard.
Xiang Li 0003, Jianle Chen, Marta Karczewicz
DCC2
2017 Multiple direct mode for intra coding
abstract
In this paper, a multiple direct mode (MDM) method is presented for chroma intra coding. The main contributions of the proposed MDM method include two aspects: selection of multiple luma intra prediction modes from co-located luma blocks, and the derivation of chroma intra prediction modes from spatial neighbouring blocks. With the proposed method, both the cross-component correlation and spatial correlation of intra prediction modes can be better utilized for more efficient chroma intra coding. Simulation results have validated the efficiency of MDM especially under the decoupled luma-chroma partition trees. The proposed method has been adopted in the Joint Exploration Model (JEM) which is the test platform for future video coding technology exploration in Joint Video Exploration Team (JVET).
Li Zhang 0006, Wei-Jung Chien, Jianle Chen, Xin Zhao 0003, Marta Karczewicz
VCIP3
2017 Multi-model based cross-component linear model chroma intra-prediction for video coding
abstract
Cross-component Linear Model (CCLM) chroma intra prediction assumes a linear correlation between the luma and chroma components in a coding block. With this assumption, the chroma components can be predicted by LM mode, which utilizes the reconstructed neighbouring samples to derive parameters of the linear model by linear regression. This paper presents a multi-model CCLM (MM-CCLM) approach, which applies more than one linear models in a coding block. With MM-CCLM, reconstructed neighbouring luma and chroma samples of the current block are classified into several groups and each group is used as a training set to derive its own linear model. The reconstructed luma samples of the current block are also classified to use corresponding linear model to predict the associated chroma samples. Simulation results show that 0.26%, 1.89% and 1.96% BD rate savings on Y, Cb and Cr components are achieved for All Intra (AI) configurations in average. The proposed method has been adopted in the Joint Exploration Model (JEM) by Joint Video Exploration Team (JVET).
Kai Zhang 0007, Jianle Chen, Li Zhang 0006, Xiang Li 0003, Marta Karczewicz
VCIP2
2016 Enhanced Multiple Transform for Video Coding
abstract
The Discrete Cosine Transform (DCT), and in particular the DCT type II, has been widely used for image and video compression. Although DCT efficiently approximates the optimal Karhunen–Loève transform under first-order Markov conditions with low complexity, the energy packing efficiency is still limited since a fixed transform cannot always capture the highly dynamic statistics of natural video content. In this paper, to further improve the transform efficiency, an Enhanced Multiple Transform (EMT) scheme is proposed. In the proposed EMT, a few sinusoidal transforms, other than DCT, have also been utilized for coding both Intra and Inter prediction residuals. The best transform, as selected from a pre-defined transform subset specified by prediction mode, is explicitly signaled in a joint coding block level manner. Moreover, to accelerate encoding process, fast methods have also been proposed by skipping unnecessary transform rate-distortion evaluations using previously encoding statistics. The proposed method has been implemented on top of High-Efficiency Video Coding (HEVC) reference software, and significant coding gain has been verified.
Xin Zhao 0003, Jianle Chen, Marta Karczewicz, Li Zhang 0006, Xiang Li 0003, Wei-Jung Chien
DCC2
2016 Position dependent prediction combination for intra-frame video coding
abstract
Intra-frame prediction in the High Efficiency Video Coding (HEVC) standard can be empirically improved by applying sets of recursive two-dimensional filters to the predicted values. However, this approach does not allow (or complicates significantly) the parallel computation of pixel predictions. In this work we analyze why the recursive filters are effective, and use the results to derive sets of non-recursive predictors that have superior performance. We present an extension to HEVC intra prediction that combines values predicted using non-filtered and filtered (smoothed) reference samples, depending on the prediction mode, and block size. Simulations using the HEVC common test conditions show that a 2.0% bit rate average reduction can be achieved compared to HEVC, for All Intra (AI) configurations.
Amir Said, Xin Zhao 0003, Marta Karczewicz, Jianle Chen
ICIP4
2016 Highly efficient non-separable transforms for next generation video coding
abstract
For the last few decades, the application of signal-adaptive transform coding to video compression has been stymied by the large computational complexity of matrix-based solutions. In this paper, we propose a novel parametric approach to greatly reduce the complexity without degrading the compression performance. In our approach, instead of following the conventional technique of identifying full transform matrices that yield best compression efficiency, we look for the best transform parameters defining a new class of transforms, called HyGTs, which have low complexity implementations that are easy to parallelize. The proposed HyGTs are implemented as an extension of High Efficiency Video Coding (HEVC), and our comprehensive experimental results demonstrate that proposed HyGTs improve average coding gain by 6% bit rate reduction, while using 6.8 times less memory than KLT matrices.
Amir Said, Xin Zhao 0003, Marta Karczewicz, Hilmi E. Egilmez, Vadim Seregin, Jianle Chen
PCS6
2016 NSST: Non-separable secondary transforms for next generation video coding
abstract
In traditional image and video coding schemes, separable transforms are typically employed due to their low-complexity implementations. However, the compression efficiency of separable transforms is limited for most natural image/video blocks which generally have arbitrarily directed edge and texture patterns. It is well known that non-separable transforms can achieve better compression efficiency for directional texture patterns, yet they are computationally complex, especially for larger block sizes. In order to achieve higher transform coding gains with relatively low-complexity implementations, in this paper, we propose non-separable secondary transforms (NSSTs). The proposed approach applies a secondary non-separable transform on a sub-block of low frequency coefficients generated using a primary separable transform, such as discrete cosine transform (DCT). Since the proposed NSST is a non-separable transform applied on low frequency coefficients in a much smaller block size, which typically captures most of the signal energy, better coding gains can be achieved with at a relatively low-computational cost. Experimental results show that, compared to the latest HEVC reference software (HM16.6), the proposed method achieves up to a significant 12% coding gain for Intra coding.
Xin Zhao 0003, Jianle Chen, Amir Said, Vadim Seregin, Hilmi E. Egilmez, Marta Karczewicz
PCS2
2016 Overview of SHVC: Scalable Extensions of the High Efficiency Video Coding Standard
abstract
This paper provides an overview of Scalable High efficiency Video Coding (SHVC), the scalable extensions of the High Efficiency Video Coding (HEVC) standard, published in the second version of HEVC. In addition to the temporal scalability already provided by the first version of HEVC, SHVC further provides spatial, signal-to-noise ratio, bit depth, and color gamut scalability functionalities, as well as combinations of any of these. The SHVC architecture design enables SHVC implementations to be built using multiple repurposed single-layer HEVC codec cores, with the addition of interlayer reference picture processing modules. The general multilayer high-level syntax design common to all multilayer HEVC extensions, including SHVC, MV-HEVC, and 3D HEVC, is described. The interlayer reference picture processing modules, including texture and motion resampling and color mapping, are also described. Performance comparisons are provided for SHVC versus simulcast HEVC and versus the scalable video coding extension to H.264/advanced video coding.
Jill M. Boyce, Yan Ye 0003, Jianle Chen, Adarsh K. Ramasubramonian
IEEE Trans. Circuits Syst. Video Technol.3
2015 Resampling Process of the Scalable High Efficiency Video Coding
abstract
SHVC is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC) and spatial resampling process is inevitable module to support spatial scalability. This paper describes in details the resampling process, including both texture and motion data resampling in SHVC, and using experimental evidence, demonstrate their benefits in terms of coding efficiency.
Jianle Chen, Elena Alshina, Xiang Li 0003, Marta Karczewicz, Alexander Alshin
DCC1
2015 Asymmetric 3D Lookup Table Based Color Gamut Scalability in SHVC
abstract
SHVC is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC). Color Gamut Scalability (CGS) refers to a scalable use case in which base layer and enhancement layer have different color gamuts. In this case, special inter-layer prediction is needed to improve coding efficiency in SHVC. In this paper, a solution based on asymmetric 3D lookup table is presented for color gamut scalability. Compared to SHVC without CGS coding tool, the proposed solution provides 9.6% - 16.1% overall luma BD-rate reduction in different test cases.
Xiang Li 0003, Jianle Chen, Marta Karczewicz, Yuwen He, Yan Ye 0003, Cheung Auyeung
DCC2
2015 Adaptive Color-Space Transform for HEVC Screen Content Coding
abstract
This paper presents an in-loop adaptive color-space transform for the HEVC Screen Content Coding extension. In the proposed method, the prediction residual is adaptively converted into a different color space to reduce the cross-component redundancy. After the ACT, the signal is coded following the existing HEVC framework. To keep the complexity as low as possible, fixed color-space transforms that are easily implemented with shift and add operations are utilized. Significant coding gains are achieved by this method in the current HEVC Screen Content Coding reference software with no increase of decoding runtime. The proposed method has been adopted to the HEVC Screen Content Coding extension.
Li Zhang 0006, Jianle Chen, Joel Sole, Marta Karczewicz, Xiaoyu Xiu, Ji-Zheng Xu
DCC2
2014 Region based inter-layer cross-color filtering for scalable extension of HEVC
abstract
Inter-layer filtering is a key module of the emerging Scalable Extension of High Efficiency Video Coding Standard (SHVC). In SHVC, up-sampled based layer reconstructed pictures are used as inter-layer references to predict enhancement layer frames such that inter-layer redundancy is reduced. To improve the coding performance of inter-layer filtering, luma plane based chroma plane enhancement was proposed at picture level. However, the efficiency of the picture level adaptation is not very promising when picture resolution is high. To address this issue, region based inter-layer cross-color filtering is proposed in this paper. Simulations under the common test conditions defined by Joint Collaborative Team on Video Coding (JCT-VC) showed that significant chroma coding gain and moderate luma improvement were achieved by the proposed method. When compared to the luma plane based chroma plane enhancement method, the coding gain over SHVC reference software SHM-2.0 is about doubled while the decoding complexity is kept even lower. Moreover, the proposed method outperforms other tools studied in SHVC core experiment on inter-layer filtering.
Xiang Li 0003, Jianle Chen, Marta Karczewicz, Elena Alshina, Alexander Alshin, Yongjin Cho
ICIP3
2014 Cross component decorrelation for HEVC range extension standard
abstract
This paper presents a new coding tool named cross component decorrelation in the emerging High Efficiency Video Coding Range Extension (HEVC RExt) standard. Color video is generally composed of three color components, e.g., RGB or YCbCr. It has been known for over a decade that the three color components have correlation among each other. Although global out-of-loop color space conversion, e.g., RGB-to-YCbCr, can reduce the cross component correlation, local correlation still exists in YCbCr signal. Many methods have been developed in the literature to exploit such redundancy to improve coding efficiency. However, existing methods introduce high computational or implementation complexity, which makes them never be included in mainstream video coding standard such as H.264/AVC and HEVC version 1. In this research, a new hardware friendly cross component decorrelation method is presented which reduces implementation cost while achieving significant BD-rate reduction. For example, based on JCT-VC common test condition for HEVC RExt standardization, the proposed method results in (17.3%, 18.1%, 16.6%) BD-rate reduction for the three color components of the RGB test sequences, in the case of All Intra configuration.
Woo-Shik Kim, Jianle Chen, Joel Sole, Marta Karczewicz
ICIP3
2013 Scalable Video Coding Extension for HEVC
abstract
This paper describes a scalable video codec that was submitted as a response to the joint call for proposals issued by ISO/IEC MPEG and ITU-T VCEG on HEVC scalable extension. The proposed codec uses a multi-loop decoding structure. Several inter-layer texture prediction methods are employed to remove the inter-layer redundancy. Inter-layer prediction is also used when coding enhancement layer syntax elements such as motion parameter and intra prediction mode, to further reduce bit overhead. Additionally, alternative transforms as well as adaptive coefficients scanning are used to code the prediction residues more efficiently. Experimental results are presented to demonstrate the effectiveness of the proposed scheme. When compared to HEVC single-layer coding, the additional rate overhead for the proposed scalable extension is 1.2% to 6.4% to achieve two layers of SNR and spatial scalability.
Jianle Chen, Krishnakanth Rapaka, Xiang Li 0003, Vadim Seregin, Marta Karczewicz, Geert Van der Auwera, Joel Sole, Xianglin Wang, Chengjie Tu, Ying Chen 0011, Rajan L. Joshi
DCC1
2013 Generalized inter-layer residual prediction for scalable extension of HEVC
abstract
Scalable video coding extension of HEVC (SHVC) is being developed by Joint Collaborative Team on Video Coding (JCT-VC) of ISO/IEC MPEG and ITU-T VCEG. Different from scalable video coding extension of H.264/AVC (SVC), SHVC employs a multi-loop decoding framework so that the inter-layer residual prediction in SVC does not perform well in SHVC. In this paper, a method called generalized inter-layer residual prediction (GILRP) is proposed. To improve prediction accuracy, the residual predictor is derived with the information from both base and enhancement layers. Moreover, three additional weighting types are introduced on top of inter coding modes to further compensate errors caused by base layer quantization. Simulations under SHVC common test conditions defined by JCT-VC show that 2.9%, 5.1% and 4.9% overall luma BD-rate reduction on average were obtained over SHVC reference software for configurations of random access, low delay with P slices, and low delay with B slices, respectively.
Xiang Li 0003, Jianle Chen, Krishnakanth Rapaka, Marta Karczewicz
ICIP2
2013 Inter-layer filtering for scalable extension of HEVC
abstract
This paper introduces inter-layer filters for the scalable extension of High Efficiency Video Coding (SHVC) standard, which is being developed by the Joint Collaborative Team on Video Coding (JCT-VC). The major new coding tool in SHVC is inter-layer texture prediction. It provides about 18% average BD-rate reduction compared with HEVC two-layer simulcast. In the case of spatial scalability, base layer reconstructed pictures are up-sampled to the enhancement layer resolution to generate inter-layer texture prediction. A set of 2D separable 8 taps (luma) and 4 taps (chroma) DCT based interpolation filters, which follow the design principles of HEVC motion compensation interpolation filter, are used in the up-sampling process. In the case of SNR scalability, the up-sampling process is not needed since the reference layer has the same spatial resolution as the current layer but encoded with lower quality. This paper proposes a novel inter-layer filter with denoising effect for SNR scalability to improve enhancement layer coding efficiency and equalize the number of stages in inter-layer processing between SNR and spatial scalabilities. Experimental results show that the usage of inter-layer de-noising filter in SNR scalability provides up to 7.5% BD-rate reduction and has observable improvement on subjective visual quality.
Elena Alshina, Alexander Alshin, Yongjin Cho, Jeong-Hoon Park, Jianle Chen, Xiang Li 0003, Vadim Seregin, Marta Karczewicz
PCS6
2013 High Frequency SAO for scalable extension of HEVC
abstract
Scalable extension of HEVC, a.k.a. SHVC, is being standardized by the Joint Collaborative Team on Video Coding (JCT-VC). SHVC employs one of the most important coding tools called interlayer prediction, in which. Reconstructed base layer pictures can be used as reference pictures to predict enhancement layer pictures. Therefore, how to efficiently generate interlayer reference pictures to improve coding efficiency is one of the core research topics for the new international standard. In this paper, we presented High Frequency SAO filter (HF-SAO), which extends Sample Adaptive Offset filter (SAO) in HEVC to SHVC. Experimental results based on SHVC reference software version 1.0 show that HFSAO achieves 1.2% (Luma), 1.4% (Cb), 1.4% (Cr) average BD-rate reduction for the enhancement layer coding, which makes itself one of the most promising candidate interlayer filters to the new generation of scalable video coding standard.
Jianle Chen, Krishnakanth Rapaka, Xiang Li 0003, Marta Karczewicz
PCS2
2013 Efficient key picture and single loop decoding scheme for SHVC
abstract
Scalable video coding has been a popular research topic for many years. As one of its key objectives, it aims to support different receiving devices connected through a network structure using a single bitstream. Scalable video coding extension of HEVC, also called as SHVC, is being developed by Joint Collaborative Team on Video Coding (JCT-VC) of ISO/IEC MPEG and ITU-T VCEG. Compared to previous standardized scalable video coding technologies, SHVC employs multi-loop decoding design with no low-level changes within any given layer compared to HEVC. With such a simplified extension it aims at solving some of the problems of previous scalable extensions that haven't been successful, and at the same time, aims at supporting all design features that are of vital importance for the success of SHVC. Supporting lightweight and finely tunable bandwidth adaptation is one such vital design feature important for the success of SHVC. This paper proposes novel high level syntax mechanism for SHVC quality scalability to support: (a) using the decoded pictures from higher quality layer as reference for lower layer pictures and key pictures concept to reduce drift; (b) single loop decoding design with encoder only constraints without introducing any normative low-level changes to the normal multi-loop decoding process. Experimental results based on SHVC reference software (SHM 2.0) show that the proposed key picture method achieves an average of 2.9% luma BD-rate reduction in multi-loop framework and an average of 4.4% luma BD-rate loss to attain the capability of single loop decoding.
Krishnakanth Rapaka, Jianle Chen, Marta Karczewicz
VCIP2
2010 Improved Video Compression Efficiency Through Flexible Unit Representation and Corresponding Extension of Coding Tools
abstract
This paper proposes a novel video compression scheme based on a highly flexible hierarchy of unit representation which includes three block concepts: coding unit (CU), prediction unit (PU), and transform unit (TU). This separation of the block structure into three different concepts allows each to be optimized according to its role; the CU is a macroblock-like unit which supports region splitting in a manner similar to a conventional quadtree, the PU supports nonsquare motion partition shapes for motion compensation, while the TU allows the transform size to be defined independently from the PU. Several other coding tools are extended to arbitrary unit size to maintain consistency with the proposed design, e.g., transform size is extended up to 64 × 64 and intraprediction is designed to support an arbitrary number of angles for variable block sizes. Other novel techniques such as a new noncascading interpolation Alter design allowing arbitrary motion accuracy and a leaky prediction technique using both open-loop and closed-loop predictors are also introduced. The video codec described in this paper was a candidate in the competitive phase of the high-efficiency video coding (HEVC) standardization work. Compared to H.264/AVC, it demonstrated bit rate reductions of around 40% based on objective measures and around 60% based on subjective testing with 1080 p sequences. It has been partially adopted into the first standardization model of the collaborative phase of the HEVC effort.
Woojin Han 0001, Junghye Min, Il-Koo Kim, Elena Alshina, Alexander Alshin, Tammy Lee, Jianle Chen, Vadim Seregin, Sunil Lee, Yoon Mi Hong, Min-Su Cheon, Nikolay Shlyakhov, Ken McCann, Thomas Davies 0002, Jeong-Hoon Park
IEEE Trans. Circuits Syst. Video Technol.7
2009 Adaptive linear prediction for block-based lossy image coding
abstract
Linear prediction model has been well investigated and applied in lossless image and video coding. In this paper, we investigate the linear prediction method for block-based lossy image coding and propose a method that merges linear prediction technique into H.264/AVC video coding framework. A block-based linear prediction method is designed instead of pixel-based one in order to cooperate with transform module. Furthermore, line-based linear prediction with 1D transform is developed by considering coding gain tradeoff between prediction and transform. Linear prediction model coefficients are derived by using neighboring reconstructed data with least square error method. The model coefficients implicitly embed the local texture characteristics and no bits overhead is needed for signaling the coefficients since we can derive them with same process at decoder side. We insert block-based and line-based linear prediction modes into H.264/AVC as additional intra prediction modes and select the best mode by minimum rate-distortion sense. Experimental results show that the proposed technique improves coding efficiency of H.264/AVC intra picture with average 4.3% bit saving and up to 7.0% bit saving.
Jianle Chen, Woojin Han 0001
ICIP1
2005 Modified edge-oriented spatial interpolation for consecutive blocks error concealment
abstract
In this paper, a low-complexity spatial-domain error concealment algorithm is proposed for recovering consecutive blocks error in still images or intra-coded (I) frames of video sequences. The proposed algorithm works with the following steps. Firstly the Sobel operator is performed on the top and bottom adjacent pixels to detect the most likely edge direction of current block area. After that one-dimensional (1D) matching is used on the available block boundaries. Displacement between edge direction candidate and most likely edge direction is taken into consideration as an important factor to improve stability of ID boundary matching. Then the corrupted pixels are recovered by linear weighting interpolation along the estimated edge direction. Finally the interpolated values are merged to get last recovered picture. Simulation results demonstrate that the proposed algorithms obtain good subjective quality and higher PSNR than the methods in literatures for most images.
Jianle Chen, Jilin Liu, Xingguo Wang, Guobin Chen
ICIP (3)1