EDBT 2026 Demo / reviewers in the wild / expert
Shanshe Wang
dblp:126/4540
· DBLP profile ↗
24ranked-venue papers in the field
0as first author
17since 2021 · last 2025
0000-0002-7665-7434ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 23Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Comprehensively Introduction and Analysis of AVS3 Entropy CodingabstractThis paper elaborates the binary bin distribution of each syntax element in AVS3. The overall syntax elements in AVS3 can be divided into Context-Modeled(CM) Bin, Bypass(BPS) Bin, Weighted Context-Modeled(WCM) Bin. A comprehensive and quantitative analysis on these bitstreams are conducted and the result is shown in Fig. 1. Moreover, Fig. 2 illustrates the distribution of CM Bin, BPS Bin, and WCM Bin across various resolutions and QPs. The analysis reveals insightful correlations between the types of binary symbols and syntax elements within the AVS3 bitstream. Through this analysis, we aim to contribute to the implementation of AVS3 codecs. Jianchao Wei, Shanshe Wang, Jiaqi Zhang 0007 |
DCC | 3 |
| 2025 | Image Coding for Machine with Visual-Language Mimic Feature LearningabstractThis paper propose a Image Coding for Machine (ICM) framework with Visual-Language Mimic Feature Learning (VLM-ICM). VLM-ICM decouples the position and semantic information into language modality and extracts universal features from the input image. Language, inherently more semantically compact, helps reduce the bitrate. Meanwhile, the universal features in VLM-ICM, guided by the language at the decoder side, allow for flexible domain adaptation, thereby enhancing versatility and practicality. Zhimeng Huang, Junlong Gao, Jiaqi Zhang 0007, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia |
DCC | 4 |
| 2025 | A Hardware-Friendly AVS3 Entropy Decoder Architecture for 8K Ultra-High-Definition VideoabstractLogarithmic binary arithmetic coding (LBAC) is a entropy coding method first introduced in AVS2 and now used in the third generation of audio video coding standard (AVS3). While LBAC provides high coding efficiency, the data dependencies in AVS3 make it challenging to parallelize parsing syntax element. To enhance the throughput of AVS3 LBAC entropy decoding engine, a sophisticated hardware implementation architecture is meticulously designed in the paper. The proposed method decouples the bitstream decoding process into four distinct modules. To ensure the decoding efficiency and performance, an advanced parallel pipeline is carefully designed. Furthermore, we develop sub-branch state prediction mechanism and grouped decoding method to further improve the decoding efficiency. The proposed AVS3 entropy decoder has been implemented on the S10 FPGA System Board, and experimental results show that the proposed method could achieve real-time and seamless decoding efficiency for 8K AVS3 bitstreams, even at bitrates as high as 207.1Mbps. Jiaqi Zhang 0007, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001 |
DCC | 4 |
| 2025 | MoRLACS: A Monocular RGBD-based Locomotion Approach for CAVE SystemsabstractNavigation within Cave Automatic Virtual Environment (CAVE) systems often faces challenges due to limited physical space and the necessity for seamless user interaction. Traditional solutions typically rely on multi-view tracking systems or constrained locomotion techniques, which can interrupt immersion and hinder usability. In this paper, we introduce MoRLACS, a novel locomotion approach for CAVE systems that leverages a single RGBD camera. This hybrid framework integrates small-scale physical walking with controller-based large-scale exploration through a tailored guidance method. By accurately tracking the user's head position in the real world and synchronizing it with the virtual camera, MoRLACS enables natural walking within confined CAVE spaces and supports extended interaction in larger virtual environments. Preliminary user experiments demonstrate the approach's effectiveness, revealing improvements in usability and a heightened sense of presence. These findings underscore the potential of MoRLACS to enrich user experiences in immersive CAVE settings and offer valuable design insights for integrating 3D sensor data into multimedia interaction frameworks. Haopeng Lu, Qian Yin 0002, Li Song 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
ICMR | 6 |
| 2024 | A Fast Four-Parameter Affine Motion Compensation Algorithm for Video CodingabstractThis paper proposes a fast four-parameter Affine Motion Compensation (AMC) algorithm. As shown in Fig. 1, the translation Motion Vector (MV) is derived by reusing the AMC sub-block MV derivation method firstly, which is used to conduct translation pre-transform. Secondly, a coordinate system whose coordinate origin is located on its top-left control point is established for the transformed block. Finally, the geometric relationship between two control point motion vectors (CPMVs) of the transformed block can be described as follows,\begin{equation*}\delta = \left| {\left({m{v_{0x}} - m{v_{1x}}}\right) \times H - \left({m{v_{0y}} + m{v_{1y}}}\right) \times W} \right| = 0\tag{1}\end{equation*} Jiaqi Zhang 0007, Ivan V. Bajic, Shanshe Wang, Songlin Sun |
DCC | 4 |
| 2023 | Rate-Distortion Optimization for Cross Modal CompressionabstractRecently, cross modal compression (CMC) is proposed to compress highly redundant visual data into a compact, common, human-comprehensible domain (such as text) to preserve semantic fidelity for semantic-related applications. However, CMC only achieves a certain level of semantic fidelity at a constant rate, and the model aims to optimize the probability of the ground truth text but not directly semantic fidelity. To tackle the problems, we propose a novel scheme named rate-distortion optimized CMC (RDO-CMC). Specifically, we model the text generation process as a Markov decision process and propose rate-distortion reward which is used in reinforcement learning to optimize text generation. In rate-distortion reward, the distortion measures both the semantic fidelity and naturalness of the encoded text. The rate for the text is estimated by the sum of the amount of information of all the tokens in the text since the amount of information of each token is a lower bound of coding bits. Experimentally, RDO-CMC effectively controls the rate in the CMC framework and achieves competitive performance on MSCOCO dataset. Junlong Gao, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2023 | An Efficient Rate Control Scheme for Video Compression in Low-latency Interoperable InterfacesabstractLightweight video compression has effectively alleviated the tension between growing transmission demands and expensive integration upgrades. Effective rate control algorithms are believed to be the crucial bottleneck for quality improvement during those ultra-high throughput coding processes. This paper proposes a novel rate control (RC) scheme that constructs a contextual adaptive bit estimation model through clustering historical compression information into block-gradient complexity categories. A buffer-aware tuning method and a flexible quantization parameter (QP) mapping algorithm are designed to determine the Luma/Chroma QP distribution where a simplified Lagrangian multiplier is further defined to preserve the stability of the overall compression process. As a result, the constant-bitrate compression towards low-latency interoperable ASICs is implemented with a promising RC performance. Huiwen Ren, Zetian Song, Yan Wang 0011, Shanshe Wang, Fangdong Chen, Shiliang Pu, Siwei Ma 0001, Wen Gao 0001 |
DCC | 5 |
| 2023 | An Adaptive Intra-frame Quantization Parameter Derivation Model Jointing with Inter-frame AnalysisabstractThis paper proposes a novel quantization parameter (QP) derivation module that constructs several spatiotemporal characteristics into key-frame QP determination through an efficient pre-analysis progress. A series of simplified prediction modes and a histogram statistic are employed to model the reference quality that key-frames provide to subsequent frames. An adaptive delta-QP value is generated to address the conflict between the low compression efficiency of intra-only frames and the critical predictive basis of temporal-underlying frames. The experimental result shows that the proposed method reduces 41.91% of peak-to-valley bitrate difference while leading to a 0.02% BDBR performance change, indicates that high-quality key-frames may not be indispensable in nowadays video compression frameworks. The proposed method has shown that the pre-analysis based QP optimization for intra-only frames is promising for the enhancement of transmission bandwidth utilization, which may hopefully provide new inspiration for bit allocation and rate control designs. Huiwen Ren, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2022 | A Smart Reference Picture Resampling Approach for VVCabstractResampling-based coding, i.e. down-sampling before encoding and up-sampling after decoding, has been recognized to be an effective tool for compressing high-resolution videos at low bitrates. The newest video coding standard, Versatile Video Coding (VVC), supports resampling-based coding via a mechanism named Reference Picture Resampling (RPR), where the spatial resolution can be changed without inserting an intra frame. Intuitively, it is not wise to utilize a single resolution throughout the whole video, because frames with different contents may prefer different coding resolutions. In this paper, we propose a smart reference picture resampling approach, namely smart-RPR, where the coding-resolution of a frame is determined based on the property of the frame without multiple-pass encoding. Specifically, we first down- and up-sample a frame without considering compression and compare the up-sampled frame with the original frame to obtain the resampling distortion, which is then compared with a threshold to decide whether to code the frame in a resampling way. Then, we build up an exponential model to approximate the optimal threshold. In addition, we also study how to derive the coding parameters of the down-sampled frame to achieve better performance. Simulation results on the VTM-12.0 show that the proposed method could achieve 2.72%, 5.29%, and 10.82% BD-rate reductions for Y, Cb, and Cr components, respectively, with lower encoding and decoding complexity. Tianliang Fu, Kai Zhang 0007, Yue Li 0015, Li Zhang 0006, Shanshe Wang, Siwei Ma 0001 |
DCC | 5 |
| 2022 | Rate Distortion Characteristic Modeling for Neural Image CompressionabstractEnd-to-end optimized neural image compression (NIC) has obtained superior lossy compression performance recently. In this paper, we consider the problem of rate-distortion (R-D) characteristic analysis and modeling for NIC. We make efforts to formulate the essential mathematical functions to describe the R-D behavior of NIC using deep networks. Thus arbitrary bit-rate points could be elegantly realized by leveraging such model via a single trained network. We propose a plugin-in module to learn the relationship between the target bit-rate and the binary representation for the latent variable of auto-encoder. The proposed scheme resolves the problem of training distinct models to reach different points in the R-D space. Furthermore, we model the rate and distortion characteristic of NIC as a function of the coding parameter$\lambda$respectively. Our experiments show our proposed method is easy to adopt and realizes state-of-the-art continuous bit-rate coding performance, which implies that our approach would benefit the practical deployment of NIC. Chuanmin Jia, Ziqing Ge, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2022 | Coarse-to-fine Prediction With Local and Nonlocal Correlations for Intra CodingabstractRecently many efforts have been devoted to learning non-linear predictions from neighboring samples with deep neural networks. However, existing methods mainly generate predictions with local reference samples, regardless of nonlocal self-similarity. In this paper, we aim to incorporate local and nonlocal correlations for intra prediction and propose a two-stage coarse-to-fine network (CTFN), which is integrated into VVC codec as an optional intra prediction mode. The prediction process of CTFN is decomposed into two stages. In the first stage, we train a set of networks to generate a coarse result with local reference samples. In the second stage, we extract sufficient features from nonlocal region using the coarse result as priors and transform the features into a fine prediction result. In particular, a patch-wise attention layer (PAL) is designed in the second stage that can fully explore nonlocal correlations in feature domain and assign weights to each nonlocal feature adaptively, as shown in Fig. 1. As such, the proposed CTFN can not only learn a non-linear mapping from local context, but also explicitly borrow similar features from nonlocal region in a weighted form. Different from image inpainting tasks, the patch synthesis problem is converted to patch matching problem with the CTFN, yielding more reliable predictions. More-over, we construct a classified dataset based on Pearson Correlation Coefficient for network training to better handle contents that are highly correlated. Experiments on VTM-11.0 show that the proposed network achieves 1.77% ED-rate reductions under all intra configuration, which outperforms the state-of-the-art methods. Meng Lei, Xuewei Meng, Chuanmin Jia, Shanshe Wang, Zhipeng Cheng, Siwei Ma 0001 |
DCC | 4 |
| 2022 | Parametric Non-local In-loop Filter for Future Video CodingabstractIn-loop filter has been comprehensively explored during the development of video coding standards to suppress compression artifacts. However, the existing in-loop filters in Versatile Video Coding (VVC) mainly take advantage of the image local similarity. Although some non-local based in-loop filters can make up for this short-coming, the unsupervised parameter selection scheme, which is widely used by non-local filters, limits the content adaptability. Given this, we propose a parametric non-local in-loop filter (PNLF) that fully considers the non-local characteristics and trains the filter coefficients based on the video content. In the filtering process, the reference samples based on the non-local similarity are first derived for each to-be-filtered sample. Then to-be-filtered samples are grouped into specific classes based on multiple features. For each class, filter coefficients are online trained in the encoder and transmitted to the decoder. Finally, the filtering process is conducted using the online-selected coefficients. Simulation results reveal that the proposed approach achieves 0.70%, 1.43%, and 2.09% bit-rate savings on average compared to VTM-11.0 under All Intra (AI), Random Access (RA), and Low-Delay B (LDB) configurations, respectively. The sequences used in the experiment include Class AI, A2, B, C, D, E, F, and SCC. Compared to the non-local structure-based filter (NLSF) [1], our proposed PNLF with fast block matching scheme [2] applied on B-frames and P-frames can achieve better performance gain with lower software and hardware complexity under RA and LDB configurations. Xuewei Meng, Chuanmin Jia, Xinfeng Zhang 0001, Meng Lei, Shanshe Wang, Lin Li 0062, Siwei Ma 0001 |
DCC | 5 |
| 2022 | Fast Partition Mode Decision via a Plug-in Fully Connected Network for Video CodingabstractFlexible coding unit partitioning such as quad-tree nested binary-tree and ternary-tree adopted by the emerging enhanced compression model (ECM) brings promising coding performance improvement. Meanwhile, the computational complexity increases dramatically, which may block the exploration and validation of new coding tools. This paper investigates a partition mode early pruning scheme via a fully connected network to reduce the encoding complexity for the ECM. In particular, we carefully select features and devise the fully connected network, which could seamlessly cooperate with the encoder, revealing promising learning and inference capability. Experimental results demonstrate that the proposed method achieves 15%~50% encoding time savings with moderate bit-rate increasing on the ECM, and the extra complexity regarding the fully connected network and feature extraction is negligible. Jiaqi Zhang 0007, Meng Wang 0017, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 5 |
| 2021 | Intra Block Partition Structure Prediction via Convolutional Neural NetworkabstractIn video coding, block partition segments images into non-overlap blocks for individual coding, the structure of which is becoming more and more flexible along with the development of video coding standards. Multiple types of tree structures have been proposed recently, which extensively improved the complexity of the encoding process due to the recursive rate-distortion search for the optimal partition. In this paper, a two-stage Convolutional Neural Network (CNN) based partition structure prediction method is proposed to bypass the decision process of the block size in intra frame coding. Specifically, the Coding Unit (CU) partition is first represented in sub-block granularity and predicted by the end-to-end trained CNNs. Then, the final partition structure compatible with the coding standard is derived from the prediction results directly. Experimental results show that the proposed CNN based partition method achieves about 56 times speedup (97% time-saving) with 9% BD-rate degradation against the reference software of the latest AVS3 coding standard (IEEE Standard 1857.10). Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2021 | Quad-Treea Based Sample Refinement Filter for Video CodingabstractIn-loop filter is a crucial module in video coding, which can improve both subjective and object quality of reconstructed videos. In this paper, a new sample-based classification method is first proposed using features extracted from different stages of the existing in-loop filter process. Based on this method, an adaptive three-layer Quad-tree Based Sample Refinement Filter (QSRF) algorithm is designed to further improve the coding efficiency. Experimental results show that the proposed QSRF algorithm achieves 0.39%, 0.77% and 0.70% BD-rate savings for random access, lowdelay B and lowdelay P configurations compared to AVS3 reference software, respectively. Moreover, the proposed method can also improve visual quality of reconstructed videos significantly. Yunrui Jian, Jiaqi Zhang 0007, Chuanmin Jia, Suhong Wang, Shanshe Wang, Siwei Ma 0001 |
DCC | 5 |
| 2021 | Optimized Adaptive Loop Filter in Versatile Video CodingabstractIn the Versatile Video Coding (VVC) standard, adaptive loop filter (ALF), including Geometry transformation-based Adaptive Loop Filter (GALF) and Cross Component Adaptive Loop Filter (CCALF), plays an essential role in reducing compression artifacts. However, it also has high coding complexity and requires many picture buffer accesses in the encoder that will increase external memory access and is unfriendly to the software and hardware design. Therefore, we propose an optimized ALF framework, including the parallel design of GALF and CCALF, the adaptive parameter decision of GALF, and one-pass CCALF scheme by effectively estimating the CCALF filtering distortion without conducting filter operation. Compared to VTM-8.0, the proposed method can reduce the picture buffer access from 152 to 1 and achieve roughly 25% time-savings of the ALF module with negligible coding performance change under RA configuration. Some of the proposed methods have been adopted in the VVC reference software. Xuewei Meng, Jiaqi Zhang 0007, Chuanmin Jia, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
DCC | 5 |
| 2021 | Flow-Grounded Dynamic Texture Synthesis for Video CompressionabstractThe basic ingredients of modern video coding standards are block-based prediction and transforms. However, when dealing with video contents containing dynamic textures (DT), the existing prediction schemes usually failed due to temporal variability and randomness of DT, which results in more bit cost on residual coding compared with other contents. In view of this point, a novel video compression scheme for DT is proposed in this work. In particular, wavelet-based analysis on motion characteristics of DT is firstly presented and based on the analysis, we introduce a flow-grounded texture synthesis method for video compression. Instead of conventional inter prediction, synthesized DT contents are used for reconstruction at the decoder. The proposed scheme has been fully integrated into the test model of Versatile Video Coding standard, VTM-10.0, for validation and a subjective test has also been carried out. Experimental results show that bitrate savings can be achieved by 40% on average at comparable visual quality. Suhong Wang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2020 | Sub-Sampled Cross-Component Prediction for Chroma Component CodingabstractCross-component prediction, which takes advantage of inter-channel correlations, predicts the chroma block with the luma reconstructed block according to associated linear model. Instead of involving all available reference samples in building the linear model, in this paper, we propose a sub-sampled approach that utilizes at most four neighboring chroma samples and their corresponding down-sampled luma samples, leading to significantly reduced operations in the derivation of model parameters at both encoder and decoder. The proposed scheme is hardware friendly in terms of the overheads of memory access and clock cycles, and greatly benefits the practical implementations of the emerging video coding standard in real applications. Extensive experiments reveal that the proposed sub-sampled method provides simple operations and robust coding performance, leading to the adoption by Versatile Video Coding (VVC) Standard and the third generation Audio Video Coding Standard (AVS3). Meng Wang 0017, Li Zhang 0006, Kai Zhang 0007, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 6 |
| 2019 | Perceptual Video Coding Based on Visual Saliency Modulated Just Noticeable DistortionabstractTo reduce the perceptual redundancy in the video coding process, human visual system (HVS)-based visual attention and visual sensitivity can be utilized due to their intrinsic natures. Just Noticeable Distortion (JND) is one of widely used models to simulate human visual sensitivity, while visual saliency map has been popular for years in image processing to describe the visual attention feature, which has been proved by the ability to enhance the visual sensitivity effect. In this paper, we proposed a perceptual video coding (PVC) scheme with visual saliency modulated JND model to suppress the DCT coefficient without resulting in noteworthy subjective quality degradation. The experimental results show that the PVC scheme with the proposed VS-JND model can save bit rates up to 35.58% in high bit rates case with the similar subjective quality compared with that of HEVC software reference code HM 16.12. Ruiqin Xiong, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
DCC | 4 |
| 2019 | Adaptive Wavelet Domain Filter for Versatile Video Coding (VVC)abstractOwing to the ability of removing compression artifacts, extensive in-loop filters have been proposed for video coding standards. They are performed after the reconstruction of all coding units (CUs), however, none of them has been taken into account in the mode decision when coding each CU. To address this issue and make the rate-distortion optimization (RDO) more precise for each CU, we introduce a low-pass filter when checking the rate-distortion cost after the reconstruction of each CU. Specifically, based on Haar wavelet, the reconstructed block is transformed to the frequency domain, and then an adaptive wavelet domain filter (AWF) is proposed to suppress the quantization noises in coded blocks. To be adaptive, the filter strength varies from CU to CU according to the texture complexity and quantization parameters (QPs). Experimental results show that the proposed method can reduce the compression artifacts and improve both the objective and subjective quality. Suhong Wang, Xiang Zhang 0004, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2018 | Locally Refined Motion Compensation for Future Video CodingabstractMotion compensation plays a key role in high efficiency video coding. The popular video compression standards, such as H.264/AVC and HEVC, adopt block based motion compensation technique due to its high compression efficiency and relatively low computational complexity. However, block based motion compensation may not be in accordance with the actual object boundary, potentially leading to low prediction accuracy especially in the high-texture areas. In this paper, we propose a locally refined motion compensation method to address this issue. In particular, the image segmentation is applied on the prediction block indicated by a motion vector rather than the original block to avoid explicit signaling. Furthermore, the local content is analyzed to select one segmented region and subsequently the prediction of this region is generated based on the local motion filed. Experimental results show that the proposed algorithm can achieve 0.8%, 1.1% and 1.7% bitrate savings for Random Access, Lowdelay-B and Lowdelay-P configurations respectively without introducing noticeable computational complexity. Zhao Wang 0004, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
DCC | 4 |
| 2017 | Effective Quadtree Plus Binary Tree Block Partition Decision for Future Video CodingabstractBlock partition structure has been recognized as a crucial module in video coding scheme. Recently, a quadtree plus binary tree (QTBT) block partition structure has been proposed in the Joint Video Exploration Team (JVET) development. Compared to the quadtree structure in HEVC, QTBT can achieve better coding performance with hugely increased encoding complexity. Here, we propose an effective QTBT partition decision algorithm to achieve a good trade-off between computational complexity and coding performance. In particular, at the Coding Tree Unit level, the partition parameters of QTBT are dynamically derived to adapt to the local characteristics without transmitting any overhead. Subsequently, at the Coding Unit level, a joint-classifier decision tree structure is designed to eliminate unnecessary iterations and meanwhile control the risk of false prediction. Experimental results show that the proposed algorithm can achieve 64% encoding time reduction on average with only 1.26% increase in terms of bit rate. This greatly benefits the practical implementations of QTBT in real application scenarios. Zhao Wang 0004, Shiqi Wang 0001, Jian Zhang 0018, Shanshe Wang, Siwei Ma 0001 |
DCC | 4 |
| 2016 | From Visual Search to Video Compression: A Compact Representation Framework for Video Feature DescriptorsabstractVisual feature descriptors have been successfully deployed in a wide range of applications, e.g. visual retrieval and analysis. To transmit these descriptors over bandwidth-limited networks, a high efficiency feature coding technique is highly desired to maximize compression capability and achieve compact feature representations. In this paper, a hybrid visual feature descriptor compression framework is presented and implemented in the encoding and decoding loops of texture videos. In particular, the multiple-hypothesis prediction is employed to effectively remove redundancies originated not only from spatial and temporal similarities, but also from reconstructed video frames. As the ultimate purpose of the transmitted descriptors is retrieval, the rate-accuracy optimization (RAO) technique is proposed to obtain the best tradeoff between the rate and retrieval performance. Such paradigm enables the conventional video stream to achieve high efficient retrieval/analysis with very low bitrate consumption. Moreover, we also demonstrate that texture video compression can also benefit from the additional information provided by the transmitted descriptors, leading to significantly improvement of coding efficiency on top of the high efficiency video coding (HEVC) standard. Extensive simulations have shown that the proposed method can offer significant bitrate reduction in representing both the descriptors and texture video frames, and meanwhile providing desirable retrieval performance. Xiang Zhang 0004, Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Xinfeng Zhang 0001, Wen Gao 0001 |
DCC | 4 |
| 2013 | Low Complexity Rate Distortion Optimization for HEVCabstractThe emerging High Efficiency Video Coding (HEVC) standard has improved the coding efficiency drastically, and can provide equivalent subjective quality with more than 50% bit rate reduction compared to its predecessor H.264/AVC. As expected, the improvement on coding efficiency is obtained at the expense of more intensive computation complexity. In this paper, based on an overall analysis of computation complexity in HEVC encoder, a low complexity rate distortion optimization (RDO) coding scheme is proposed by reducing the number of available candidates for evaluation in terms of the intra prediction mode decision, reference frame selection and CU splitting. With the proposed scheme, the RDO technique of HEVC can be implemented in a low-complexity way for complexity-constrained encoders. Experimental results demonstrate that, compared with the original HEVC reference encoder implementation, the proposed algorithms can achieve about 30% reduced encoding time on average with ignorable coding performance degradation (0.8%). Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Liang Zhao 0007, Qin Yu 0003, Wen Gao 0001 |
DCC | 3 |