Junyan Huo

dblp:67/8798 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Coefficients Energy Guided Fast Transform Algorithm in Beyond VVC
abstract
The Non-Separable Transform (NST) is an effective transform algorithm for improving coding efficiency in H.266/Versatile Video Coding (VVC) and Beyond VVC. In this paper, we propose an adaptive reordering strategy and fast pruning techniques on the encoder side to optimize the transform kernel selection of NST. The absolute sum values of forward transform coefficients (CoefAbsSum) are used as a coefficients energy criterion for evaluating transform kernels: a smaller CoefAbsSum indicates a better transformation of the kernel. Two key technical efforts are designed: we first propose an adaptive reordering for NST transform kernels, which sorts the three transform kernels in each prediction mode's transform set according to CoefAbsSum. It thereby prioritizes kernels with better potential Rate-Distortion (RD) performance. Subsequently, we design three fast algorithms for the sorted kernels, including adjusted RD cost threshold pruning, CoefAbsSum-based pre-screening, and historical RD performance comparison to skip modes with poor preestimates. Experimental results, conducted based on ECM 17.0, show that the proposed framework achieves an average encoding time reduction of 1.3 % (with the encoding time ratio of 98.7 %); meanwhile, it maintains a BD-rate change of 0.00 % for the Y component, 0.03 % for the Cb component, and 0.01 % for the Cr component. The Beyond VVC is an exploration platform with extremely high coding complexity, and our work reduces the complexity while introducing nearly no coding loss.
Minzhe Chen, Junyan Huo, Fuzheng Yang 0001
DCC3
2026 Optimized Adaptive Loop Filter Based on the Refined Adaptation Parameter Sets in H.266/VVC
abstract
Adaptive loop filter (ALF), including luma ALF, chroma ALF, and cross-component adaptive loop filter (CCALF), has been adopted in H.266/versatile video coding (VVC). It can enhance the quality of reconstructed videos based on the Wiener filtering principle. In-depth analyses reveal that the efficiency of chroma ALF and CCALF is significantly lower than that of luma ALF, primarily due to the insufficient number of chroma-related filter sets in the adaptation parameter sets (APSs). To address this limitation, this paper focuses on the efficient design of ALF APSs. Specifically, we propose a luma-guided rate-distortion optimization (RDO) criterion for chroma components to improve the efficiency of the new ALF APS. Furthermore, an improved ALF APS list management is introduced to extend the lifetime of chroma-related filter sets in the ALF APS list.
Junyan Huo, Wenjie Zou, Fuzheng Yang 0001, Shuai Wan
DCC2
2026 Multiple Candidates Derivation of Dominant Intra Prediction Mode in Beyond VVC
Junyan Huo, Yanzhuo Ma, Wei Zhang 0072, Fuzheng Yang 0001, Jiarun Song
ISCAS2
2026 Slim Matrix-Based Intra Prediction for H.266/VVC: A Prediction-Diversity-Driven Design
Junyan Huo, Yundi Gao, Fuzheng Yang 0001
IEEE Signal Process. Lett.2
2026 Text and Non-Text Latent Feature Disentanglement for Screen Content Image Compression
abstract
With the growing prevalence of screen content images in multimedia communication, efficient compression has become increasingly crucial. Unlike natural scene images, screen content typically contains rich text regions that exhibit unique characteristics and low correlation with surrounding non-text elements. The intricate mixture of text and non-text within images poses significant challenges for existing learned compression networks, as the text and non-text features are severely entangled in the latent domain along the channel dimension, leading to compromised reconstruction quality and suboptimal entropy estimation. In this paper, we propose a novel Disentangled Image Compression Architecture (DICA) that enhances the analysis module and the entropy model of existing compression architectures to address these limitations. First, we introduce a Disentangled Analysis Module (DAM) by augmenting original analysis modules with an additional text approximation branch and a disentangling network. They work in concert to disentangle latent features into text and non-text classes along the channel dimension, resulting in a more structured feature distribution that better aligns with compression requirements. Second, we propose a Disentangled Channel-Conditional Entropy Model (DCEM) that efficiently leverages the feature distribution bias introduced by DAM, thereby further improving compression performance. Experimental results demonstrate that the proposed DICA, along with DAM and DCEM can be integrated into various channel-conditional compression backbones, significantly improving their performance in screen content compression—particularly in hard-to-compress text regions. When integrated with an advanced WACNN backbone, our method achieves a 13% overall BD-Rate gain and a 16% BD-Rate gain in text regions on the SIQAD dataset.
Hao Wang 0184, Junyan Huo, Fei Yang 0004, Shuai Wan, Gaoxing Chen, Luis Herranz, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Fast Encoding Method for Cloud Gaming Videos Based on Effective Validation of RG-MVs
abstract
Cloud gaming is a popular paradigm in the gaming industry, allowing real-time gaming experiences powered by cloud computing. Due to a strict latency sensitivity, efficient video encoding is crucial for its wide application. Cloud gaming videos are produced by game engines, which provide auxiliary data such as rendered generative motion vectors (RG-MVs). These motion vectors, calculated by 3D geometry and representing ground-truth motion, offer significant potential for fast video coding design.
Junyan Huo, Fuzheng Yang 0001
DCC4
2025 Customizing Image Codecs for Text-Rich Screen Content with Plugin Processing Networks
abstract
With the rapid growth of remote education, telemedicine, and cloud gaming, screen content images have become prevalent in these applications. They differ significantly from natural scene images, making learning-based image codecs optimized with natural scenes inefficient when compressing them. Through empirical analysis, we observe the textual region in screen content is not only hard to compress in itself but also impacts the compression efficiency of the non-textual region. To customize the image codecs to screen content without altering their parameters, we introduced plugin pre- and post-processing modules. Specifically, we designed a filtering network in the pre-processing module to remove compression-unfriendly information from textual regions and a restoration network in the post-processing module to recover it. Additionally, we implemented a multi-scale fuse approach to enhance the high-frequency details in images. Experiments on public datasets demonstrated that our plugin solution can be seamlessly integrated into learning-based image codecs, significantly improving compression performance.
Hao Wang 0184, Junyan Huo, Shuai Wan, Gaoxing Chen, Fuzheng Yang 0001
ICME2
2025 OMR-Net+: A Frequency-Aware Feature Refinement and Entropy Modeling Method for Efficient Screen Content Image Compression
abstract
Screen content image (SCI) compression faces challenges due to distinct characteristics such as sharp edges and repetitive structures. Existing learned image compression methods encounter two key issues: 1) insufficient frequency-aware processing, and 2) suboptimal entropy modeling for mixed-frequency components. To this end, we propose OMR-Net+, a novel SCI compression method that incorporates frequency-aware feature characteristics, including a frequency-aware refinement network (FARN) and a frequency-aware entropy model (FAEM). The proposed FARN uses an invertible neural network to preserve critical high-frequency details and a transformer-based model to reduce redundancy in low-frequency features. Additionally, the proposed FAEM provides tailored conditional probability estimation based on a parallel context model for high- and low-frequency features, respectively, to improve both coding performance and computational efficiency. Experimental results on the SCID and SIQAD datasets show that OMR-Net+ significantly outperforms the previous OMR-Net and other state-of-the-art methods in rate-distortion performance, demonstrating its potential for efficient SCI compression.
Shiqi Jiang 0006, Ting Ren, Hui Yuan 0001, Junyan Huo, Xin Lu 0001
IEEE Signal Process. Lett.4
2025 Adaptive Enhanced Global Intra Prediction for Efficient Video Coding in Beyond VVC
abstract
Global intra prediction (GIP), including intra-block copy and template matching prediction (TMP), exploits the global correlation of the same image to improve the coding efficiency. In Beyond VVC, TMP uses template matching to determine the reference blocks for efficient prediction. There usually exists an error between the coding block and reference blocks, caused by the content mismatch or the coding distortion of the reference blocks. We propose an enhancement over the reference blocks, namely enhanced GIP (EGIP). Specifically, we design an enhanced filter according to the templates of the coding block and the reference blocks, with the reconstructed template of the coding block as the label for supervised learning. To support different enhancements, we design two types of inputs, i.e., EGIP based on neighboring samples (N-EGIP) and EGIP based on multiple hypothesis references (M-EGIP). Experimental results show that, based on enhanced compression model (ECM) version 8.0, N-EGIP achieves BD-rate reductions of 0.37%, 0.42%, and 0.40%, and M-EGIP brings 0.34%, 0.37%, and 0.34% BD-rate savings for Y, Cb, and Cr components, respectively. A higher coding gain, 0.46%, 0.54%, and 0.52% BD-rate savings, can be achieved by integrating N-EGIP and M-EGIP together. Owing to the coding gain and small complexity increase, the proposed EGIP has been adopted in the exploration of Beyond VVC and integrated into its reference software.
Junyan Huo, Yanzhuo Ma, Zhenyao Zhang, Hui Yuan 0001, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Adaptive Chroma Block Vector Derivation from Luma for Screen Content Coding
abstract
Intra Block Copy (IBC) and Intra Template Matching Prediction (IntraTMP) are two efficient algorithms to sufficiently exploit the correlation in the same picture. Block Vector (BV) is used to represent the displacement between the current block and its reference within the same picture. The BV information of luma can be employed to help the chroma coding efficiently. Based on this feature, an adaptive chroma prediction is proposed to derive the BV of the chroma block from the luma. Two strategies are designed to improve the coding performance, including multiple positions’ check and template-based BV refinement. Compared with Enhanced Compression Model (ECM) of beyond VVC, 0.43%, 0.35%, and 0.60% BD-rate savings for Y, Cb, and Cr components are achieved for Class F, and 2.23%, 2.31%, and 2.93% BD-rate savings are provided for Class TGM. We also integrated the proposed method into the VVC Test Model (VTM). A similar coding improvement can be observed. Due to the coding gain and low complexity, the proposed method has been adopted into the beyond VVC exploration and integrated into the latest version of ECM.
Junyan Huo, Xue Hao, Shuai Wan, Fuzheng Yang 0001
ICASSP1
2024 A Transformer-Based Intra Luma Enhancement for H.266/VVC
abstract
Intra prediction is essential in reducing spatial domain correlation in video coding. To improve intra prediction accuracy, we introduce a transformer-based quality enhancement method aiming atimproving the luma quality of reconstructed coding tree units (CTUs). Our transformer-based model, namely Enhanceformer, utilizes multi-head attention for comprehensive feature extraction across multiple stages, levels, and scales. By integrating the model into H.266/VVC codec, it not only improves the luma quality of the current reconstructed CTU, but also provides more accurate references for intra prediction of subsequent CTUs. Experimental results show that this method achieves average BD rate savings of 2.63%, 0.21% and 0.48% for Y, Cb and Cr components respectively in all intra configuration, outperforming H.266/Versatile Video Coding (VVC) anchor.
Wenrui Lv, Hui Yuan 0001, Congrui Fu, Shiqi Jiang 0006, Junyan Huo
PCS5
2024 Refined Chroma From Luma Prediction in AV1 Based on Color Component Grouping
abstract
Chroma from luma (CfL) in AOMedia Video 1 (AV1) utilizes the correlation between color components to derive the predicted chroma from the reconstructed luma. The mean of the predicted chroma, i.e., dc, is set to the chroma average of the neighboring regions. Since the neighboring regions have limited reference samples, the distribution of the coding block is hard to be predicted from the neighboring regions, resulting in an error in the predicted dc. A new feature of CfL is that the luma of the coding block has been reconstructed. Using the luma information, the ground-truth distribution of the coding block can be built. Based on this observation, a refined dc prediction is proposed based on color component grouping (GDC). We design an adaptive grouping scheme and use the number of luma and the chroma average in each group to derive a refined dc. The offline experiment verifies that the proposed method provides a dc with high prediction accuracy. Compared with the libaom of AV1, the proposed GDC achieves BD-rate reductions of 0.15%, 0.90%, and 0.98% for the Y, Cb, and Cr components. With the increased group numbers, additional coding gains can be provided. The proposed GDC can be implemented with a small complexity, which is hardware-friendly.
Junyan Huo, Zhenyao Zhang, Jiarun Song, Yanzhuo Ma, Fuzheng Yang 0001
IEEE Trans. Ind. Informatics1
2024 Joint Rate-Distortion Optimization for Video Coding and Learning-Based In-Loop Filtering
abstract
Learning-based in-loop filters (ILFs) have recently been widely deployed in the video codec to remove compression artifacts and to obtain better-quality reconstructed videos. However, in the existing codec, the impact of the learning-based ILF is not considered in the Rate-Distortion optimization (RDO) process. With the learning-based ILF, the set of coding parameters selected by the conventional RDO process may no longer be the best one, and the best overall Rate-Distortion (R-D) performance can not be guaranteed. In this article, we propose a joint RDO (JRDO) for Video Coding and learning-based in-loop filtering, which incorporates the effect of the learning-based ILF on the reconstructed video into the RDO process, aiming to achieve the best overall R-D performance of the reconstructed video after in-loop filtering. Furthermore, to realize the proposed JRDO in a standardized video codec, we propose practical strategies to efficiently estimate the effect of learning-based ILF during the RDO process, i.e., efficiently estimate the distortion of the reconstructed block after in-loop filtering during the RDO process. Extensive experiments demonstrate that the proposed joint RDO is standard-compliant and can improve the R-D performance without increasing the decoding time. Besides, the superiority of joint RDO is achieved in various ILFs, indicating the generality of the proposed work.
Mingyi Yang, Junyan Huo, Xile Zhou, Wenhan Qiao, Shuai Wan, Hao Wang 0184, Fuzheng Yang 0001
IEEE Trans. Multim.2
2023 Enhancing Video Encoding for Cloud Virtual Reality Gaming Based on User Types
abstract
Cloud Virtual Reality (VR) gaming is a novel technology that allows users to enjoy complex games on their thin clients by offloading the graphics rendering to cloud servers. The thin clients only need to perform basic decoding functions, which reduces the hardware requirements and costs. However, cloud VR gaming also faces the challenge of high bandwidth consumption when transmitting high-resolution game video streams. This paper presents a cloud VR gaming system that can transmit users’ gaze point data to the server in real time to identify users’ regions of interest. With this system, we verify the difference in spatial visual sensitivity caused by the different types of users. Then, a user-type-based video encoding method is proposed. Through conducting the subjective test experiment, the proposed video encoding method can reduce the bitrate for players and viewers by at least 71% and 69%, respectively, without compromising the perceptual quality.
Junyan Huo, Fuzheng Yang 0001, Gaoxing Chen
VCIP3
2023 Adaptive Chroma Prediction Based on Luma Difference for H.266/VVC
abstract
Cross-component chroma prediction plays an important role in improving coding efficiency for H.266/VVC. We use the differences between reference samples and the predicted sample to design an attention model for chroma prediction, namely luma difference-based chroma prediction (LDCP). Specifically, the luma differences (LDs) between reference samples and the predicted sample are employed as the input of the attention model, which is designed as a softmax function to map LDs to chroma weights nonlinearly. Finally, a weighted chroma prediction is conducted based on the weights and chroma reference samples. To provide adaptive weights, the model parameter of the softmax function can be determined based on the template (T-LDCP) or offline learning (L-LDCP), respectively. Experimental results show that the T-LDCP achieves BD-rate reductions of 0.34%, 2.02%, and 2.34% for the Y, Cb, and Cr components, and the L-LDCP brings 0.32%, 2.06%, and 2.21% BD-rate savings for Y, Cb, and Cr components, respectively. The L-LDCP introduces slight encoding and decoding time increments, i.e., 2% and 1%, when integrated into the latest VVC test model version 18.0. Besides, the LDCP can be implemented by a pixel-level parallelization which is hardware-friendly.
Junyan Huo, Danni Wang, Hui Yuan 0001, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Image Process.1
2022 Unified Matrix Coding for NN Originated MIP in H.266/VVC
abstract
Matrix-based Intra Prediction (MIP) is an effective coding algorithm in H.266/Versatile Video Coding (VVC) which is originated by Neural Networks (NN). With the requirement of low complexity, MIP is conducted by a matrix-vector multiplication. To handle with the diversity of video content, 30 matrices are trained and stored to derive predicted samples. Since matrices from training are usually floating-point values, which should be avoided in H.266/VVC, two parameters, shift and offset, are introduced for each matrix to convert floating-point values to integers. This paper designs an efficient algorithm to determine the input vector of MIP, with which the range of the matrices can be minimized, and all matrices can be converted to integers with a unified shift and a unified offset. The proposed algorithm removes the matrix-dependent parameters for integer conversion and saves the memory for storing MIP parameters. Experimental results demonstrate that the proposed algorithm has a similar coding performance with VVC reference software. Due to the unified operation, memory reduction, and no coding loss, the proposed algorithm has been adopted into H.266/VVC.
Junyan Huo, Shuai Wan, Fuzheng Yang 0001
ICASSP1
2022 Unified Cross-Component Linear Model in VVC Based on a Subset of Neighboring Samples
abstract
To compress industrial video content efficiently, H.266/Versatile Video Coding (VVC) introduces cross-component linear model (CCLM) prediction as a new coding tool, in which chroma components are predicted from the luma component based on a linear model. In this article, we propose a subset-based CCLM (S-CCLM), in which the model parameters are derived based on a subset of neighboring samples. To choose the most proper subset, we build the relationship between the prediction error and the geometric distance and resolve the optimal subset construction problem by minimizing the geometric distance. With the well-designed subset, a weight-guided parameter derivation algorithm is further proposed to improve the accuracy of the model parameters. The experimental results show that the proposed S-CCLM can achieve Bjontegaard delta bitrate (BD-rate) reductions of 0.14%, 0.64%, and 0.75% for the Y, Cb, and Cr components, respectively, when the number of samples in the subset,$N$, is 4 and BD-rate reductions of 0.22%, 0.80%, and 0.95% when$N$is 8. Given a small fixed$N$, fewer memory access operations are needed during the CCLM calculation, and a unified CCLM process can be achieved for coding blocks with different sizes and different modes. Due to its hardware-friendly architecture, the S-CCLM has been partially adopted by H.266/VVC.
Junyan Huo, Hongqing Du, Shuai Wan, Hui Yuan 0001, Yanzhuo Ma, Fuzheng Yang 0001
IEEE Trans. Ind. Informatics1
2021 Improved Chroma from Luma Prediction in AV1 Based On Virtual Chroma Block Generation
abstract
Chroma from Luma (CfL) prediction is an efficient coding tool in AV1 which builds chroma prediction by implementing a linear model on luma pixels. To avoid transmission, the off-set factor in the linear model is set to the average of neighboring chroma pixels. An improved CfL algorithm is proposed to derive the offset factor based on a virtual chroma block. Such a block is constructed by using the chroma of the matched pixel which is determined according to the luma difference of neighboring regions and the co-located luma. Compared with libaom, the proposed CfL algorithm provides 0.31% and 0.15% weighted PSNR BD-rate saving under AI and RA configuration, respectively. Experimental results show that over 1.00% and 0.80% BD-rate saving can be achieved for chroma components under AI and RA configuration. With the proposed algorithm, the percentage of pixels with CfL as the optimal coding mode is increased.
Junyan Huo, Menglin Zhang, Wenhan Qiao, Fuzheng Yang 0001, Hui Su, Debargha Mukherjee
ICME1
2019 A Fidelity-Assured Rate Distortion Optimization Method for Perceptual-Based Video Coding
abstract
Rate-distortion optimization (RDO) is one of the essential method to improve the video coding efficiency. The main target of RDO is to find the optimal tradeoff between the reconstructed video quality and encoding rate. In traditional video coding standards, e.g. the emerging H.266/Versatile Video Coding(VVC), the H.265/High Efficiency Video Coding (HEVC) and H.264/Advanced Video Coding (AVC), sum of squared error (SSE) is used as the distortion criterion because SSE can represent the image fidelity efficiently. Based on the existing research on human visual characteristic, the perceptual visual quality is not consistent with image fidelity. Accordingly, we propose a video coding method to improve the subjective quality while avoiding great fidelity degradation. Experimental results demonstrate that the proposed method is efficient in preserving subjective qualities with only a little fidelity degradation when comparing to existing subjective quality based rate distortion optimization methods.
Hui Yuan 0001, Junyan Huo
ICIP3
2019 A 3D Haar Wavelet Transform for Point Cloud Attribute Compression Based on Local Surface Analysis
abstract
Point cloud is a main representation of 3D scenes. It is widely applied in many fields including autonomous driving, heritage reconstruction, virtual reality and augmented reality. The data size of this type of media is massive since it contains numerous points with each associated with a large amount of information including geometric coordinate, color, reflectance, and normal. It is thus of great significance to investigate the compression of point cloud data to boost its application. However, developing efficient point cloud compression method is challenging mainly due to the unstructured nature and nonuniform distribution of the data. In this paper, we propose a novel point cloud attribute compression algorithm based on Haar Wavelet Transform (HWT). More specifically, the transform is performed taking into account the surface orientation of point cloud. Experimental results demonstrate that the proposed method outperforms other state-of-the-art transforms.
Sujun Zhang, Wei Zhang 0072, Fuzheng Yang 0001, Junyan Huo
PCS4
2018 A novel distortion criterion of rate-distortion optimization for depth map coding
Ziqi Zheng, Junyan Huo, Hui Yuan 0001, Weisi Lin
J. Vis. Commun. Image Represent.2
2018 Adaptive Lagrangian Multiplier derivation model for depth map coding
Ziqi Zheng, Junyan Huo, Hui Yuan 0001
Signal Process. Image Commun.2
2018 Fine Virtual View Distortion Estimation Method for Depth Map Coding
abstract
In three-dimensional (3-D) video coding systems, depth maps represent the geometric information of a 3-D scene. Since depth maps are not displayed to viewers but to generate virtual views, the quality of depth maps needs to be measured by its effect on the virtual view quality, which is indicated by the virtual view distortion (VVD) in depth map coding. In this letter, a fine VVD estimation method is proposed based on the analysis of the VVD. Specifically, the depth distortion of the current pixel, the texture gradient of the colocated color video and the depth distortions of adjacent pixels are all taken into consideration to estimate the VVD accurately. Experimental results demonstrate that the proposed method can improve 13.1% bitrate saving compared with the sum of squared differences based depth distortion calculation method and can improve 1.2% bitrate saving compared with the VVD estimation method in three dimensional high efficiency video coding (3D-HEVC) reference software.
Ziqi Zheng, Junyan Huo, Hui Yuan 0001
IEEE Signal Process. Lett.2
2011 Model-Based Joint Bit Allocation Between Texture Videos and Depth Maps for 3-D Video Coding
abstract
In 3-D video coding, texture videos and depth maps need to be jointly coded. The distortion of texture videos and depth maps can be propagated to the synthesized virtual views. Besides coding efficiency of texture videos and depth maps, joint bit allocation between texture videos and depth maps is also an important research issue in 3-D video coding. First, we present comprehensive analyses on the impacts of the compression distortion of texture videos and depth maps on the quality of the virtual views, and then derive a concise distortion model for the synthesized virtual views. Based on this model, the joint bit allocation problem is formulated as a constrained optimization problem, and is solved by using the Lagrangian multiplier method. Experimental results demonstrate the high accuracy of the derived distortion model. Meanwhile, the rate-distortion (R-D) performance of the proposed algorithm is close to those of search-based algorithms which can give the best R-D performance, while the complexity of the proposed algorithm is lower than that of search-based algorithms. Moreover, compared with the bit allocation method using fixed texture and depth bits ratio (5:1), a maximum 1.2 dB gain can be achieved by the proposed algorithm.
Hui Yuan 0001, Yilin Chang, Junyan Huo, Fuzheng Yang 0001, Zhaoyang Lu
IEEE Trans. Circuits Syst. Video Technol.3
2009 Scalable Prediction Structure for Multiview Video Coding
abstract
Both temporal prediction and inter-view prediction are employed to improve the coding efficiency in multiview video coding. Hierarchical B pictures are usually used as the basic structure for temporal prediction. The inter-view prediction in each temporal hierarchy level brings different improvement to the entire coding efficiency. We propose a scalable prediction structure in which inter-view prediction would be disabled if the picture redundancy can be almost exploited by temporal prediction and intra prediction. In this way, time-consuming computation of disparity estimation can be saved. Experimental results show that lower encoding complexity, smaller decoded picture buffer size and better random access ability can be achieved with a slightly gain loss.
Junyan Huo, Yilin Chang, Yanzhuo Ma
ISCAS1
2009 Fine-Granular Motion Matching for Inter-View Motion Skip Mode in Multiview Video Coding
abstract
As motion of neighboring views is highly correlated in multiview video systems, the inter-view motion skip mode has been proposed to improve the efficiency of multiview video coding (MVC) by reusing the motion information of neighboring views. To fully exploit the inter-view motion correlation, a fine-granular motion matching algorithm for the motion skip mode is presented in this paper. The proposed algorithm searches for the best matching motion information in neighboring views with fine granularity, and then uses the motion information for the motion-compensated coding in the motion skip mode. The rate-distortion optimization criterion is applied to find the best matching motion information. Experimental results show that the proposed algorithm can improve the performance of the motion skip mode, and further improve the efficiency of MVC.
Haitao Yang 0001, Yilin Chang, Junyan Huo
IEEE Trans. Circuits Syst. Video Technol.3