Xiaozhong Xu

dblp:23/2240 · DBLP profile ↗
← Back
54ranked-venue papers
12as first author
42since 2021 · last 2026
0000-0002-1309-5470ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 11 first-author · 40 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Mamba-Based Perceptual Loss Function for Learning-Based UGC Transcoding
Zihao Qi, Chen Feng 0008, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001
QoMEX4
2025 Gop-Level Adaptive Resampling with CNN-based Super Resolution
abstract
Recently, deep learning-based super resolution methods have been studied for resampling-based video coding to compress high resolution images with limited bandwidth. In this paper, a group of picture-level (GOP-level) adaptive resampling method with convolutional neural network-based (CNN-based) super resolution is proposed to improve the coding gains beyond the latest video coding standard, named Versatile Video Coding (VVC). Specifically, to better restore the detailed information of high resolution video, a super resolution network using multiple side information is first proposed for generating the up-sampled videos. Besides, to further improve the overall performance, an encoder decision strategy is proposed to adaptively select the best scale factor from ×1.0 (original size) and ×2.0 (half size) to determine the encoding resolution at the GOP level. Experimental results demonstrate that the proposed method achieves {-5.34%, -2.38%, -2.08%} and {-3.35%, -8.02%, -4.98%} BD-rate savings for {Y, U, V} under random access and all intra configurations, respectively. This proposed method has been adopted into the reference software of JVET-NNVC.
Renjie Chang, Xiaozhong Xu, Shan Liu 0001
ICIP3
2025 Enhancing HDR Video Compression based on Deep Effective Bit Depth Adaptation
abstract
It is well known that high dynamic range (HDR) videos enhance immersive visual experiences compared to conventional standard dynamic range content. However, HDR content is typically more challenging to encode due to the increased detail associated with the wider dynamic range. In this work, we improve HDR compression performance using an Effective Bit Depth Adaptation approach (EBDA), which reduces the effective bit depth of the original video content before encoding and reconstructs the full bit depth using a CNN-based up-sampling method at the decoder. The up-sampling deep network is based on a new version of Multi-frame MFRNet, MF-MFRNet. This approach has been integrated into the EBDA framework with two Versatile Video Coding (VVC) reference models: VTM 16.2 and the Fraunhofer Versatile Video Encoder (VVenC 1.4.0). The proposed approach has been evaluated under the JVET HDR Common Test Conditions using the Random Access configuration. The results show evident coding gains over both the original VTM 16.2 and VVenC 1.4.0 on all JVET HDR tested sequences, with average bitrate savings of 3.1% and 4.8% based on PSNR and 7.8% and 9.6% based on VMAF against VTM and VVenC respectively. The source code of multi-frame MFRNet has been released at https://github.com/fan-aaron-zhang/MF-MFRNet.
Chen Feng 0008, Zihao Qi, Duolikun Danier, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001
ISCAS5
2025 GeodesicPSIM: Predicting the Quality of Static Mesh With Texture Map via Geodesic Patch Similarity
abstract
Static meshes with texture maps have attracted considerable attention in both industrial manufacturing and academic research, leading to an urgent requirement for effective and robust objective quality evaluation. However, current model-based static mesh quality metrics (i.e., metrics that directly use the raw data of the static mesh to extract features and predict the quality) have obvious limitations: most of them only consider geometry information, while color information is ignored, and they have strict constraints for the meshes' geometrical topology. Other metrics, such as image-based and point-based metrics, are easily influenced by the prepossessing algorithms, e.g., projection and sampling, hampering their ability to perform at their best. In this paper, we propose Geodesic Patch Similarity (GeodesicPSIM), a novel model-based metric to accurately predict human perception quality for static meshes. After selecting a group keypoints, 1-hop geodesic patches are constructed based on both the reference and distorted meshes cleaned by an effective mesh cleaning algorithm. A two-step patch cropping algorithm and a patch texture mapping module refine the size of 1-hop geodesic patches and build the relationship between the mesh geometry and color information, resulting in the generation of 1-hop textured geodesic patches. Three types of features are extracted to quantify the distortion: patch color smoothness, patch discrete mean curvature, and patch pixel color average and variance. To the best of our knowledge, GeodesicPSIM is the first model-based metric especially designed for static meshes with texture maps. GeodesicPSIM provides state-of-the-art performance in comparison with image-based, point-based, and video-based metrics on a newly created and challenging database. We also prove the robustness of GeodesicPSIM by introducing different settings of hyperparameters. Ablation studies also exhibit the effectiveness of three proposed features and the patch cropping algorithm. The code is available at https://multimedia.tencent.com/resources/GeodesicPSIM.
Qi Yang 0003, Joël Jung, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Image Process.3
2025 TDMD: A Database for Dynamic Color Mesh Quality Assessment Study
abstract
Dynamic colored meshes (DCM) are widely used in various applications. However, this kind of meshes may undergo different processes, such as compression or transmission, which can distort them and degrade their quality. To facilitate the development of objective metrics for DCMs and study the influence of typical distortions on their perception, we create the Tencent - Dynamic colored Mesh Database (TDMD) containing eight reference DCM objects with six typical distortions. Using processed video sequences (PVS) derived from the DCM, we conduct a large-scale subjective experiment that resulted in 303 distorted DCM samples with mean opinion scores, making the TDMD the largest available DCM database to our knowledge. This database enables us to study the impact of different types of distortion on human perception and offers recommendations for DCM compression and related tasks. Additionally, we have evaluated three types of state-of-the-art objective metrics on the TDMD, including image-based, point-based, and video-based metrics, on the TDMD. Our experimental results highlight the strengths and weaknesses of each metric, and we provide suggestions about the selection of metrics in practical DCM applications.
Qi Yang 0003, Joël Jung, Timon Deschamps, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Vis. Comput. Graph.4
2024 Transferable Learned Image Compression-Resistant Adversarial Perturbations
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
BMVC5
2024 Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
abstract
No-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference, which have achieved tremendous improvements due to the utilization of deep neural networks. However, learning-based NR-PCQA methods suffer from the scarcity of labeled data and usually perform suboptimally in terms of generalization. To solve the problem, we propose a novel contrastive pre-training framework tailored for PCQA (CoPA), which enables the pre-trained model to learn quality-aware representations from unlabeled data. To obtain anchors in the representation space, we project point clouds with different distortions into images and randomly mix their local patches to form mixed images with multiple distortions. Utilizing the generated anchors, we constrain the pretraining process via a quality-aware contrastive loss following the philosophy that perceptual quality is closely related to both content and distortion. Furthermore, in the model fine-tuning stage, we propose a semantic-guided multi-view fusion module to effectively integrate the features of projected images from multiple perspectives. Extensive experiments show that our method outperforms the state-of-the-art PCQA methods on popular benchmarks. Further investigations demonstrate that CoPA can also benefit existing learning-based PCQA models.
Ziyu Shan, Qi Yang 0003, Haichen Yang, Yiling Xu, Jenq-Neng Hwang, Xiaozhong Xu, Shan Liu 0001
CVPR7
2024 Transferable Learned Image Compression-Resistant Adversarial Perturbations
abstract
With the rapid evolution of advanced image compression, DNN-based learned image compression has emerged as the promising approach for transmitting images in many security-critical applications, such as cloud-based face recognition and autonomous driving, due to its superior performance over traditional compression. There is a pressing need to fully investigate the robustness of a classification system post-processed by learned image compression. To bridge this research gap, we explore the adversarial attack on Learned Image Compression Classification System (LICCS) that targets image classification models that utilize learned image compressors as preprocessing modules. To perform an adversarial attack on an image within the LICCS, the goal is to introduce the adversarial perturbation δ to the source image X that causes the reconstructed adversarial examples gs(Q(ga(X+δ))) to be misclassified by the classification model, which can be formulated as follows:\begin{equation*}\begin{array}{ll} {\mathop {\arg \max }\limits_i f{{\left({{g_s}\left({Q\left({{g_a}\left({{\mathbf{X + \delta }}}\right)}\right)}\right)}\right)}_i} \ne y,}&{{\text{s}}{\text{.t}}{\text{.}}\parallel \delta {\parallel _p} \leq \varepsilon .} \end{array}\tag{1}\end{equation*}
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
DCC5
2024 Reconstruction Distortion of Learned Image Compression with Imperceptible Perturbations
abstract
In this paper, we introduce an imperceptible adversarial attack approach designed to effectively degrade the reconstruction quality of LIC, resulting in the reconstructed image being severely disrupted by noise where identifying any object in the reconstructed image is virtually impossible. More specifically, we generate adversarial examples by introducing a Frobenius norm-based loss function to maximize the discrepancy between original images and reconstructed images from adversarial examples in order to corrupt the reconstructed image severely.
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
DCC5
2024 SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map
abstract
In recent years, static meshes with texture maps have become one of the most prevalent digital representations of 3D shapes in various applications, such as animation, gaming, medical imaging, and cultural heritage applications. However, little research has been done on the quality assessment of textured meshes, which hinders the development of quality-oriented applications, such as mesh compression and enhancement. In this paper, we create a large-scale textured mesh quality assessment database, namely SJTU-TMQA, which includes 21 reference meshes and 945 distorted samples. The meshes are rendered into processed video sequences and then conduct subjective experiments to obtain mean opinion scores (MOS). The diversity of content and accuracy of MOS has been shown to validate its heterogeneity and reliability. The impact of various types of distortion on human perception is demonstrated. 13 state-of-the-art objective metrics are evaluated on SJTU-TMQA. The results report the highest correlation is around 0.6, indicating the need for more effective objective metrics. The SJTU-TMQA is available at https://ccccby.github.io
Bingyang Cui, Qi Yang 0003, Kaifa Yang, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
ICASSP5
2024 Full-Reference Video Quality Assessment for User Generated Content Transcoding
abstract
Unlike video coding for professional content, the delivery pipeline of User Generated Content (UGC) involves transcoding where unpristine reference content needs to be compressed repeatedly. In this work, we observe that existing full-/no-reference quality metrics fail to accurately predict the perceptual quality difference between transcoded UGC content and the corresponding unpristine references. Therefore, they are unsuited for guiding the rate-distortion optimisation process in the transcoding process. In this context, we propose a bespoke full-reference deep video quality metric for UGC transcoding. The proposed method features a transcoding-specific weakly supervised training strategy employing a quality ranking-based Siamese structure. The proposed method is evaluated on the YouTube-UGC VP9 subset and the LIVE-Wild database, demonstrating state-of-the-art performance compared to existing VQA methods. The source code of the developed quality metric and the associated training data are available from https://zihaoq1:github/io/FRUGC/.
Zihao Qi, Chen Feng 0008, Duolikun Danier, Fan Zhang 0017, Xiaozhong Xu, Shan Liu 0001, David Bull 0001
PCS5
2024 Improvements of the BD-Rate Metrics Using Monotonic Curve-Fitting Methods
abstract
The Bj⊘ntegaard Delta rate (BD-rate) measurements have been used as the primary metrics to evaluate performance of video codecs. However, current BD-rate calculation methods are only applicable under the condition that the rate-distortion (R-D) values maintain a monotonic relationship, as this prerequisite is essential for computing integral along the distortion axis. To address this limitation, we propose a curve-fitting based BD-rate solution that guarantees the reconstructed R-D curve to be monotonic. Considering different use cases, we provide a four parameters logistic curve and a constraint cubic curve to approximate the underlying R-D curve. Computation of BD-rate and BD-metric using fitted R-D curve are elaborated in detail. Experimental results indicate that the proposed solutions work well on non-monotonic data. Furthermore, we verified through quantitative analysis that curve-fitting solutions provide more precise measurements of coding efficiency compared to interpolation methods. This improved accuracy contributed by the proposed methods is attributed to the higher resilience to the inherent randomness present in observed data. The proposed method has been adopted by the MPEG WG4 VCM study group for standardization activities. The source code was released at https://multimedia.tencent.com/resources/tvd.
Haiqiang Wang, Xin Zhao 0003, Ding Ding 0004, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001
PCS6
2024 Corner-to-Center long-range context model for efficient learned image compression
Yang Sui 0001, Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001, Bo Yuan 0001, Zhenzhong Chen 0001
J. Vis. Commun. Image Represent.4
2024 Hierarchical Image Feature Compression for Machines via Feature Sparsity Learning
abstract
Recently, Video Coding for Machines (VCM) has gained more and more attention due to its efforts in machine vision tasks. As a crucial track in VCM, feature compression preserves and transmits critical feature information for machine vision. Most existing studies employ dimensionality reduction to the raw multi-scale feature before compression. However, feature sparsity is left insufficiently considered in removing redundancy in compressed features. In this letter, we propose a novel framework for image feature compression for machines, where the multi-scale feature is hierarchically transformed into a sparse representation for compression. The multi-scale feature is first fused by convolutional neural networks and the attention mechanism. To introduce sparsity into the fused feature, informative channels are identified by a channel-wise binary mask where activated elements are sampled from the importance distribution of channels learned from feature content. Then, the fused feature is masked to generate a sparse representation for compression. Experiments conducted on two machine tasks show significant improvements in our model over state-of-the-art methods.
Ding Ding 0004, Zhenzhong Chen 0001, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001
IEEE Signal Process. Lett.4
2024 Deep Reference Frame Generation Method for VVC Inter Prediction Enhancement
abstract
In video coding, inter prediction aims to reduce temporal redundancy by using previously encoded frames as references. The quality of reference frames is crucial to the performance of inter prediction. This paper presents a deep reference frame generation method to optimize the inter prediction in Versatile Video Coding (VVC). Specifically, reconstructed frames are sent to a well-designed frame generation network to synthesize a picture similar to the current encoding frame. The synthesized picture serves as an additional reference frame inserted into the reference picture list (RPL) to provide a more reliable reference for subsequent motion estimation (ME) and motion compensation (MC). The frame generation network employs optical flow to predict motion precisely. Moreover, an optical flow reorganization strategy is proposed to enable bi-directional and uni-directional predictions with only a single network architecture. To reasonably apply our method to VVC, we further introduce a normative modification of the temporal motion vector prediction (TMVP). Integrated into the VVC reference software VTM-15.0, the deep reference frame generation method achieves coding efficiency improvements of 5.22%, 3.61%, and 3.83% for the Y component under random access (RA), low delay B (LDB), and low delay P (LDP) configurations, respectively. The proposed method has been discussed in Joint Video Exploration Team (JVET) meeting and is currently part of Exploration Experiments (EE) for further study.
Jianghao Jia, Yuantong Zhang, Han Zhu 0003, Zhenzhong Chen 0001, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 GPA-Net:No-Reference Point Cloud Quality Assessment With Multi-Task Graph Convolutional Network
abstract
With the rapid development of 3D vision, point cloud has become an increasingly popular 3D visual media content. Due to the irregular structure, point cloud has posed novel challenges to the related research, such as compression, transmission, rendering and quality assessment. In these latest researches, point cloud quality assessment (PCQA) has attracted wide attention due to its significant role in guiding practical applications, especially in many cases where the reference point cloud is unavailable. However, current no-reference metrics which based on prevalent deep neural network have apparent disadvantages. For example, to adapt to the irregular structure of point cloud, they require preprocessing such as voxelization and projection that introduce extra distortions, and the applied grid-kernel networks, such as Convolutional Neural Networks, fail to extract effective distortion-related features. Besides, they rarely consider the various distortion patterns and the philosophy that PCQA should exhibit shift, scaling, and rotation invariance. In this paper, we propose a novel no-reference PCQA metric named the Graph convolutional PCQA network (GPA-Net). To extract effective features for PCQA, we propose a new graph convolution kernel, i.e., GPAConv, which attentively captures the perturbation of structure and texture. Then, we propose the multi-task framework consisting of one main task (quality regression) and two auxiliary tasks (distortion type and degree predictions). Finally, we propose a coordinate normalization module to stabilize the results of GPAConv under shift, scale and rotation transformations. Experimental results on two independent databases show that GPA-Net achieves the best performance compared to the state-of-the-art no-reference PCQA metrics, even better than some full-reference metrics in some cases.
Ziyu Shan, Qi Yang 0003, Rui Ye 0001, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Vis. Comput. Graph.6
2024 TCDM: Transformational Complexity Based Distortion Metric for Perceptual Point Cloud Quality Assessment
abstract
The goal of objective point cloud quality assessment (PCQA) research is to develop quantitative metrics that measure point cloud quality in a perceptually consistent manner. Merging the research of cognitive science and intuition of the human visual system (HVS), in this article, we evaluate the point cloud quality by measuring the complexity of transforming the distorted point cloud back to its reference, which in practice can be approximated by the code length of one point cloud when the other is given. For this purpose, we first make space segmentation for the reference and distorted point clouds based on a 3D Voronoi diagram to obtain a series of local patch pairs. Next, inspired by the predictive coding theory, we utilize a space-aware vector autoregressive (SA-VAR) model to encode the geometry and color channels of each reference patch with and without the distorted patch, respectively. Assuming that the residual errors follow the multi-variate Gaussian distributions, the self-complexity of the reference and transformational complexity between the reference and distorted samples are computed using covariance matrices. Additionally, the prediction terms generated by SA-VAR are introduced as one auxiliary feature to promote the final quality prediction. The effectiveness of the proposed transformational complexity based distortion metric (TCDM) is evaluated through extensive experiments conducted on five public point cloud quality assessment databases. The results demonstrate that TCDM achieves state-of-the-art (SOTA) performance, and further analysis confirms its robustness in various scenarios.
Qi Yang 0003, Xiaozhong Xu, Le Yang 0001, Yiling Xu
IEEE Trans. Vis. Comput. Graph.4
2023 Surface-Sampling Based Objective Quality Assessment Metrics for Meshes
abstract
In this paper, we prove that it is feasible to perform mesh quality assessment by sampling it into point cloud. We propose a general and efficient surface-sampling based framework that can deal with various types and levels of distortions with less complexity. In this method, the original and distorted meshes are first converted into point clouds by sampling the triangle surfaces. Then, the geometry and attribute quality of the distorted mesh can be evaluated by the well-defined point cloud quality metrics. The final objective score can be obtained by fusing multiple quality metrics to get a more accurate prediction of the subjective quality. In addition, we compare the performance in terms of different sampling methods and sampling resolutions on a large public dataset, thus being able to suggest the best sampling configurations.
Chunyang Fu, Xiang Zhang 0004, Thuong Nguyen-Canh, Xiaozhong Xu, Ge Li 0002, Shan Liu 0001
ICASSP4
2023 Exploring the Influence of View and Camera Path Selection for Dynamic Mesh Quality Assessment
abstract
With the development of 3D mesh processing and applications, the quality assessment of dynamic mesh sequences attracts more and more attention. One prevalent strategy for performing the subjective experiment and designing objective quality metrics is to convert the 3D dynamic mesh into 2D images or videos via projection and to collect subjective scores or calculate objective indexes based on these images or videos. In this paper, we study the influence of the view, or camera path selection for the projection, for both subjective and objective dynamic mesh quality assessment, and compare the performance of image-based metrics and point-based metrics corresponding to the collected subjective scores. First, we use the dynamic mesh sequences proposed by MPEG as anchors and generate videos corresponding to different coding configurations and different camera paths. Then, we conduct subjective experiments to collect the ground truth of mean opinion scores. Besides, we calculate the state-of-the-art objective metric scores for each sequence. We analyze the differences between subjective scores with respect to different camera paths and the correlation between subjective scores and objective metrics. The results show that different camera paths tend to generate close subjective perceptions and that the selection of views can influence some objective metrics.
Kaifa Yang, Qi Yang 0003, Joël Jung, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
ICME5
2023 Towards Deep Reference Frame in Versatile Video Coding NNVC
abstract
In this paper, we propose a deep reference frame generation method that aims to enhance bi-direction inter prediction under random access configuration in the latest video coding standard, Versatile Video Coding. Specifically, a pair of neighboring reconstructed frames are selected from decoded picture buffer and put into an optical-flow-based interpolation network to synthesize a new frame, similar to the current to-be-coded frame. Subsequently, this synthesized frame is incorporated into two-sided picture reference lists as additional reference frames. The proposed method is employed in both the encoding and decoding processes to eliminate bitstream signaling for supplementary information. The Small Ad-hoc Deep-Learning Library is utilized for implementing the proposed method. Experimental results demonstrate 3.67%/7.34%/6.51% coding efficiency improvements for Y/U/V components under the random access configuration when compared to the Versatile Video Coding NNVC reference software VTM-11_NNVC-5.0.
Weijie Bao, Jianghao Jia, Wenhui Meng, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
VCIP5
2023 Symmetric Geometry Coding for Static Meshes
abstract
Mesh compression plays an important role in the efficient storage and transmission of 3D models. While techniques like Video-based Dynamic Mesh Coding (V-DMC) mainly exploit the local structure of mesh for enhanced compression efficiency, they often overlook the global structure as well as a bias towards scanned 3D meshes. This work explores the unique characteristics of computer-generated (CG) meshes and presents a novel Symmetric Geometry Coding (SGC) tool. SGC takes advantage of the global symmetric property to encode just half of the symmetric mesh and predicts the other. We further extended SGC to accommodate partial symmetric meshes by processing submesh with multiple symmetries. Our results demonstrate that SGC achieves up to 40.90%/40.41% bitrate reduction compared to V-DMC 1.1 in D1/D2-PSNR metrics for symmetric CG meshes.
Thuong Nguyen-Canh, Fang-Yi Chao, Xiaozhong Xu, Shan Liu 0001
VCIP4
2023 Low-complexity Transform Network Architecture for JPEG AI Image Codec
abstract
Learning-based image coding schemes, exemplified by JPEG AI, have shown potential by greatly exceeding the conventional image compression standards in rate-distortion (RD) performance. However, their widespread applications are hindered by high decoding complexity, particularly from the upsampling and attention modules. Existing works sought to reduce this complexity, but their solutions are not fully effective, leaving considerable complexity unaddressed. In this paper, we present a simplified transform network architecture that employs an optimized attention module, a streamlined upsampling module, and a pared-down activation function to tackle this issue. Simulation results show that the simplified decoder sees its complexity (measured by kMACs/pixel) reduced by 80% (from 833 to 172), while the gain over Versatile Video Coding (VVC) increases slightly (from 27.3% to 27.5%). Partial methods in this paper have been integrated into the JPEG AI Verification Model (VM) software.
Ding Ding 0004, Xiaozhong Xu, Shan Liu 0001
VCIP4
2023 Immersive experience via 6DoF adaptive streaming
abstract
The highest immersive experience can be enjoyed by enabling 6DoF (6-degree of freedom) content delivery and playback. On the other hand, the ultra-high data rate associated with 6DoF content imposes quite a challenge. To achieve that, highly capable compression of such contents, adaptive and efficient streaming, as well as real-time rendering of such contents are required. In this demo session, we will introduce a 6DoF streaming system and player that can demonstrate true 6DoF immersion experience with smooth streaming and playback capability of high-resolution, dynamic 3D contents.
Xiaozhong Xu, Ningjun Dou, Guiquan Feng, Wen Gao 0001, Shan Liu 0001
VCIP1
2023 TSMD: A Database for Static Color Mesh Quality Assessment Study
abstract
Static meshes with texture map are widely used in modern industrial and manufacturing sectors, attracting considerable attention in the mesh compression community due to its huge amount of data. To facilitate the study of static mesh compression algorithm and objective quality metric, we create the Tencent – Static Mesh Dataset (TSMD) containing 42 reference meshes with rich visual characteristics. 210 distorted samples are generated by the lossy compression scheme developed for the Call for Proposals on polygonal static mesh coding, released on June 23 by the Alliance for Open Media Volumetric Visual Media group. Using processed video sequences, a large-scale, crowdsourcing-based, subjective experiment was conducted to collect subjective scores from 74 viewers. The dataset undergoes analysis to validate its sample diversity and Mean Opinion Scores (MOS) accuracy, establishing its heterogeneous nature and reliability. State-of-the-art objective metrics are evaluated on the new dataset. Pearson and Spearman correlations around 0.75 are reported, deviating from results typically observed on less heterogeneous datasets, demonstrating the need for further development of more robust metrics. The TSMD, including meshes, PVSs, bitstreams, and MOS, is made publicly available at the following location: https://multimedia.tencent.com/resources/tsmd.
Qi Yang 0003, Joël Jung, Haiqiang Wang, Xiaozhong Xu, Shan Liu 0001
VCIP4
2023 An Adaptive Predictive Tree for Point Cloud Geometry Compression
abstract
Point cloud has become widely applied in the real-time presentation of 3D objects and scenes. Efficient point cloud compression is a challenging task due to the massive number of points and non-uniform sampling structures. In this paper, we introduce an adaptive predictive tree for point cloud geometry compression. We provide a global-local adaption mechanism to add valid predictors to the list, largely reducing the searching complexity and ensuring the encoding efficiency for low-delay services. Joint optimization of point distance, prediction mode, and number of children nodes is employed in the predictor determination. We adjust the tree construction process from a comprehensive viewpoint, avoiding extreme outliers in the overall prediction. Compared with the latest geometry compression method, our approach can offer substantial coding efficiency improvement with notable encoding complexity reduction. The effectiveness of our approach is demonstrated using a variety of common test sequences employed by the standard committee.
Xiaozhong Xu, Wen Gao 0019, Shan Liu 0001
VCIP2
2023 Dynamic gaussian deep belief network design and stock market application
abstract
Stock price forecasting has been an important topic for investors, researchers, and analysts. In this paper, a prediction model of Dynamic Gaussian Deep Belief Network (DGDBN) is proposed. Generally, the network structure of traditional Deep Belief Network (DBN) determines the performance of its time series prediction. Most previous research uses artificial experience to adjust the network structure, it is difficult to ensure performance and time efficiency by constantly trying. In addition, the accuracy of the traditional DBN stacked by binary Restricted Boltzmann Machines(RBM) needs to be improved when solving the time series problem. The DGDBN designed in this paper contains two points: The first point is to add Gaussian noise to the RBM. The second point is to realize the increase or decrease branch algorithm of hidden layer structure according to the connection weights and average percentage error (MAPE). Finally, the forecast for the stocks of United Technologies Corporation and Unisys Corp, DGDBN is compared with DBN and LSTM. The root means square error (RMSE) increases by 15% and 65%. The interesting thing we found is that the number of neurons in the last layer of the DGDBN network has a greater effect than other layers.
Shuyue Xi, Xiaozhong Xu
Intell. Data Anal.2
2022 Rate Control for Learned Video Compression
abstract
Rate control is a critical part for video compression, especially in bandwidth-limited tasks such as live and broadcast. The newly-rising learned video compression has shown advantageous rate-distortion (RD) performance in previous research, but lack of rate control heavily limits its usage in real coding scenarios. In this work, we present the first rate control scheme tailored for learned video compression. Specifically, we explore the inter-frame dependency of learned video compression and propose a novel R-D-λ model accordingly for efficient rate allocation. Additionally, a staged update algorithm is developed for robust parameter estimation. Experiments on public datasets show that, the proposed rate control scheme achieves low rate error while maintaining equal or even higher RD performance, without introducing coding time overhead.
Yanghao Li, Jisheng Li, Jiangtao Wen, Yuxing Han 0001, Shan Liu 0001, Xiaozhong Xu
ICASSP7
2022 An Open Dataset for Video Coding for Machines Standardization
abstract
In recent years, video has dominated internet traffic and has become one of the major media formats. Besides its consumption by humans, videos are now often consumed by machines for analysis tasks such as object detection, segmentation, tracking, etc. Thus, efficient video coding for machines (VCM) becomes an important topic in academia and industry. Datasets are essential in the evaluation of a variety of coding tools for VCM. However, most of the publicly available datasets are only for academic usage, which prohibits their usage for many participants in the field of VCM study. In this paper, an open dataset with permissive license terms will be introduced. This dataset has been accepted by MPEG-VCM group as a test dataset. It is based on a larger video dataset, of which annotations for a subset of images are used to evaluate the performance of object detection and instance segmentation tasks. In addition, annotations for object tracking for a subset of videos are provided. The characteristics of the dataset and details of the annotations will be described. Evaluation results for multiple machine vision tasks will be presented to demonstrate the usage of the dataset.
Wen Gao 0017, Xiaozhong Xu, Matthew Qin, Shan Liu 0001
ICIP2
2022 Neural Network Based in-Loop Filter with Constrained Memory
abstract
In this paper, neural network based in-loop filters are designed to explore the potential performance improvement beyond the latest video coding standard Versatile Video Coding (VVC). Besides the performance, the model memory size is also taken into consideration. There are three filters which are designed with different sizes of memory requirement. Along with the un-filtered image, the prediction image and the partitioning image are also fed into the network. Additionally, the quantized parameter (QP) or slice type is fed into the network for classifying different situations when the model is shared by multiple sequence level base QPs or slice types. In the experiments, these filters are integrated into the latest reference software of VVC. An average of 12.26% YUV BD-rate saving is achieved for the filter with the largest memory size, and an average of 10.20% YUV BD-rate saving is achieved even if the required memory is reduced to 1/22.
Xiaozhong Xu, Shan Liu 0001
ICME2
2022 Learning-based Intra-Prediction For Point Cloud Attribute Transform Coding
abstract
The attached attributes of each point in point cloud are fairly valuable but aggravate the burden for storage and transmission. In this paper, we propose a learning-based intra-prediction method for region adaptive hierarchical transform (RAHT) to compress point cloud attributes efficiently. First, we design an adaptive neighbor selection (ANS) module to produce the most correlated neighbors for child nodes. Then, the correlated neighbors obtained by ANS, the corresponding distance weights, and some additional auxiliary information are concatenated and fed to a multi-layer perception (MLP) based network to estimate the child node attributes precisely. Besides, residual learning is introduced to accelerate the network convergence. The predicted child node attributes are finally transformed by RAHT, and the residual of transform coefficients are then quantized and entropy coded. Experimental results demonstrate that our proposed methods can significantly improve attribute coding efficiency with average 10.2% BD-Rate gains compared with MPEG G-PCC reference software TMC13v14.0 on MPEG PCC dataset.
Lizhi Hou, Linyao Gao, Yiling Xu, Zhu Li 0001, Xiaozhong Xu, Shan Liu 0001
MMSP5
2022 Boundary-Preserved Geometry Video for Dynamic Mesh Coding
abstract
In this paper, we present a Boundary-Preserved Geometry Video (BPGV) framework for dynamic mesh coding (DMC) with time-varying geometry, connectivity and attributes. The geometry video is generated by interpolating the 3D XYZ coordinates in the sampled 2D UV charts, and can be coded by any video codec to remove the spatial and temporal redundancies. However, the reconstruction from the geometry video itself may suffer from serious distortions because of the missing boundary information of the UV charts. Therefore it is proposed to code the boundary information of the UV charts in a separate sub-bitstream by efficient prediction and residual coding. The connectivity information can be inferred from the decoded geometry images and the boundary information via triangulation with linear complexity on the decoder side. Better coding performance can be achieved by leveraging the trade-offs between bitrate and quality from the proposed coding tools, including the adaptive chart sampling and the raw-chart coding mode. The proposed BPGV framework was submitted as a response to the MPEG’s CfP on DMC and the results demonstrated its superior performance compared to a state-of-the-art mesh codec.
Xiang Zhang 0004, Xiaozhong Xu, Shan Liu 0001
PCS4
2022 Substitutional Neural Image Compression
abstract
We describe Substitutional Neural Image Compression (SNIC), a general approach for enhancing any neural image compression model, that requires no data or additional tuning of the trained model. It boosts compression performance toward a flexible distortion metric and enables bit-rate control using a single model instance. The key idea is to replace the image to be compressed with a substitutional one that outperforms the original one in a desired way. Finding such a substitute is inherently difficult for conventional codecs, yet surprisingly favorable for neural compression models thanks to their fully differentiable structures. With gradients of a particular loss back-propogated to the input, a desired substitute can be efficiently crafted iteratively. We demonstrate the effectiveness of SNIC, when combined with various neural compression models and target metrics, in improving compression quality and performing bit-rate control measured by rate-distortion curves.
Xiao Wang 0028, Ding Ding 0004, Wei Jiang 0001, Wei Wang 0311, Xiaozhong Xu, Shan Liu 0001, Brian Kulis, Sang (Peter) Chin
PCS5
2022 Optimize neural network based in-loop filters through iterative training
abstract
The latest video coding standard, named Versatile Video Coding (VVC), had been finalized in 2020. In our previous work, several neural network based in-loop filters are proposed to improve the compression performance beyond VVC. However, the effect of inter-frame referencing mechanism is not taken into the consideration, which leads to an inconsistency between the training process and the final testing process. To solve this problem, an iterative training method is proposed in this paper to further optimize the neural network based in-loop filters. Based on the proposed method, up to 1.74% additional YUV BD-rate saving can be achieved. Compared with VVC, the experiments show that on average 14.00% YUV BD-rate saving is achieved by the filter with 22 models while 11.21% YUV BD-rate saving is achieved by the filter with one single model. Moreover, the subjective evaluation has verified that the performance of the single model filter is better than VVC by a noticeable margin.
Xiaozhong Xu, Shan Liu 0001
PCS2
2022 An Open Video Dataset For Screen Content Coding
abstract
In recent years, screen content video is becoming increasingly popular in several major video applications, such as video recording and video conferencing. Due to the unique features of screen content videos that are not captured by camera sensors but produced artificially, dedicated coding tools have been developed for achieving significant compression efficiency gain. In recognition of the popularity of screen content applications, an open video dataset for screen content is proposed in this paper for the development of screen content coding technologies. The proposed video dataset consists of 12 typical screen content type video clips that are publicly available. In addition, to better understand the characteristics of the proposed video dataset, several major screen content coding tools in AOMedia Video 1 (AV1) have been evaluated on this dataset and analyzed in this paper.
Yingbin Wang, Xin Zhao 0003, Xiaozhong Xu, Shan Liu 0001, Zhijun Lei, Mariana Afonso, Andrey Norkin, Thomas Daede
PCS3
2022 Deep Reference Frame Interpolation based Inter Prediction Enhancement for Versatile Video Coding
abstract
In video coding, bi-directional inter prediction aims to remove temporal redundancy via two-sided previously coded frames as reference. High-quality reference frames are essential to reduce the prediction residuals and improve coding efficiency performance. In this paper, we propose a deep learning-based reference frame interpolation method to enhance bi-prediction by introducing a synthetic frame to reference picture lists. Specifically, reconstructed frames are fed into a well-designed interpolation and filtering network to synthesize a picture which can be regarded as an additional reference of to-be-coded frame. Then, the picture is inserted at the appropriate place in reference picture lists to provide a more reliable reference for subsequent motion estimation and motion compensation. Experimental results show that the proposed method achieves 2.03%/6.96%/6.40% coding efficiency improvements for Y/U/V components under random access configuration, when compared with the VVC reference software VTM-15.0.
Jianghao Jia, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001, Zhenzhong Chen 0001
VCIP3
2022 Overview of Screen Content Coding in Recently Developed Video Coding Standards
abstract
In recent years, computer-generated texts, graphics, and animations have drawn more attention than ever. These types of media, also known as screen content, have become increasingly popular due to their widespread applications. To address the need for efficient coding of such content, several coding tools have been developed and have made great advances in terms of coding efficiency. The inclusion of screen content coding features in some recently developed video coding standards (namely, HEVC SCC, VVC, AVS3, AV1 and EVC) demonstrates the importance of supporting such features. This paper provides an overview and comparative study of screen content coding technologies, as well as discussions on the performance and complexity of the tools developed in these standards.
Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Enhanced Implicit Selection of Transform Skip in AVS3
abstract
As the demand of remote desktop sharing grows, Screen Content Coding (SCC) is rapidly drawing attention. Several efficient coding tools have been adopted into AVS3 for improving the performance of SCC. In this paper, enhanced implicit selection of transform skip (EISTS) is proposed to improve the performance of the transform module. When transform skip mode is introduced, its usage will be indicated in an implicit manner, instead of coding one explicit flag in the bitstream. More specifically, the transform type (whether it is transform skip or regular transform) can be derived by checking the parity of the number of even quantized coefficients at the decoder side. Correspondingly, one coefficient may need to be adjusted at the encoder side to match the assumed transform selection. Verified by experiments, EISTS can improve the performance of SCC efficiently. Therefore, this technology has been adopted into the AVS3 standard.
Xiaozhong Xu, Shan Liu 0001
ICME2
2021 A Real-Time H.266/VVC Software Decoder
abstract
The new Versatile Video Coding Standard (VVC) was finalized in July 2020. This new international standard is able to provide a higher compression efficiency by up to 50% for the same subjective quality compared to its predecessor HEVC but at the cost of an increased computational load. This paper investigates the complexity of VVC decoder processing blocks and presents a highly optimized decoder implementation that can achieve 4K 60fps VVC real-time decoding on an x86 based CPU using SIMD instruction extensions of the processor and additional parallel processing including data and task-level parallelism.
Shan Liu 0001, Hualong Jiao, Xiaozhong Xu, Xianguo Zhang, Chenchen Gu
ICME9
2021 Video Coding Tool Analysis and Dataset for Gaming Content
abstract
The gaming market has kept growing significantly in recent years. Driven by multiple technology advances, such as cloud computing and video technologies, new gaming applications, e.g. AR, VR and cloud gaming, are becoming more and more practical and popular. Among different types of gaming, the emergence of cloud gaming is driving the market with enhanced gamer experience as well as new challenges to the services. One of the key technological challenges of gaming applications is the video coding, which is the foundation of several popular gaming applications, including cloud gaming and game live streaming. Comparing to the typical camera captured content and screen content, gaming content presents unique features that directly lead to different preferences on the selection of coding tool sets. To better understand the behaviors of known video coding tools and provide test materials for research and development on future coding tools that benefits more on gaming content, in this paper, a dataset consists of a set of gaming video is proposed together with analysis of the performances of existing coding tools on these materials. It is observed that, several known coding tools are exceptionally beneficial for gaming content and the rational is analyzed in this paper.
Xin Zhao 0003, Shan Liu 0001, Xiang Li 0003, Guichun Li, Xiaozhong Xu
PCS5
2021 No-reference Quality Assessment of Panoramic Video based on Spherical-domain Features
abstract
As one of the most important parts of virtual reality applications, panoramic video has become very popular. Differing from the traditional plane video, panoramic video is projected onto the 2D plane for processing, while viewed in the spherical domain. Quality of the projected plane cannot represent the real visual quality in VR viewing. In this paper, a no-reference (NR) quality assessment method is proposed for panoramic video. The proposed method extracts spatial and temporal video features from spherical domain to better model the perceived visual quality, alleviating the influence of non-uniform projection. Experiments conducted on a subjective quality database for panoramic video show that the proposed method can achieve better performance compared with the NR method designed for traditional 2D video.
Yingxue Zhang 0004, Zizheng Liu, Zhenzhong Chen 0001, Xiaozhong Xu, Shan Liu 0001
PCS4
2021 A Video Dataset for Learning-based Visual Data Compression and Analysis
abstract
Learning-based visual data compression and analysis have attracted great interest from both academia and industry recently. More training as well as testing datasets, especially good quality video datasets are highly desirable for related research and standardization activities. A UHD video dataset, referred to as Tencent Video Dataset (TVD), is established to serve various purposes such as training neural network-based coding tools and testing machine vision tasks including object detection and segmentation. This dataset contains 86 video sequences with a variety of content coverage. Each video sequence consists of 65 frames at 4K (3840x2160) spatial resolution. In this paper, the details of this dataset, as well as its performance when compressed by VVC and HEVC video codecs, are introduced.
Xiaozhong Xu, Shan Liu 0001, Zeqiang Li
VCIP1
2021 Overview of the Screen Content Support in VVC: Applications, Coding Tools, and Performance
abstract
In an increasingly connected world, consumer video experiences have diversified away from traditional broadcast video into new applications with increased use of non-camera-captured content such as computer screen desktop recordings or animations created by computer rendering, collectively referred to as screen content. There has also been increased use of graphics and character content that is rendered and mixed or overlaid together with camera-generated content. The emerging Versatile Video Coding (VVC) standard, in its first version, addresses this market change by the specification of low-level coding tools suitable for screen content. This is in contrast to its predecessor, the High Efficiency Video Coding (HEVC) standard, where highly efficient screen content support is only available in extension profiles of its version 4. This paper describes the screen content support and the five main low-level screen content coding tools in VVC: transform skip residual coding (TSRC), block-based differential pulse-code modulation (BDPCM), intra block copy (IBC), adaptive color transform (ACT), and the palette mode. The specification of these coding tools in the first version of VVC enables the VVC reference software implementation (VTM) to achieve average bit-rate savings of about 41% to 61% relative to the HEVC test model (HM) reference software implementation using the Main 10 profile for 4:2:0 screen content test sequences. Compared to the HM using the Screen-Extended Main 10 profile and the same 4:2:0 test sequences, the VTM provides about 19% to 25% bit-rate savings. The same comparison with 4:4:4 test sequences revealed bit-rate savings of about 13% to 27% for$Y'C_{B}C_{R}$and of about 6% to 14% for$R'G'B'$screen content. Relative to the HM without the HEVC version 4 screen content coding extensions, the bit-rate savings for 4:4:4 test sequences are about 33% to 64% for$Y'C_{B}C_{R}$and 43% to 66% for$R'G'B'$screen content.
Tung Nguyen 0001, Xiaozhong Xu, Félix Henry, Ru-Ling Liao, Mohammed Golam Sarwer, Marta Karczewicz, Yung Hsuan Chao, Jizheng Xu, Shan Liu 0001, Detlev Marpe, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.2
2020 Screen Content Coding in Recently Developed Video Coding Standards
abstract
In the recently years, screen content video including computer generated text, graphics and animations, have drawn more attention than ever, as many related applications become very popular. However, conventional video codecs are typically designed to handle the camera-captured, natural video. Screen content video on the other hand, exhibits distinct signal characteristics and varied levels of the human's visual sensitivity to distortions. To address the need for efficient coding of such contents, a number of coding tools have been specifically developed and achieved great advances in terms of coding efficiency.The importance of screen content applications is well addressed by the fact that all of the recently developed video coding standards have included screen content coding (SCC) features. Nevertheless, the inclusion considerations of SCC tools in these standards are quite different. Each standard typically adopts only a subset of the known tools. Further, for one particular coding tool, when adopted in more than one standard, its technical features may various quite a lot from one standard to another.All these caused confusions to both researchers who want to further explore SCC on top of the state-of-the-art and engineers who want to choose a codec particularly suitable for their targeted products. Information of SCC technologies in general and specific tool designs in these standards are of great interest. This tutorial provides an overview and comparative study of screen content coding (SCC) technologies across a few recently developed video coding standards, namely HEVC SCC, VVC, AVS3, AV1 and EVC. In addition to the technical introduction, discussions on the performance and design/implementation complication aspects of the SCC tools are followed up, aiming to provide a detailed and comprehensive report. The overall performances of these standards are also compared in the context of SCC. The SCC tools in discussion are listed as follows:
Xiaozhong Xu, Shan Liu 0001
VCIP1
2020 An Optimized Video Encoder Implementation with Screen Content Coding Tools
abstract
Screen content video applications require efficient coding of computer-generated materials. The new screen content coding tools such as intra block copy (IBC) and palette mode (PLT) have addressed this requirement. However, the added computational complexity on top of the existing sophisticated video encoders is also challenging. In this paper, we focus on the fast and efficient encoder implementation of these screen content coding tools. Improvements on hash-based IBC search, PLT optimization, mode decision between PLT and intra mode, and other general encoder accelerations towards screen content applications are studied and discussed. Experimental results show that with these methods added, the encoder can achieve some faster runtime performance than before while the compression efficiency is almost doubled with screen content coding tools.
Xiaozhong Xu, Shitao Wang, Yiming Li 0001, Yushan Zheng, Shan Liu 0001
VCIP1
2019 Intra block copy in Versatile Video Coding with Reference Sample Memory Reuse
abstract
Screen contents such as online gaming streaming, remote desktop and WIFI display, become popular in current mainstream video applications. In versatile video coding (VVC), the most recent international video coding standard development, coding tools have been evaluated for optimizing screen content materials. Intra block copy (IBC) has shown its effectiveness in coding of computer-generated contents such as texts and graphics. Therefore, it has been previously included into the HEVC standard version 4, extensions for screen content coding (SCC). A constrained version of IBC mode has also been adopted in the VVC standard where the compensation range is limited within the current coding-tree unit (CTU), assuming a 1-CTU size of memory is allocated for storing IBC’s reference samples. In this paper, methods are proposed to efficiently utilize this reference sample memory for IBC mode such that effectively the search range for IBC mode can be increased without requiring more memory to store the reference samples. As a result, significant coding efficiency improvement over the traditional 1-CTU search range setting can be achieved. One of the proposed memory reuse strategies is considered practical for implementation and therefore has been included in the VVC standard.
Xiaozhong Xu, Xiang Li 0003, Shan Liu 0001
PCS1
2015 Block Vector Prediction for Intra Block Copying in HEVC Screen Content Coding
abstract
In screen content video, the spatial correlation among pixels shows different characteristics as compared to natural content video. Intra-picture motion compensation (or intra block copy in HEVC) plays a key role in reducing the bit rate of representing high resolution video with contents such as text and graphics. The efficiency of intra block copy is highly related to the accuracy of block vector prediction. In this paper, several block vector prediction methods are proposed to improve the performance of intra block copy technology in HEVC screen content coding extension. Simulation results show that an average bit rate reduction of 8.0% can be achieved for typical 1080p text and graphics sequences in all intra configuration, when compared to the standard's test model SCM-1.0. As a result, some of the proposed methods have been adopted into the standard working draft and reference software.
Xiaozhong Xu, Shan Liu 0001, Tzu-Der Chuang, Shawmin Lei
DCC1
2014 Screen content coding using non-square intra block copy for HEVC
abstract
To achieve high coding performance for screen content, the intra block copy (IntraBC) performs block matching within a limited area of the reconstructed samples inside the current picture. We further extend its notion from the unit of square coding unit (CU) to non-square prediction unit (PU) partitions. This design is then justified by theoretical and empirical analyses which reveal the same fact that blocks coded by IntraBC mode tend to enable more at smaller partition levels. Besides, the syntax design of the proposed method is fully aligned with that of inter partition modes. Therefore the architecture-wise change in video codec design can be minimized. The experimental results justify the effectiveness of the proposed mode for coding of screen content video. In particular, up to 19.5% rate reduction (with an average of 12.7%) relative to the HM-12.0+RExt-4.1 anchor can be achieved on top of the usage of CU-based IntraBC prediction.
Chun-Chi Chen, Xiaozhong Xu, Ru-Ling Liao, Wen-Hsiao Peng, Shan Liu 0001, Shawmin Lei
ICME2
2012 Two-way wireless video communication using Randomized cooperation, Network Coding and packet level FEC
abstract
Two-way real-time video communication in wireless networks requires high bandwidth, low delay and error resiliency. This paper addresses these demands by proposing a system with the integration of Network Coding (NC), user cooperation using Randomized Distributed Space-time Coding (R-DSTC) and packet level Forward Error Correction (FEC) under a one-way delay constraint. Simulation results show that the proposed scheme significantly outperforms both conventional direct transmission as well as R-DSTC based two-way cooperative transmission, and is most effective when the distance between the users is large.
Xiaozhong Xu, Özgü Alay, Elza Erkip, Yao Wang 0001, Shivendra S. Panwar
ICC1
2012 Predictive coding of intra prediction modes for high efficiency video coding
abstract
The High Efficiency Video Coding (HEVC) standardization process currently underway includes many tools for the coding of intra pictures. HEVC allows for many more intra prediction modes or directions as compared to previous standards. Efficient coding of these modes is therefore important because the modes consume a non-negligible portion of the total bit-stream used for coding intra pictures. In this paper, a predictive coding method is proposed to reduce the number of bits needed for signaling the intra prediction modes, where the spatial angular correlation between the intra prediction mode of the current Prediction Unit (PU) and the neighboring PUs is computed using a few modulo-N arithmetic operations that do not impact encoder or decoder run-times. The proposed method provides similar or greater improvements in compression efficiency as compared to competing tools, while requiring no changes to the existing bit-stream syntax.
Xiaozhong Xu, Robert A. Cohen, Anthony Vetro, Huifang Sun
PCS1
2010 Pattern-based Assembled DCT scheme with DC prediction and adaptive mode coding
abstract
A Pattern-based Assembled DCT (PADCT) scheme is proposed to fully utilize textual structure in images in a more flexible manner. Different patterns can be designed or extracted according to texture structures such as directionality and distribution patterns. Pixels in the block are divided into sub-partitions, followed by assembling pixels in each sub-partition together before the succeeding transformations. Compatible DC prediction and multiple mode coding schemes are also proposed. Compared with existed directional transform schemes, PADCT has the advantages of advanced compact energy concentration, efficient usage of memory, fewer categories of one dimensional transform, and adaptability for hardware optimization. Experimental results show that the bitrate reduction over the 2-D DCT is up to 40%. It also outperforms significantly than previous typical directional transform schemes in comparison with PSNR or subjective quality.
Zhibo Chen 0001, Xiaozhong Xu
ICIP2
2010 Pattern-based assembled DCT scheme for image coding
abstract
A Pattern-based Assembled DCT (PADCT) scheme is proposed as a general framework to fully utilize texture structure property in images in a more feasible and flexible manner. Different patterns can be designed or extracted according to texture structures such as directionality and distribution patterns. Pixels in the block are divided into sub-partitions, followed by assembling pixels in each sub-partitions together into a compact and efficient form for the succeeding separated one dimensional transform operations, suitable DC prediction and entropy coding schemes are also introduced. PADCT has the advantages of advanced compact energy concentration, efficient usage of memory, fewer categories of one dimensional transform, and adaptability for hardware optimization. Experimental results show that the bitrate reduction over the 2-D DCT is up to 35%. It also outperforms significantly than previous typical directional transform schemes in comparison with PSNR or subjective quality.
Zhibo Chen 0001, Xiaozhong Xu
VCIP2
2010 A refined motion estimation strategy for adaptive interpolation filter
Xiaozhong Xu, Zhongmou Wu
Signal Process. Image Commun.1
2008 Fast disparity motion estimation in MVC based on range prediction
abstract
In order to achieve a better coding performance in Multi- view video coding (MVC), the search range for both motion estimation (ME) and disparity estimation (DE) are set twice as large as necessary for ME alone, which brings great computational complexity to the encoder. Based on the behavior analysis of disparity estimation, a simple yet effective fast search algorithm, inter-view search range prediction (ISRP), is proposed in this paper. By applying this technique to full search or other fast search algorithms, a majority portion of the disparity search time could be removed while the coding performance is still maintained.
Xiaozhong Xu
ICIP1
2008 Improvements on Fast Motion Estimation Strategy for H.264/AVC
abstract
This paper analyzes the statistical characters of rate-distortion performances of the fast motion estimation (ME) algorithms in H.264/AVC and introduces an additional early termination scheme with adaptive thresholds for UMHS and CBFPS. This keeps the R-D performance of the original UMHS and CBFPS, while reducing the computational complexity of encoding process considerably. Simulation results in different conditions (image sizes from QCIF to HD and various quantization parameters) are shown. The proposed method is suitable on other video coding platforms and could be integrated with ease. Partial methods in this paper have been adopted into the JM software and the JVT test model.
Xiaozhong Xu
IEEE Trans. Circuits Syst. Video Technol.1