Hui Yuan 0001

dblp:21/780-1 · DBLP profile ↗
← Back
126ranked-venue papers
12as first author
96since 2021 · last 2026
0000-0001-5212-3393ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 106 · 9 first-author · 81 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Computer networks · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 An Information-Guided Learned Framework for Free-View Image Coding
abstract
In this paper, we propose an information-guided learned image compression framework, which achieves efficient free-view image compression. Specifically, we model the inter-view correlations as inter-view prior based on the proposed feature transform module, which is used to guide the encoding and decoding process of different views. Specifically, we design the feature transform module to generate the inter-view prior. Different from prior information modeled in pixel domain, the inter-view prior in feature domain has higher dimensions and can provide richer and more correlated condition information, which can guide the encoding and decoding process of different view images effectively. Furthermore, we design a multi-prior fusion module, which can fuse the information from different reference views to achieve more accurate guidance. We evaluate the proposed framework by comparing the rate-distortion performance averaged over the BlendedMVS dataset. Extensive experiments demonstrate that the proposed framework can achieve the superior performance in compressing free-view images.
Wenhong Duan, Hui Yuan 0001, Siwei Ma 0001
DCC3
2026 Feature Compression for Cloud-Edge Multimodal 3D Object Detection
abstract
Machine vision systems, which can efficiently manage extensive visual perception tasks, are becoming increasingly popular in industrial production and daily life. Due to the challenge of simultaneously obtaining accurate depth and texture information with a single sensor, multimodal data captured by cameras and LiDAR is commonly used to enhance performance. Additionally, cloud-edge cooperation has emerged as a novel computing approach to improve user experience and ensure data security in machine vision systems. This paper proposes a pioneering solution to address the feature compression problem in multimodal 3D object detection. Given a sparse tensor-based object detection network at the edge device, we introduce two modes to accommodate different application requirements: Transmission-Friendly Feature Compression (T-FFC) and Accuracy-Friendly Feature Compression (A-FFC). In T-FFC mode, only the output of the last layer of the network's backbone is transmitted from the edge device. The received feature is processed at the cloud device through a channel expansion module and two spatial upsampling modules to generate multi-scale features. In A-FFC mode, we expand upon the T-FFC mode by transmitting two additional types of features. These added features enable the cloud device to generate more accurate multi-scale features. Experimental results on the KITTI dataset using the VirConv-L detection network showed that T-FFC was able to compress the features by a factor of 4933 with less than a 3% reduction in detection performance. On the other hand, A-FFC compressed the features by a factor of about 733 with almost no degradation in detection performance. We also designed optional residual extraction and 3D object reconstruction modules to facilitate the reconstruction of detected objects. The reconstructed objects effectively reflected the shape, occlusion, and details of the original objects.
Chongzhen Tian, Hui Yuan 0001, Raouf Hamzaoui, Liquan Shen, Sam Kwong
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 SCVCD-5K: The First Large Scale Screen Content Video Dataset for Compression
Chaofei Li, Jinglan Tian, Hui Yuan 0001
IEEE Signal Process. Lett.5
2026 Deep Joint-Source Channel Coding-Based Receiver-Driven Replay Protocol for 3D Point Clouds
Huda Adam Sirag Mekki, Hui Yuan 0001, Mohanad M. G. Hassan, Zejia Chen
IEEE Signal Process. Lett.2
2026 Virtual Reference Frame-Based Inter Prediction for MPEG Enhanced G-PCC
abstract
As the demand for 3D point clouds grows, the data volume is growing dramatically. To tackle this challenge, the Moving Picture Expert Group (MPEG) is developing the enhanced geometry-based point cloud compression (Enhanced G-PCC) standard, which uses Region-Adaptive Hierarchical Transform (RAHT) for highly efficient attribute coding. However, since the geometry of the current frame and the reference frame is different, the octree structure between them does not match, which affects the performance of inter prediction. Therefore, we propose a virtual reference frame-based inter prediction method by aligning the geometry of the reference frame and the current frame. Specifically, the geometry of the virtual reference frame comes from the current frame, while its attribute information comes from the reference frame. Experimental results show that the proposed method can significantly increase the proportion of inter predicted RAHT coefficients and thus achieve average Bjøntegaard Delta Rates (BD-rates) of-6.3%,-8.9%, and-8.4% for the Luma, Cb, and Cr components, respectively, under the lossless geometry and lossy attribute coding condition, compared to the state-of-the-art Enhanced G-PCC reference software version 28 release candidate 2 (TMC13v28.0-rc2). For the coding condition of lossy geometry and lossy attribute, the corresponding BD-rates are-6.5%,-11.3%, and-7.7%, respectively.
Yuxuan Wei, Hui Yuan 0001
IEEE Signal Process. Lett.5
2026 A Noise Constrained Diffusion (NC-Diffusion) Framework for High-Fidelity Image Compression
abstract
With the great success of diffusion models in image generation, diffusion-based image compression is attracting increasing interests. However, due to the random noise introduced in the diffusion learning, they usually produce reconstructions with deviation from the original images, leading to suboptimal compression results. To address this problem, in this paper, we propose a Noise Constrained Diffusion (NC-Diffusion) framework for high fidelity image compression. Unlike existing diffusion-based compression methods that add random Gaussian noise and direct the noise into the image space, the proposed NC-Diffusion formulates the quantization noise originally added in the learned image compression as the noise in the forward process of diffusion. Then a noise constrained diffusion process is constructed from the ground-truth image to the initial compression result generated with quantization noise. The NC-Diffusion overcomes the problem of noise mismatch between compression and diffusion, significantly improving the inference efficiency. In addition, an adaptive frequency-domain filtering module is developed to enhance the skip connections in the U-Net based diffusion architecture, in order to enhance high-frequency details. Moreover, a zero-shot sample-guided enhancement method is designed to further improve the fidelity of the image. Experiments on multiple benchmark datasets demonstrate that our method can achieve the best performance compared with existing methods.
Yanbo Gao, Shuai Li 0005, Hui Yuan 0001, Mao Ye 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Dependability Feature Learning Based on Sample Generation for Unsupervised Text-to-Image Person Re-Identification
abstract
Text-to-image person re-identification (TIReID) aims to retrieve the target pedestrians according to specific textual descriptions. Benefiting from abundant annotated training data, current supervised TIReID methods have achieved impressive performance. However, annotating cross-modality data is extremely time-consuming, which limits their application in real-world scenarios. Several methods attempt to generate text descriptions or pseudo-labels but neglect the dependability of image-text matching relationships or identity information. To this end, we propose a Dependability Feature Learning based on Sample Generation (DFLSG) for unsupervised TIReID. First, we introduce a dependable text generation method that leverages multimodal large language models to generate diverse texts and further filtrate dependable texts for establishing image-text matching relationships. Second, we design an Error Sample Filtering Module (ESFM) to eliminate abnormal samples and obtain reliable identity labels. Furthermore, we develop a Multilevel Triplet Joint Learning (MTJL) process, which continuously optimizes the cross-modality dependable feature from center and instance views. Extensive experiments are implemented to assess the proposed DFLSG on four mainstream TIReID databases. Experimental results demonstrate that DFLSG achieves state-of-the-art performance compared with other unsupervised methods. Code will be available at: https://github.com/CLS-2001/DFLSG.
Chenglong Shao, Tongzhen Si, Hui Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Layer-Based Rate-Distortion Optimized Attribute Coding for Solid Geometry-Based Point Cloud Compression
abstract
In recent years, three-dimensional (3D) point clouds, which are applicable in various fields such as the metaverse and immersive communication, are attracting increasing attention. Under constrained storage and bandwidth conditions, efficient point cloud compression (PCC) plays a crucial role. To address these challenges, the Moving Picture Experts Group has been actively developing the geometry-based point cloud compression (G-PCC) standard and has recently proposed a test model for dynamic solid point clouds called Solid G-PCC. However, several issues still hinder the coding efficiency of attributes, such as inaccurate prediction, redundant coding bits, and accumulated distortion due to dependencies between frames. To tackle these challenges, we propose a layer-based rate-distortion optimized (RDO) attribute coding (L-RDOAC) method. This approach incorporates a layer-based RDO prediction (L-RDOP) to enhance prediction accuracy, a layer-based RDO quantization (L-RDOQ) to minimize redundant coding bits, and a layer-based RDO Wiener filter (L-RDOWF) to reduce distortion. Experimental results demonstrate that the coding efficiency of the proposed method significantly outperforms the state-of-the-art G-PCC reference software, as assessed through both objective and subjective evaluations. Specifically, compared to the state-of-the-art GeS-TM version 7.0, the proposed L-RDOAC achieves average Bjøntegaard-delta (BD) rates of -7.94%, -10.95%, and -8.16% for Luma, Cr, and Cb, respectively, under the C1 configuration (lossless geometry with lossy attributes), while under the C2 configuration (lossy geometry with lossy attributes), the average BD-rates are -7.88%, -8.09%, and -4.87%, respectively, when octree-based geometry coding is used.
Zexing Sun, Yuxuan Wei, Hao Liu 0044, Hui Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Perceptual Geometry Distortion Assessment of Compressed 3D Meshes
abstract
The evaluation of perceptual quality in 3D mesh compression, particularly for Video-based Dynamic Mesh Coding (V-DMC), is challenged by the scarcity of subject-rated datasets and the high computational cost of full-mesh decoding and the sophisticated visual feature extraction steps. To bridge this gap, we first introduce a novel V-DMC distortion dataset, comprising 16 high-quality original meshes and 400 compressed, textureless variants. We conducted a subjective quality assessment study with 30 participants using the Double Stimulus Impairment Scale (DSIS) method to collect reliable Mean Opinion Scores (MOS). We then propose streamMQ, the first-of-its-kind no-reference, bitstream-layer model for perceptual quality assessment of V-DMC compressed meshes. By extracting key geometric features such as quantization parameters and triangle count directly from the compressed bitstream, streamMQ predicts perceptual quality without full decoding. Experimental evaluation and comparison with state-of-the-art methods demonstrate that streamMQ achieves highly competitive quality assessment performance at tiny fractions of computational and storage costs, facilitating real-time and low-storage application environments. The dataset and source code will be made publicly available at https://github.com/HFL01/QDU-GDM.
Fanglin Hou, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Zhou Wang 0001
IEEE Trans. Image Process.4
2026 FD-SCU: Frequency Decomposition-Based Spectrum Collaborative Upsampling for Point Cloud Color Attribute
abstract
Existing point cloud color upsampling methods typically treat color upsampling as an interpolation problem within a local color or implicit feature domain. This largely overlooks the ability of the frequency domain to capture color correlations in local point sets. To address this limitation, we propose a spectrum collaborative strategy that uses frequency decomposition on voxel blocks (VBs) to enhance point cloud color reconstruction. We first voxelize the low-resolution (LR) color point cloud to generate multiple VBs and introduce a virtual filling strategy that adaptively assigns colors to empty voxels in each VB, ensuring that the irregularly distributed color information fully occupies the VB. We then apply the discrete cosine transform, known for its strong frequency-domain representation of locally smooth signals, to each color-filled VB to obtain frequency coefficients. These frequency coefficients are separated into high-frequency (HF) and low-frequency (LF) components. The LF coefficients, together with the LR color point cloud, are fed into a multi-scale cross-domain feature extraction module to capture deep features. Next, a Gaussian perturbation-based feature expansion generates upsampled color features, which are used to regress a coarse upsampled color point cloud. Finally, a high-frequency-guided residual refinement module uses the HF coefficients to refine the coarse upsampled result and produce a high-fidelity color point cloud. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art methods. Our code will be publicly available at https://github.com/wangwenchaoxx/FD-SCU.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan, Junhui Hou
IEEE Trans. Image Process.3
2026 LPCM: Learning-Based Predictive Coding for LiDAR Point Cloud Compression
abstract
In recent years, LiDAR point clouds have been widely used in many applications. Since the data volume of LiDAR point clouds is very huge, efficient compression is necessary to reduce their storage and transmission costs. However, existing learning-based compression methods do not exploit the inherent angular resolution of LiDAR and ignore the significant differences in the correlation of geometry information at different bitrates. The predictive geometry coding method in the geometry-based point cloud compression (G-PCC) standard uses the inherent angular resolution to predict the azimuth angles. However, it only models a simple linear relationship between the azimuth angles of neighboring points. Moreover, it does not optimize the quantization parameters for residuals on each coordinate axis in the spherical coordinate system. To address these issues, we propose a learning-based predictive coding method (LPCM) with both high-bitrate and low-bitrate coding modes. LPCM converts point clouds into predictive trees using the spherical coordinate system. In high-bitrate coding mode, we use a lightweight Long-Short-Term Memory-based predictive (LSTM-P) module that captures long-term geometry correlations between different coordinates to efficiently predict and compress the elevation angles. In low-bitrate coding mode, where geometry correlation degrades, we introduce a variational radius compression (VRC) module to directly compress the point radii. Then, we analyze why the quantization of spherical coordinates differs from that of Cartesian coordinates and propose a differential evolution (DE)-based quantization parameter selection method, which improves rate-distortion performance without increasing coding time. Experimental results show that LPCM achieved a D1-PSNR BD-rate reduction of 21.2% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 5.6% compared with the PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0.
Hui Yuan 0001, Shiqi Jiang 0006, Da Ai, Wei Zhang 0072, Raouf Hamzaoui
IEEE Trans. Image Process.2
2026 Inter-LPCM: Learning-Based Inter-Frame Predictive Coding for LiDAR Point Cloud Compression
abstract
Because LiDAR sensors acquire point clouds with a fixed angular resolution, the resulting data can be systematically parameterized and efficiently compressed in the spherical coordinate system. Traditional spherical coordinate-based point cloud compression methods have shown strong rate-distortion (RD) performance, with the predictive geometry coding (PredGeom) method in the geometry-based point cloud compression (G-PCC) standard being a prominent example. While PredGeom includes an inter-frame prediction mode, it relies on a simple linear model, which limits its ability to capture complex motion patterns or structural dependencies. On the other hand, existing learning-based compression methods in the spherical domain do not exploit inter-frame correlations to reduce geometry redundancy. To address these limitations, we propose a learning-based inter-frame predictive coding method (Inter-LPCM). For azimuth prediction, we use a delta coding strategy based on the predefined angular resolution. To improve compression for radii, we introduce an inter-frame radius predictive (Inter-RP) model that estimates the current point's radius using neighboring points from both the current frame and the registered reference frame. In addition, we design a lightweight attention-based prediction (LAEP) model to predict elevation angles by capturing long-range geometric correlations across different coordinates. For quantization, we propose an RD-optimized method to select the quantization steps in the spherical coordinate system. For entropy coding, we design distinct models for each spherical coordinate component. These models are adapted to the statistical priors of each coordinate, which enables more accurate probability estimation. Experimental results show that Inter-LPCM, in its best RD configuration, achieved a D1-PSNR BD-rate reduction of 26.1% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 8.3% compared with the inter-frame prediction mode of PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0. Our source code is publicly available at https://github.com/SDUChangSun/Inter-LPCM.
Hui Yuan 0001, Shiqi Jiang 0006, Chongzhen Tian, Raouf Hamzaoui
IEEE Trans. Image Process.2
2026 Character Transformation Artifact: Perception, Prediction, and Representative Application in Image Quality Assessment
abstract
Text screen content images (TSCIs) have been extensively applied in multimedia applications. When a TSCI is compressed by an encoder, the reconstructed image often exhibits changes of text stroke structures, leading to a novel and intriguing distortion, namely character transformation artifact (CTA). Specifically, CTA makes the original characters being transformed into other characters with varying degrees, causing inaccurate or even misleading perception of text semantic information, thereby affecting the quality of TSCIs. This paper systematically focuses on the quantitative representation of CTA for the first time, taking the advanced Versatile Video Coding (H.266/VVC) standard and English text as examples. First, the perceptual forms of CTA and the variation of CTA perceptual degree (CTA-PD) with the quantization parameter (QP) of H.266/VVC are explored. Second, a method for predicting the CTA-PD, referred to as P-CTA-PD, is formulated by extracting and quantifying the stroke features of distorted characters. Finally, P-CTA-PD is further applied to form a TSCI quality assessment method, namely CTA-based image quality assessment (CTA-IQA). Experimental results demonstrate that P-CTA-PD can achieve high accuracy in predicting perceptual degree of CTA, with an accuracy of 95.64%. Meanwhile, CTA-IQA scores show a stronger correlation with the mean opinion score compared to state-of-the-art quality assessment methods. This paper lays the groundwork for CTA, providing fundamental support for future exploration on CTA for other video coding standards, text types, and application scenarios.
Kaifang Yang, Yanchao Gong, Xuemin Chao, Qinqin Meng, Hui Yuan 0001
IEEE Trans. Image Process.5
2026 UGAE: Unified Geometry and Attribute Enhancement for G-PCC Compressed Point Clouds
abstract
Lossy compression of point clouds reduces storage and transmission costs; however, it inevitably leads to irreversible distortion in geometry structure and attribute information. To address these issues, we propose a unified geometry and attribute enhancement (UGAE) framework, which consists of three core components: post-geometry enhancement (PoGE), pre-attribute enhancement (PAE), and post-attribute enhancement (PoAE). In PoGE, a Transformer-based sparse convolutional U-Net is used to reconstruct the geometry structure with high precision by predicting voxel occupancy probabilities. Building on the refined geometry structure, PAE introduces an innovative enhanced geometry-guided recoloring strategy, which uses a detail-aware K-Nearest Neighbors (DA-KNN) method to achieve accurate recoloring and effectively preserve high-frequency details before attribute compression. Finally, at the decoder side, PoAE uses an attribute residual prediction network with a weighted mean squared error (W-MSE) loss to enhance the quality of high-frequency regions while maintaining the fidelity of low-frequency regions. UGAE significantly outperformed existing methods on three benchmark datasets: 8iVFB, Owlii, and MVUB. Compared to the latest G-PCC test model (TMC13v29), in terms of total bitrate setting, UGAE achieved an average BD-PSNR gain of 9.98 dB and -90.54% BD-bitrate for geometry under the D1 metric, as well as a 3.34 dB BD-PSNR improvement with -55.53% BD-bitrate for attributes. Additionally, it improved perceptual quality significantly. Our source code will be released on GitHub at: https://github.com/yuanhui0325/UGAE.
Hui Yuan 0001, Chongzhen Tian, Raouf Hamzaoui
IEEE Trans. Image Process.2
2026 PACE: A Multi-Round Cell-Enhanced Prefetching Strategy in Volumetric Video Streaming
abstract
In recent years, volumetric video streaming has emerged as a key application in the field of virtual reality (VR) and augmented reality (AR), attracting growing research and industry interest. Despite its potential, the extremely high bandwidth requirements of volumetric content far exceed the capacity of current networks to support full-resolution streaming. To address this limitation, industry practitioners often employ field-of-view (FoV) prediction to downscale the streaming content based on user gaze direction. Although this approach makes the streaming feasible, our empirical measurement reveals that, working with the FoV prediction, the commonly used sequential prefetching mechanism severely constrains streaming efficiency. To overcome this bottleneck, we introduce PACE, a smart prefetching strategy built on multi-round cell-level download scheduling. PACE divides the prefetching process of each group-of-frames (GoF) into several rounds, where distinct video cells are downloaded in each round according to periodically updated FoV predictions. This design decouples FoV prediction from the constraint of prefetch length, enabling more adaptive and responsive streaming. Moreover, PACE integrates a greedy quality selection policy that dynamically adjusts video quality to make full use of available bandwidth and enhance the overall quality of experience (QoE). Comprehensive evaluations show that PACE improves FoV prediction accuracy by up to 23.1% and boosts QoE by as much as 54.8%, while maintaining robust performance across diverse network and user scenarios.
Shuquan Liu, Mengbai Xiao, Hui Yuan 0001, Dongxiao Yu, Xiuzhen Cheng
IEEE Trans. Mob. Comput.4
2026 CWRNN-INVR: A Coupled WarpRNN Based Implicit Neural Video Representation
abstract
Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found athttps://github.com/yiyang-sdu/CWRNN-INVR.git.
Yanbo Gao, Shuai Li 0005, Jinglin Zhang 0001, Hui Yuan 0001, Mao Ye 0001, Xingyu Gao 0001
IEEE Trans. Multim.6
2026 Temporal Consistency-Aware Dynamic Point Clouds Color Attribute Enhancement
abstract
Dynamic point clouds, widely used in virtual reality and autonomous driving systems, often suffer from distortions due to quantization in the process of compression. These distortions significantly degrade the visual quality of dynamic point clouds, especially temporal inconsistency. To address this issue, a temporal consistency-aware dynamic point clouds color attribute enhancement method is proposed in this work. Specifically, a 3D Spatial-Temporal Search (STS) module is designed to adaptively search point cloud patches in the temporal domain for feature alignment. These matched patches are then individually fed into Single Frame Feature Extraction (SFFE) module that comprises of multi-head attention and graph convolution to exploit latent features of point cloud color attribute. In addition, to further capture both the spatial and temporal dependencies, a Convolutional Point cloud Long Short-Term Memory (Conv-PointLSTM) network is applied, which integrates convolution and max pooling with LSTM mechanism to facilitate the color attribute correspondents across the spatial-temporal latent features. Experimental results demonstrate that the proposed method can achieve 0.44 dB gains on average in terms of Peak Signal-to-Noise Ratio (PSNR) and 1.50%/5.31%bit rate reductions at the low/high bit rate, which outperforms the state-of-the-art works. The source code and trained models are available athttps://github.com/xu-coder-666/DPC.
Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Hui Yuan 0001, Sam Kwong
IEEE Trans. Multim.4
2026 Perceptual Quality Assessment of Trisoup-Lifting Encoded 3D Point Clouds
abstract
No-reference bitstream-layer point cloud quality assessment (PCQA) can be deployed without full decoding at any network node to achieve real-time quality monitoring. In this work, we develop the first PCQA model dedicated to Trisoup-Lifting encoded 3D point clouds by analyzing bitstreams without full decoding. Specifically, we investigate the relationship among texture bitrate per point (TBPP), texture complexity (TC) and texture quantization parameter (TQP) while geometry encoding is lossless. Subsequently, we estimate TC by utilizing TQP and TBPP. Then, we establish a texture distortion evaluation model based on TC, TBPP and TQP. Ultimately, by integrating this texture distortion model with a geometry attenuation factor, a function of trisoupNodeSizeLog2 (tNSL), we acquire a comprehensive NR bitstream-layer PCQA model named streamPCQ-TL. In addition, this work establishes a database named WPC6.0, the first PCQA database dedicated to Trisoup-Lifting encoding mode, encompassing 400 distorted point clouds with 4 geometry multiplied by 5 texture distortion levels. Experiment results on M-PCCD, ICIP2020 and the proposed WPC6.0 database suggest that the proposed streamPCQ-TL model exhibits robust and notable performance in contrast to existing advanced PCQA metrics, particularly in terms of computational cost.
Juncheng Long, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Wei Gao 0003, Jiarun Song, Zhou Wang 0001
IEEE Trans. Vis. Comput. Graph.4
2025 MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representation
abstract
This paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with linear indexing for continuous-space points can only provide degenerate linear latent space representations, and such representations cannot be adequately compensated to represent complex nonlinear signals by the following compact decoder. To address this problem while keeping the simplicity of a regular grid structure, our approach builds upon the standard grid-based paradigm by constructing multiple elementary metric grids as high-order terms to approximate complex nonlinearities, following the Taylor expansion principle. Furthermore, we enhance model compactness with hash encoding based on different sparsities of the grids to prevent detrimental hash collisions, and a high-order extrapolation decoder to reduce explicit grid storage requirements. experimental results on both 2D and 3D reconstructions demonstrate the superior fitting and rendering accuracy of the proposed method across diverse signal types, validating its robustness and generalizability. Code is available at https://github.com/wangshu31/MetricGrids.
Yanbo Gao, Shuai Li 0005, Chong Lv, Chuankun Li, Hui Yuan 0001, Jinglin Zhang 0001
CVPR7
2025 EEPNet: Efficient Edge Pixel-based Matching Network for Cross-Modal Dynamic Registration between LiDAR and Camera
abstract
Multisensor fusion is essential for autonomous vehicles to accurately perceive, analyze, and plan their trajectories within complex environments. This typically involves the integration of data from LiDAR sensors and cameras, which necessitates high-precision and real-time registration. Current methods for registering LiDAR point clouds with images face significant challenges due to inherent modality differences and computational overhead. To address these issues, we propose an efficient edge pixel-based matching network (EEPNet), an advanced network that leverages reflectance maps obtained from point cloud projections to enhance registration accuracy. The introduction of point cloud projections substantially mitigates cross-modality differences at the network input level, while the inclusion of reflectance data improves performance in scenarios with limited spatial information of point cloud within the camera’s field of view. Furthermore, by employing edge pixels for feature matching and incorporating an efficient matching optimization layer, EEPNet markedly accelerates real-time registration tasks. Experimental validation demonstrates that EEPNet achieves superior accuracy and efficiency compared to state-of-the-art methods. Our contributions offer significant advancements in autonomous perception systems, paving the way for robust and efficient sensor fusion in real-world applications.
Yuanchao Yue, Hui Yuan 0001, Shuai Li 0005
ISCAS2
2025 PCAC-GAN: A Sparse-Tensor-Based Generative Adversarial Network for 3D Point Cloud Attribute Compression
abstract
Learning-based methods have proven successful in compressing geometric information for point clouds. For attribute compression, however, they still lag behind non-learning-based methods such as the MPEG G-PCC standard. To bridge this gap, we propose a novel deep learning-based point cloud attribute compression method that uses a generative adversarial network (GAN) with sparse convolution layers. Our method also includes a module that adaptively selects the resolution of the voxels used to voxelize the input point cloud. Sparse vectors are used to represent the voxelized point cloud, and sparse convolutions process the sparse tensors, ensuring computational efficiency. To the best of our knowledge, this is the first application of GANs to compress point cloud attributes. Our experimental results show that our method outperforms existing learning-based techniques and rivals the latest G-PCC test model (TMC13v23) in terms of visual quality.
Xiaolong Mao, Hui Yuan 0001, Xin Lu 0001, Raouf Hamzaoui, Wei Gao 0003
Comput. Vis. Media2
2025 Adaptive local neighborhood search and dual attention convolution network for complex semantic segmentation towards indoor point clouds
Da Ai, Siyu Qin, Zihe Nie, Dianwei Wang, Hui Yuan 0001, Ying Liu 0026
Expert Syst. Appl.5
2025 DeepJSCC-PCG: Deep Joint Source Channel Coding for Point Cloud Geometry
abstract
Recently, three-dimensional (3D) point clouds have become more and more popular as they can represent 3D scenes and objects conveniently for applications like immersive communication and autonomous driving, etc. However, the huge data volume of 3D point clouds and the variable channel band width prevent its efficient and reliable transmission. We propose a deep joint source-channel coding method for point cloud geometry, namely DeepJSCC-PCG, by leveraging DeepJSCC and the octree-based adaptive voxelization to overcome the “cliff effect” caused by traditional separate source and channel coding (SSCC). Specifically, DeepJSCC-PCG first partitions the point cloud into non-overlapping voxel blocks via octree decomposition and prunes empty regions for efficient computation. Second, a 3D convolution-based neural network is specially designed to extract compact semantic features from the sparse data structure of the voxel blocks efficiently. Finally, the extracted features are transmitted through fading channels corrupted by Gaussian or Rayleigh noise, and subsequently decoded for reconstruction. The neural network is trained in an end-to-end fashion, allowing for the rebuilding of a high-quality point cloud. Experiments on the ModelNet40 and 8i datasets demonstrated that DeepJSCC CPCGsignificantly outperforms existing DeepJSCC methods and traditional SSCC methods.
Zejia Chen, Hui Yuan 0001, Huda Adam Sirag Mekki, Mohanad M. G. Hassan
IEEE Signal Process. Lett.2
2025 Perception-Weighted Multi-View Point Cloud Quality Assessment With Saliency-Guided Coverage Analysis
abstract
Due to the non-uniform perception of human vision, structural or color changes in salient regions play a dominant role in point cloud quality assessment (PCQA). In this paper, we propose a perception-weighted multi-view PCQA method based on saliency-guided coverage analysis (PW-SCQA), which dynamically quantifies the contribution of different viewpoints on perceptual quality through saliency. First, multi-view projection images are generated based on a polyhedral projection mechanism and multi-scale features are extracted for constructing a 2D saliency map. Then, the 2D to 3D saliency propagation model is used to refine the point-level saliency weights and achieve point cloud saliency visualization. Subsequently, a perception-driven viewpoint optimization mechanism and a novel viewpoint saliency region coverage (SRC) index are innovatively introduced, in which the viewpoint evaluation weights are dynamically adjusted by calculating the SRC of high, medium, and low saliency under the candidate viewpoints. Finally, the multi-view information content weighting image quality assessment method is combined to predict the overall point cloud quality. PW-SCQA outperforms several state-of-the-art methods on three different PCQA datasets.
Qi Liu 0029, Honglei Su, Hao Liu 0044, Hui Yuan 0001
IEEE Signal Process. Lett.5
2025 OMR-Net+: A Frequency-Aware Feature Refinement and Entropy Modeling Method for Efficient Screen Content Image Compression
abstract
Screen content image (SCI) compression faces challenges due to distinct characteristics such as sharp edges and repetitive structures. Existing learned image compression methods encounter two key issues: 1) insufficient frequency-aware processing, and 2) suboptimal entropy modeling for mixed-frequency components. To this end, we propose OMR-Net+, a novel SCI compression method that incorporates frequency-aware feature characteristics, including a frequency-aware refinement network (FARN) and a frequency-aware entropy model (FAEM). The proposed FARN uses an invertible neural network to preserve critical high-frequency details and a transformer-based model to reduce redundancy in low-frequency features. Additionally, the proposed FAEM provides tailored conditional probability estimation based on a parallel context model for high- and low-frequency features, respectively, to improve both coding performance and computational efficiency. Experimental results on the SCID and SIQAD datasets show that OMR-Net+ significantly outperforms the previous OMR-Net and other state-of-the-art methods in rate-distortion performance, demonstrating its potential for efficient SCI compression.
Shiqi Jiang 0006, Ting Ren, Hui Yuan 0001, Junyan Huo, Xin Lu 0001
IEEE Signal Process. Lett.3
2025 Chroma Subsampling for Enhanced Geometry-Based Point Cloud Compression
abstract
Due to the huge data volume of three dimensional point clouds, efficient point cloud compression (PCC) is very important and challenging under limited storage and bandwidth conditions. The Moving Picture Experts Group (MPEG) is actively developing the geometry-based point cloud compression (G-PCC) standard and plan to release the second edition of G-PCC, namely Enhanced G-PCC. In image and video compression, chroma components are typically encoded at a lower resolution than luma, with minimal perceptual quality loss. However, chroma subsampling has not yet been explored in PCC. We investigate the characteristics of points at different level of details, and propose a chroma subsampling that can be embedded with the codec of Enhanced G-PCC. Experimental results show that the proposed method outperforms the state-of-the-art Enhanced G-PCC reference software version29.0 in terms of coding efficiency and time complexity. Due to the excellent performance, the proposed method has been adopted by the MPEG and will be integrated into the upcoming version of the reference software of Enhanced G-PCC.
Yuxuan Wei, Jongseok Lee, Hyejung Hur, Hui Yuan 0001
IEEE Signal Process. Lett.5
2025 Temporal and Spatial Perception: A Novel Perceptual Rate-Distortion Optimization Method for H.266/VVC Encoding
abstract
Introducing saliency information to mitigate perceptual redundancy and achieve superior compression represents a novel approach to the development of video compression. Existing saliency-based compression coding methods rely on the determination of saliency regions and focus too much on saliency regions while ignoring the perceptible distortion in non-saliency regions. We propose a spatiotemporal visual perceptual rate-distortion optimization (PRDO) algorithm for Versatile Video Coding (H.266/VVC) that is more in line with the human visual system (HVS). Firstly, we establish a linear weighted distortion model based on spatiotemporal and saliency features. The distortion model makes effective use of saliency features while considering image content in non-saliency regions that is still perceptible to the human eye, thereby achieving an overall visual effect that conforms to human subjective perception. Based on this distortion model, we propose a saliency adaptive quantization parameter (SAQP) selection method with a more flexible quantization parameter selection range, adaptively allocating the optimal coding unit quantization parameter according to the saliency regions of the image, ensuring a balanced bitrate allocation between saliency and non-saliency regions. The proposed method is implemented for the first time on the H.266/VVC coding standard, attaining an average bitrate saving of 19.9% across all test sequences and an average PSNR improvement of 2.65 dB in saliency regions compared to VTM16.0. The BD-EWPSNR of the proposed PRDO and SAQP method improves by 1.34 dB and 1.45 dB in the All-Intra and Lowdelay_P encoding modes, respectively. Additionally, the BD-Rate based on EWPSNR is reduced by 25.86% and 33.73%, respectively, with an overall compression coding time saving of 19.76%. The experimental results demonstrate that the proposed method can significantly reduce the bit rate and coding time while improving the subjective perceptive quality, providing a competitive solution for video compression coding.
Da Ai, Hui Yuan 0001, Ying Liu 0026, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.4
2025 DQP-PCQA: Deep Quantization Parameters Bring New Insight to Point Cloud Quality Assessment
abstract
With the rapid development of immersive multimedia technology, the growing demand for high-quality visual experiences has driven the emergence of point cloud quality assessment (PCQA). While current deep learning-based PCQA models have achieved breakthroughs in performance, problems such as high computational complexity and limited model generalization ability still need to be solved. In this study, focusing on compression distortion, we analyzed and verified that the compression quantization parameter (QP) can be used as a key feature for predicting perceptual quality. Based on this, a novel no-reference point cloud perceptual quality assessment metric, DQP-PCQA, is proposed. Unlike existing PCQA models that only use mean opinion score (MOS) as a supervisory label, this study proposes a multi-objective constrained optimization scheme that adds geometric quantization parameter (GQP) and texture quantization parameter (TQP) as auxiliary supervisory labels to help the model can learn robust perceptual features that take into account both subjective quality and objective distortion. We conducted comparative experiments with other advanced PCQA models on several mainstream PCQA datasets. The results show that the DQP-PCQA model achieves fast convergence speed, excellent and stable performance, low complexity and strong generalization. Further migration experiments show that after applying our proposed method to other advanced PCQA models, the performance of the improved model is further improved. Our discovery provides new insight for PCQA research. To facilitate future reproducible research, the source code will be publicly released at https://github.com/Dds46/DQP-PCQA.
Dongshuai Duan, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Zhou Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Adaptive Depth-Converted-Scale Convolution for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object size and relationship among different objects, are the main clues to extract the scene structure. However, previous works lack the explicit handling of the changing sizes of the object due to the change of its depth. Especially in a monocular video, the size of the same object is continuously changed, resulting in size and depth ambiguity. To address this problem, we propose a Depth-converted-Scale Convolution (DcSConv) enhanced monocular depth estimation framework, by incorporating the prior relationship between the object depth and object scale to extract features from appropriate scales of the convolution receptive field. The proposed DcSConv focuses on the adaptive scale of the convolution filter instead of the local deformation of its shape. It establishes that the scale of the convolution filter matters no less (or even more in the evaluated task) than its local deformation. Moreover, a Depth-converted-Scale aware Fusion (DcS-F) is developed to adaptively fuse the DcSConv features and the conventional convolution features. Our DcSConv enhanced monocular depth estimation framework can be applied on top of existing CNN based methods as a plug-and-play module to enhance the conventional convolution block. Extensive experiments with different baselines have been conducted on the KITTI benchmark and our method achieves the best results with an improvement up to 11.6% in terms of SqRel reduction. Ablation study also validates the effectiveness of each proposed module.
Yanbo Gao, Huibin Bai, Huasong Zhou, Xingyu Gao 0001, Shuai Li 0005, Hui Yuan 0001, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Circuits Syst. Video Technol.7
2025 Adaptive Enhanced Global Intra Prediction for Efficient Video Coding in Beyond VVC
abstract
Global intra prediction (GIP), including intra-block copy and template matching prediction (TMP), exploits the global correlation of the same image to improve the coding efficiency. In Beyond VVC, TMP uses template matching to determine the reference blocks for efficient prediction. There usually exists an error between the coding block and reference blocks, caused by the content mismatch or the coding distortion of the reference blocks. We propose an enhancement over the reference blocks, namely enhanced GIP (EGIP). Specifically, we design an enhanced filter according to the templates of the coding block and the reference blocks, with the reconstructed template of the coding block as the label for supervised learning. To support different enhancements, we design two types of inputs, i.e., EGIP based on neighboring samples (N-EGIP) and EGIP based on multiple hypothesis references (M-EGIP). Experimental results show that, based on enhanced compression model (ECM) version 8.0, N-EGIP achieves BD-rate reductions of 0.37%, 0.42%, and 0.40%, and M-EGIP brings 0.34%, 0.37%, and 0.34% BD-rate savings for Y, Cb, and Cr components, respectively. A higher coding gain, 0.46%, 0.54%, and 0.52% BD-rate savings, can be achieved by integrating N-EGIP and M-EGIP together. Owing to the coding gain and small complexity increase, the proposed EGIP has been adopted in the exploration of Beyond VVC and integrated into its reference software.
Junyan Huo, Yanzhuo Ma, Zhenyao Zhang, Hui Yuan 0001, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 PU-GSM: A Latent Geometry-Guided Self-Similarity Model for Point Cloud Upsampling
abstract
Existing point cloud upsampling methods typically treat upsampling as a local interpolation problem, neglecting the importance of global correlations within point sets, which can limit their performance. To address this limitation, we exploit the inherent self-similarity of point clouds from a global perspective and propose PU-GSM, a latent geometry-guided self-similarity model for upsampling. We first generate a lower-resolution sparse sub-point cloud (SPC) by downsampling the input point cloud (IPC). Then, we introduce a latent geometry-guided self-similarity model (LGSM) that learns a point distribution on the underlying surface of SPC by exploiting the inherent self-similarity of IPC. Next, we reuse the LGSM for the remaining points (i.e., the points left after removing SPC from IPC). Afterward, we introduce a gradient-aware dual domain refiner to generate and calibrate the upsampled point cloud from the learned point distribution. Finally, we propose an inference-free latent vector matching approach to regularize the upsampled point cloud by enhancing the feature similarity between the upsampled point cloud and the ground truth in latent space. Extensive experiments show that PU-GSM achieves better upsampling results compared to state-of-the-art methods. Our code will be available at: https://github.com/liuhaoyun/PU-GSM.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan
IEEE Trans. Circuits Syst. Video Technol.2
2025 Rate-Distortion Optimized Skip Coding of Region Adaptive Hierarchical Transform Coefficients for MPEG G-PCC
abstract
Three-dimensional (3D) point clouds are becoming more and more popular for representing 3D objects and scenes. Due to limited network bandwidth, efficient compression of 3D point clouds is crucial. To tackle this challenge, the Moving Picture Experts Group (MPEG) is actively developing the Geometry-based Point Cloud Compression (G-PCC) standard, incorporating innovative methods to optimize compression, such as the Region-Adaptive Hierarchical Transform (RAHT) nestled within a layer-by-layer octree-tree structure. Nevertheless, a notable problem still exists in RAHT, i.e., the proportion of zero residuals in the last few RAHT layers leads to unnecessary bitrate consumption. To address this problem, we propose an adaptive skip coding method for RAHT, which adaptively determines whether to encode the residuals of the last several layers or not, thereby improving the coding efficiency. In addition, we propose a rate-distortion cost calculation method associated with an adaptive Lagrange multiplier. Experimental results demonstrate that the proposed method achieves average Bjøntegaard rate improvements of -3.50%, -5.56%, and -4.18% for the Luma, Cb, and Cr components, respectively, on dynamic point clouds, when compared with the state-of-the-art G-PCC reference software under the common test conditions recommended by MPEG.
Yuxuan Wei, Hui Yuan 0001, Wei Zhang 0072
IEEE Trans. Circuits Syst. Video Technol.3
2025 High Efficiency Wiener Filter-Based Point Cloud Quality Enhancement for MPEG G-PCC
abstract
Point clouds, which directly record the geometry and attributes of scenes or objects by a large number of points, are widely used in various applications such as virtual reality and immersive communication. However, due to the huge data volume and unstructured geometry, efficient compression of point clouds is very crucial. The Moving Picture Expert Group is establishing a geometry-based point cloud compression (G-PCC) standard for both static and dynamic point clouds in recent years. Although lossy compression of G-PCC can achieve a very high compression ratio, the reconstruction quality is relatively low, especially at low bitrates. To mitigate this problem, we propose a high efficiency Wiener filter that can be integrated into the encoder and decoder pipeline of G-PCC to improve the reconstruction quality as well as the rate-distortion performance for dynamic point clouds. Specifically, we first propose a basic Wiener filter, and then improve it by introducing coefficients inheritance and variance-based point classification for the Luma component. Besides, to reduce the complexity of the nearest neighbor search during the application of the Wiener filter, we also propose a Morton code-based fast nearest neighbor search algorithm for efficient calculation of filter coefficients. Experimental results demonstrate that the proposed method can achieve average Bjøntegaard delta rates of -6.1%, -7.3%, and -8.0% for Luma, Chroma Cb, and Chroma Cr components, respectively, under the condition of lossless-geometry-lossy-attributes configuration compared to the latest G-PCC encoding platform (i.e., geometry-based solid content test model version 7.0 release candidate 2) by consuming affordable computational complexity.
Yuxuan Wei, Hao Liu 0044, Liquan Shen, Hui Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Multi-Task Learning Model for V-PCC Geometry Compression Artifact Removal
abstract
In video-based point cloud compression (V-PCC), point clouds are projected as videos using a patch projection method and then compressed using video coding techniques. However, the lossy video compression and the down-sampling of occupancy maps (OMs) can lead to geometry compression artifacts, i.e., depth errors and OM errors, respectively. These errors can significantly affect the reconstruction quality of the point clouds. Existing methods can only eliminate one type of error and therefore have limited quality improvement. In this paper, to improve the quality maximally, a multi-task learning-based geometry compression artifact removal method is proposed to reduce both types of errors simultaneously. Considering the differences between the two tasks, the proposed method deals with the challenges of shared feature extraction and heterogeneous objective optimization. First, we propose a context-aware multi-task learning (CAML) model. The proposed CAML model can extract shared features that are context-aware and satisfy both tasks. Second, an improved optimization scheme is presented to train the proposed model. The improved optimization can fix the gradient imbalance of model updating. Cross-validation experiments show that the proposed method saves an average of over 45% Bjϕntegaard Delta bitrate in terms of the D2 metric.
Jian Xiong 0005, Jiucheng Xie, Hui Yuan 0001, Hao Gao 0005
IEEE Trans. Circuits Syst. Video Technol.5
2025 Approximately Invertible Neural Network for Learned Image Compression
abstract
Learned image compression has attracted considerable interests in recent years. An analysis transform and a synthesis transform, which can be regarded as coupled transforms, are used to encode an image to latent feature and decode the feature after quantization to reconstruct the image. Inspired by the success of invertible neural networks in generative modeling, invertible modules can be used to construct the coupled analysis and synthesis transforms. Considering the noise introduced in the feature quantization invalidates the invertible process, this paper proposes an Approximately Invertible Neural Network (A-INN) framework for learned image compression. It formulates the rate-distortion optimization in lossy image compression when using INN with quantization, which differentiates from using INN for generative modelling. Generally speaking, A-INN can be used as the theoretical foundation for any INN based lossy compression method. Based on this formulation, A-INN with a progressive denoising module (PDM) is developed to effectively reduce the quantization noise in the decoding. Moreover, a Cascaded Feature Recovery Module (CFRM) is designed to learn high-dimensional feature recovery from low-dimensional ones to further reduce the noise in feature channel compression. In addition, a Frequency-enhanced Decomposition and Synthesis Module (FDSM) is developed by explicitly enhancing the high-frequency components in an image to address the loss of high-frequency information inherent in neural network based image compression, thereby enhancing the reconstructed image quality. Extensive experiments demonstrate that the proposed A-INN framework achieves better or comparable compression efficiency than the conventional image compression approach and state-of-the-art learned image compression methods.
Yanbo Gao, Shuai Li 0005, Chong Lv, Hui Yuan 0001, Mao Ye 0001
IEEE Trans. Image Process.7
2025 PCE-GAN: A Generative Adversarial Network for Point Cloud Attribute Quality Enhancement Based on Optimal Transport
abstract
Point cloud compression significantly reduces data volume but sacrifices reconstruction quality, highlighting the need for advanced quality enhancement techniques. Most existing approaches focus primarily on point-to-point fidelity, often neglecting the importance of perceptual quality as interpreted by the human visual system. To address this issue, we propose a generative adversarial network for point cloud quality enhancement (PCE-GAN), grounded in optimal transport theory, with the goal of simultaneously optimizing both data fidelity and perceptual quality. The generator consists of a local feature extraction (LFE) unit, a global spatial correlation (GSC) unit and a feature squeeze unit. The LFE unit uses dynamic graph construction and a graph attention mechanism to efficiently extract local features, placing greater emphasis on points with severe distortion. The GSC unit uses the geometry information of neighboring patches to construct an extended local neighborhood and introduces a transformer-style structure to capture long-range global correlations. The discriminator computes the deviation between the probability distributions of the enhanced point cloud and the original point cloud, guiding the generator to achieve high quality reconstruction. Experimental results show that the proposed method achieves state-of-the-art performance. Specifically, when applying PCE-GAN to the latest geometry-based point cloud compression (G-PCC) test model, it achieves an average BD-rate of -19.2% compared with the PredLift coding configuration and -18.3% compared with the RAHT coding configuration. Subjective comparisons show a significant improvement in texture clarity and color transitions, revealing finer details and more natural color gradients.
Hui Yuan 0001, Qi Liu 0029, Honglei Su, Raouf Hamzaoui, Sam Kwong
IEEE Trans. Image Process.2
2025 SPAC: Sampling-Based Progressive Attribute Compression for Dense Point Clouds
abstract
We propose an end-to-end attribute compression method for dense point clouds. The proposed method combines a frequency sampling module, an adaptive scale feature extraction module with geometry assistance, and a global hyperprior entropy model. The frequency sampling module uses a Hamming window and the Fast Fourier Transform to extract high-frequency components of the point cloud. The difference between the original point cloud and the sampled point cloud is divided into multiple sub-point clouds. These sub-point clouds are then partitioned using an octree, providing a structured input for feature extraction. The feature extraction module integrates adaptive convolutional layers and uses offset-attention to capture both local and global features. Then, a geometry-assisted attribute feature refinement module is used to refine the extracted attribute features. Finally, a global hyperprior model is introduced for entropy encoding. This model propagates hyperprior parameters from the deepest (base) layer to the other layers, further enhancing the encoding efficiency. At the decoder, a mirrored network is used to progressively restore features and reconstruct the color attribute through transposed convolutional layers. The proposed method encodes base layer information at a low bitrate and progressively adds enhancement layer information to improve reconstruction accuracy. Compared to the best anchor of the latest geometry-based point cloud compression (G-PCC) standard that was proposed by the Moving Picture Experts Group (MPEG), the proposed method can achieve an average Bjøntegaard delta bitrate of -24.58% for the Y component (resp. -21.23% for YUV components) on the MPEG Category Solid dataset and -22.48% for the Y component (resp. -17.19% for YUV components) on the MPEG Category Dense dataset. This is the first instance that a learning-based attribute codec outperforms the G-PCC standard on these datasets by following the common test conditions specified by MPEG. Our source code will be made publicly available on https://github.com/sduxlmao/SPAC.
Xiaolong Mao, Hui Yuan 0001, Shiqi Jiang 0006, Raouf Hamzaoui, Sam Kwong
IEEE Trans. Image Process.2
2025 Energy-Adaptive Bitstream-Layer Model for Perceptual Quality Assessment of V-PCC Encoded 3D Point Clouds
abstract
The scope of point cloud (PC) applications is expanding. We propose a no-reference bitstream-layer quality assessment model that eliminates the need for full decoding of the PC, providing quality evaluation scores during the V-PCC decoding process. Specifically, we illustrate the relationship between content diversity (CD) and perceptual coding distortion in lossless geometric coding. Subsequently, we model attribute distortion by predicting CD using transform energy (TE) and texture quantization parameter (TQP). By combining the geometric distortion model with geometry quantization parameters (GQP) and the attribute distortion model, we derive comprehensive quality prediction results. Our experimental results on four PC databases (WPC2.0, M-PCCD, VSENSE VVDB and VSENSE VVDB2) show that the proposed energy-adaptive bitstream-layer model (EABL) delivers competitive quality prediction performance in comparison with existing full-reference, reduced-reference and no-reference PC quality assessment models that require full decoding, and meanwhile exhibits large speed advantage. The source code will be made publicly available for repeatability research at https://github.com/arthas-sws/EABL_model.
Wusi Sang, Honglei Su, Qi Liu 0029, Hui Yuan 0001, Zhou Wang 0001
IEEE Trans. Image Process.4
2025 A Novel Spatial-Temporal Learning Method for Enhancing Generalization in Adaptive Video Streaming
abstract
Adaptive video streaming has become a fundamental technology for video delivery. With the rise of deep reinforcement learning (DRL), streaming vendors are increasingly adopting DRL-driven adaptive bitrate (ABR) algorithms. In real-world deployments, most ABR approaches are developed with the aim of maintaining good performance across a wide variety of network environments. However, contrary to this expectation, our empirical findings show that even when trained on extensive real-world network trace data, these DRL-based ABR algorithms achieve only 43.1% to 48.9% of Quality-of-Experience (QoE) under highly diverse network conditions, which falls significantly short of the 100% optimum. We termed this problem as “ABR Under-Generalization”. To overcome this problem, we introduce BETA – a novel DRL-based ABR framework that incorporates both spatial and temporal learning mechanisms: 1) Spatially, BETA features a detector that flags the network conditions likely to cause poor performance, then trains specialized ABR models tailored for those conditions; 2) Temporally, BETA enhances its learning by incorporating multi-step decision experiences at each training epoch, enabling the trained model to account for long-term environmental dynamics. Comprehensive evaluations show that BETA outperforms state-of-the-art ABR algorithms, yielding average QoE gains of 19.4% to 50.9%, and achieving improvements of up to 244.1% under severely fluctuating network conditions.
Huaren Wei, Mengbai Xiao, Hui Yuan 0001, Dongxiao Yu, Xiuzhen Cheng
IEEE Trans. Mob. Comput.5
2025 Global Spatial-Temporal Information-Based Residual ConvLSTM for Video Space-Time Super-Resolution
abstract
By converting low-frame-rate, low-resolution videos into high-frame-rate, high-resolution ones, space-time video super-resolution techniques can enhance visual experiences and facilitate more efficient information dissemination. We propose a convolutional neural network (CNN) for space-time video super-resolution, namely GIRNet. Our method combines long-term global information and short-term local information from the video to better extract complete and accurate spatial-temporal information. To generate highly accurate features and thus improve performance, the proposed network integrates a feature-level temporal interpolation module with deformable convolutions and a global spatial-temporal information-based residual convolutional long short-term memory (convLSTM) module. In the feature-level temporal interpolation module, we leverage deformable convolution, which adapts to deformations and scale variations of objects across different scene locations. This provides a more efficient solution than conventional convolution for extracting features from moving objects. Our network effectively uses forward and backward feature information to determine inter-frame offsets, leading to the direct generation of interpolated frame features. In the global spatial-temporal information-based residual convLSTM module, the first convLSTM is used to derive global spatial-temporal information from the input features, and the second convLSTM uses the previously computed global spatial-temporal information feature as its initial cell state. This second convLSTM adopts residual connections to preserve spatial information, thereby enhancing the output features. Experiments on the Vimeo90 K dataset show that the proposed method outperforms open source state-of-the-art techniques in peak signal-to-noise-ratio (by 1.45 dB, 1.14 dB, and 0.2 dB over STARnet, TMNet, and 3DAttGAN, respectively), structural similarity index(by 0.027, 0.023, and 0.006 over STARnet, TMNet, and 3DAttGAN, respectively), and visual quality.
Congrui Fu, Hui Yuan 0001, Shiqi Jiang 0006, Liquan Shen, Raouf Hamzaoui
IEEE Trans. Multim.2
2025 LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace Representation
abstract
Monocular depth estimation (MDE) has attracted increasing interest in the past few years, owing to its important role in 3D vision. MDE is the estimation of a depth map from a monocular image/video to represent the 3D structure of a scene, which is a highly ill-posed problem. To solve this problem, in this paper, we propose a LiftFormer based on lifting theory topology, for constructing an intermediate subspace that bridges the image color features and depth values, and a subspace that enhances the depth prediction around edges. MDE is formulated by transforming the depth value prediction problem into depth-oriented geometric representation (DGR) subspace feature representation, thus bridging the learning from color values to geometric depth values. A DGR subspace is constructed based on frame theory by using linearly dependent vectors in accordance with depth bins to provide a redundant and robust representation. The image spatial features are transformed into the DGR subspace, where these features correspond directly to the depth values. Moreover, considering that edges usually present sharp changes in a depth map and tend to be erroneously predicted, an edge-aware representation (ER) subspace is constructed, where depth features are transformed and further used to enhance the local features around edges. The experimental results demonstrate that our LiftFormer achieves state-of-the-art performance on widely used datasets, and an ablation study validates the effectiveness of both proposed lifting modules in our LiftFormer.
Shuai Li 0005, Huibin Bai, Yanbo Gao, Chong Lv, Hui Yuan 0001, Chuankun Li, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Multim.5
2025 EdgeRegNet: Edge Feature-Based Multimodal Registration Network Between Images and LiDAR Point Clouds
abstract
Cross-modal data registration has long been a critical task in computer vision, with extensive applications in autonomous driving and robotics. Accurate and robust registration methods are essential for aligning data from different modalities, forming the foundation for multimodal sensor data fusion and enhancing perception systems' accuracy and reliability. The registration task between 2D images captured by cameras and 3D point clouds captured by Light Detection and Ranging (LiDAR) sensors is usually treated as a visual pose estimation problem. High-dimensional feature similarities from different modalities are leveraged to identify pixel-point correspondences, followed by pose estimation techniques using least squares methods. However, existing approaches often resort to downsampling the original point cloud and image data due to computational constraints, inevitably leading to a loss in precision. Additionally, high-dimensional features extracted using different feature extractors from various modalities require specific techniques to mitigate cross-modal differences for effective matching. To address these challenges, we propose a method that uses edge information from the original point clouds and images for cross-modal registration. We retain crucial information from the original data by extracting edge points and pixels, enhancing registration accuracy while maintaining computational efficiency. The use of edge points and edge pixels allows us to introduce an attention-based feature exchange block to eliminate cross-modal disparities. Furthermore, we incorporate an optimal matching layer to improve correspondence identification. We validate the accuracy of our method on the KITTI and nuScenes datasets, demonstrating its state-of-the-art performance. Our code is publicly available on GitHub athttps://github.com/ESRSchao/EdgeRegNet.
Yuanchao Yue, Hui Yuan 0001, Qinglong Miao, Xiaolong Mao, Raouf Hamzaoui, Peter Eisert
IEEE Trans. Multim.2
2025 Enabling the Awareness of Video Perceived Quality for Short-Form Video Streaming
abstract
In recent years, fueled by the rapid advances in high-speed mobile networks, streaming short-form videos over mobile devices (e.g., TikTok) has become ubiquitous among mobile users. Despite the widespread application, our investigation based on a real video data source revealed that a large proportion of short videos watched by viewers have suboptimal video quality (e.g., with low VMAF scores), which indicates that the Quality-of-Experience (QoE) is in fact far from optimal. This problem is primarily due to the lack of awareness of video quality optimization based on the features of the video content such as the scene complexity. To tackle this problem, this work develops a novel system called Quality Aware Short Video Streaming (QASVS), which adopts machine learning techniques to learn the video content features and then automatically generate quality-driven bitrate decision models to optimize the perceived video quality and QoE. Extensive evaluations show that QASVS is able to improve the video quality by 11.1%∼27.9% while significantly reducing the playback rebuffering compared to the state-of-the-art streaming algorithms. Therefore, QASVS is able to provide an effective way for streaming vendors to deliver high-performance short-video services.
Mengbai Xiao, Hui Yuan 0001, Dongxiao Yu, Xiuzhen Cheng
IEEE Trans. Serv. Comput.4
2025 CS-Net: Contribution-Based Sampling Network for Point Cloud Simplification
abstract
Point cloud sampling plays a crucial role in reducing computation costs and storage requirements for various vision tasks. Traditional sampling methods, such as farthest point sampling, lack task-specific information and, as a result, cannot guarantee optimal performance in specific applications. Learning-based methods train a network to sample the point cloud for the targeted downstream task. However, they do not guarantee that the sampled points are the most relevant ones. Moreover, they may result in duplicate sampled points, which requires completion of the sampled point cloud through post-processing techniques. To address these limitations, we propose a contribution-based sampling network (CS-Net), where the sampling operation is formulated as a Top-$k$k operation. To ensure that the network can be trained in an end-to-end way using gradient descent algorithms, we use a differentiable approximation to the Top-$k$k operation via entropy regularization of an optimal transport problem. Our network consists of a feature embedding module, a cascade attention module, and a contribution scoring module. The feature embedding module includes a specifically designed spatial pooling layer to reduce parameters while preserving important features. The cascade attention module combines the outputs of three skip connected offset attention layers to emphasize the attractive features and suppress less important ones. The contribution scoring module generates a contribution score for each point and guides the sampling process to prioritize the most important ones. Experiments on the ModelNet40 and PU147 showed that CS-Net achieved state-of-the-art performance in two semantic-based downstream tasks (classification and registration) and two reconstruction-based tasks (compression and surface reconstruction). CS-Net also achieved high average precision for objection detection on the KITTI LiDAR point cloud dataset, demonstrating its effectiveness in three-dimensional object detection.
Chen Chen 0063, Hui Yuan 0001, Xiaolong Mao, Raouf Hamzaoui, Junhui Hou
IEEE Trans. Vis. Comput. Graph.3
2025 No-Reference Bitstream-Based Perceptual Quality Assessment of Octree-Lifting Encoded 3D Point Clouds
abstract
No-reference point cloud quality assessment (PCQA) based on bitstreams uses information extracted from the bitstream for quality monitoring at network nodes. We develop a no-reference PCQA model based on bitstreams for the perceived quality assessment of Octree-Lifting coded point clouds. At first, our research explores the essential correlation between subjective visual quality degradation and the texture quantization parameter (TQP) when using lossless geometric coding. Then, we enhance the proposed model by incorporating texture complexity (TC) while taking into account the dependence of perceptual coding distortion on the texture characteristics of a point cloud. We estimate TC by utilizing TQP and calculating the average standard deviation of the Y-component of the attribute value ($Y\_ {std}$Y_std), both of which are extracted from the bitstream. Then, a texture distortion assessment model is constructed based on TQP and $Y\_ {std}$Y_std. The integration of the texture distortion model with the position quantization scale (PQS) results in the derivation of an overall no-reference bitstream-based PCQA model, named streamPCQ-OL. The findings from the conducted experiments highlight a significant superiority of the proposed model over existing approaches in terms of performance.
Jianyu Lv, Honglei Su, Qi Liu 0029, Hui Yuan 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Progressive Knowledge Transfer Network Based on Human Visual Perception Mechanism for No-Reference Point Cloud Quality Assessment
abstract
Point cloud perceptual quality assessment plays a critical role in many applications, including compression and communication. We propose PKT-PCQA, a point-based no-reference point cloud quality assessment deep learning network that emulates the human visual system by using progressive knowledge transfer to convert coarse-grained quality classification knowledge into a fine-grained quality prediction task. PKT-PCQA exploits local and global features, as well as an attention mechanism based on spatial and channel attention modules. Experiments on three large and independent point cloud assessment datasets show that PKT-PCQA outperforms existing no-reference and reduced-reference point cloud quality assessment methods and achieves better or similar performance compared to several State-of-the-Art full-reference methods.
Honglei Su, Qi Liu 0029, Hui Yuan 0001, Raouf Hamzaoui
IEEE Trans. Vis. Comput. Graph.4
2024 A Transformer-Based Intra Luma Enhancement for H.266/VVC
abstract
Intra prediction is essential in reducing spatial domain correlation in video coding. To improve intra prediction accuracy, we introduce a transformer-based quality enhancement method aiming atimproving the luma quality of reconstructed coding tree units (CTUs). Our transformer-based model, namely Enhanceformer, utilizes multi-head attention for comprehensive feature extraction across multiple stages, levels, and scales. By integrating the model into H.266/VVC codec, it not only improves the luma quality of the current reconstructed CTU, but also provides more accurate references for intra prediction of subsequent CTUs. Experimental results show that this method achieves average BD rate savings of 2.63%, 0.21% and 0.48% for Y, Cb and Cr components respectively in all intra configuration, outperforming H.266/Versatile Video Coding (VVC) anchor.
Wenrui Lv, Hui Yuan 0001, Congrui Fu, Shiqi Jiang 0006, Junyan Huo
PCS2
2024 MGTN: Multi-scale Graph Transformer Network for 3D Point Cloud Semantic Segmentation
abstract
The structural similarity of point clouds presents challenges in accurately recognizing and segmenting semantic information at the demarcation points of complex scenes or objects. In this study, we propose a multi-scale graph transformer network (MGTN) for 3D point cloud semantic segmentation. First, a multi-scale graph convolution (MSG-Conv) is devised to address the limitations faced by existing methods when extracting local and global features of point cloud data with varying densities simultaneously. Subsequently, we employ a graph-transformer (G-T) module to enhance edge details and spatial position information in the point cloud, thereby improving recognition accuracy for small objects and confusing elements such as columns and beams. Extensive testing on ShapeNet parts and S3DIS datasets was conducted to demonstrate the effectiveness of MGTN. Compared to the baseline network DGCNN, our proposed MGTN achieves substantial performance improvements, as evidenced by notable increases in mIoU of 1.5% and 18.5% on the ShapeNet parts and S3DIS datasets respectively. Additionally, MGTN outperforms the recent CFSA- Net by 2.3% and 3.4% on OA and mIoU respectively.
Da Ai, Siyu Qin, Zihe Nie, Hui Yuan 0001, Ying Liu 0026
VCIP4
2024 RPRA: Reputation-based prioritization and resource allocation leveraging predictive analytics and vehicular fog computing
Muhammad Ilyas Khattak, Hui Yuan 0001, Ayaz Ahmad, Ajmal Khan, Inamullah
Ad Hoc Networks2
2024 OMR-NET: A Two-Stage Octave Multi-Scale Residual Network for Screen Content Image Compression
abstract
Screen content (SC) differs from natural scene (NS) with unique characteristics such as noise-free, repetitive patterns, and high contrast. Aiming at addressing the inadequacies of current learned image compression (LIC) methods for SC, we propose an improved two-stage octave convolutional residual blocks (IToRB) for high and low-frequency feature extraction and a cascaded two-stage multi-scale residual blocks (CTMSRB) for improved multi-scale learning and nonlinearity in SC. Additionally, we employ a window-based attention module (WAM) to capture pixel correlations, especially for high contrast regions in the image. We also construct a diverse SC image compression dataset (SDU-SCICD2K) for training, including text, charts, graphics, animation, movie, game and mixture of SC images and NS images. Experimental results show our method, more suited for SC than NS data, outperforms existing LIC methods in rate-distortion performance on SC images.
Shiqi Jiang 0006, Ting Ren, Congrui Fu, Shuai Li 0005, Hui Yuan 0001
IEEE Signal Process. Lett.5
2024 Aligned Intra Prediction and Hyper Scale Decoder Under Multistage Context Model for JPEG AI
abstract
Learning-based image compression has raised increasing interests in the last few years. Currently, Joint Photographic Experts Group (JPEG) is working on the standardization of learning-based image compression as JPEG AI. It adopts a deep neural network based encoder-decoder architecture with hyperprior based probability formulation for entropy coding. JPEG AI currently contains two coding profiles, including the Base Operating Point (BaseOP) and High Operating Point (HighOP). Among the various techniques developed in JPEG AI, Multistage Context Model (MCM) was adopted as the context model to perform intra prediction in HighOP. It transforms the spatially progressive context prediction into sub-image feature prediction among channels via feature down-shuffling. However, in this prediction process, sub-image features are not spatially aligned to each other, and directly using the neighboring sub-image features cannot provide accurate prediction. Moreover, the distributions of residual features generated by MCM are also not consistent with that of the hyper scale decoder, which is used to construct the probability model in the entropy coding of residual features, leading to suboptimal residual coding. To address the above problems, we propose an Aligned Intra Prediction (AIP) and Aligned Hyper Scale Decoder (AHSD) under Multistage Context Model for JPEG AI coding. AIP aligns the reference sub-image features to the to-be-predicted feature in MCM with an offset prediction network and deformable convolution. AHSD further generates hyper scale features with matched distributions to the residual features, in order to enhance the probability formulation in its entropy coding. Experimental results demonstrate that the proposed method improves the coding performance by 1.3% in terms of BD-rate saving over the JPEG AI reference software and the effectiveness of each module is verified in ablation study.
Shuai Li 0005, Yanbo Gao, Chuankun Li, Hui Yuan 0001
IEEE Signal Process. Lett.4
2024 Enhancing Octree-Based Context Models for Point Cloud Geometry Compression With Attention-Based Child Node Number Prediction
abstract
In point cloud geometry compression, most octree-based context models use the cross-entropy between the one-hot encoding of node occupancy and the probability distribution predicted by the context model as the loss. This approach converts the problem of predicting the number (a regression problem) and the position (a classification problem) of occupied child nodes into a 255-dimensional classification problem. As a result, it fails to accurately measure the difference between the one-hot encoding and the predicted probability distribution. We first analyze why the cross-entropy loss function fails to accurately measure the difference between the one-hot encoding and the predicted probability distribution. Then, we propose an attention-based child node number prediction (ACNP) module to enhance the context models. The proposed module can predict the number of occupied child nodes and map it into an 8-dimensional vector to assist the context model in predicting the probability distribution of the occupancy of the current node for efficient entropy coding. Experimental results demonstrate that the proposed module enhances the coding efficiency of octree-based context models.
Hui Yuan 0001, Xiaolong Mao, Xin Lu 0001, Raouf Hamzaoui
IEEE Signal Process. Lett.2
2024 PU-Mask: 3D Point Cloud Upsampling via an Implicit Virtual Mask
abstract
We present PU-Mask, a virtual mask-based network for 3D point cloud upsampling. Unlike existing upsampling methods, which treat point cloud upsampling as an “unconstrained generative” problem, we propose to address it from the perspective of “local filling”, i.e., we assume that the sparse input point cloud (i.e., the unmasked point set) is obtained by locally masking the original dense point cloud with virtual masks. Therefore, given the unmasked point set and virtual masks, our goal is to fill the point set hidden by the virtual masks. Specifically, because the masks do not actually exist, we first locate and form each virtual mask by a virtual mask generation module. Then, we propose a mask-guided transformer-style asymmetric auto-encoder (MTAA) to restore the upsampled features. Moreover, we introduce a second-order unfolding attention mechanism to enhance the interaction between the feature channels of MTAA. Next, we generate a coarse upsampled point cloud using a pooling technique that is specific to the virtual masks. Finally, we design a learnable pseudo Laplacian operator to calibrate the coarse upsampled point cloud and generate a refined upsampled point cloud. Extensive experiments demonstrate that PU-Mask is superior to the state-of-the-art methods. Our code will be made available at: https://github.com/liuhaoyun/PU-Mask.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Qi Liu 0029, Shuai Li 0005
IEEE Trans. Circuits Syst. Video Technol.2
2024 Enlarged Motion-Aware and Frequency-Aware Network for Compressed Video Artifact Reduction
abstract
Making full use of spatial-temporal information is the key factor for removing compressed video artifacts. Recently, many deep learning-based compression artifact reduction methods have emerged. Among them, a series of methods based on deformable convolution have shown excellent capabilities in spatio-temporal feature extraction. However, local deformable offset prediction and pixel-wise inter-frame feature alignment in the unidirectional form limit the full utilization of temporal features in the existing method. Additionally, compressed video shows inconsistent degrees of distortion on different frequency components, and their restoration difficulty is also nonuniform. For the above problems presented by existing methods, we propose anenlarged motion-aware and frequency-aware network(EMAFA) to further extract spatio-temporal information and enhance information of different frequency components. To perceive different degrees of motion artifacts between compressed frames as accurately as possible, we design a bidirectional dense propagation pattern withpixel-wise and patch-wise deformable convolution(PIPA) module in the feature domain. In addition, we propose amulti-scale atrous deformable alignment(MSADA) module to enrich spatio-temporal features in image domain. Moreover, we design amulti-direction frequency enhancement(MDFE) module with multiple direction convolution to enhance the features of different frequency components. The experimental results show that the proposed method performs better than the state-of-the-art methods in both objective evaluation and visual perception experience. Supplementary experiments for Internet Streamed Video with hybrid-distortion demonstrate that our method also exhibits considerable generalizability for quality enhancement.
Wang Liu 0001, Wei Gao 0003, Ge Li 0002, Siwei Ma 0001, Tiesong Zhao, Hui Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Dependence-Based Coarse-to-Fine Approach for Reducing Distortion Accumulation in G-PCC Attribute Compression
abstract
Geometry-based point cloud compression (G-PCC) is a state-of-the-art point cloud compression standard. While G-PCC achieves excellent performance, its reliance on the predicting transform leads to a significant dependence problem, which can easily result in distortion accumulation. This not only increases bitrate consumption but also degrades reconstruction quality. To address these challenges, we propose a dependence-based coarse-to-fine approach for distortion accumulation in G-PCC attribute compression. Our method consists of three modules: level-based adaptive quantization, point-based adaptive quantization, and Wiener filter-based refinement level quality enhancement. The level-based adaptive quantization module addresses the interlevel-of-detail (LOD) dependence problem, while the point-based adaptive quantization module tackles the interpoint dependence problem. On the other hand, the Wiener filter-based refinement level quality enhancement module enhances the reconstruction quality of each point based on the dependence order among LODs. Extensive experimental results demonstrate the effectiveness of the proposed method. Notably, when the proposed method was implemented in the latest G-PCC test model (TMC13v23.0), a Bj$\phi$ntegaard delta rate of$-$4.9%,$-$12.7%, and$-$14.0% was achieved for the Luma, Chroma Cb, and Chroma Cr components, respectively.
Hui Yuan 0001, Raouf Hamzaoui
IEEE Trans. Ind. Informatics2
2024 Low Complexity Coding Unit Decision for Video-Based Point Cloud Compression
abstract
With growing demand for point cloud coding, Video-based Point Cloud Compression (V-PCC) is released for dynamic point clouds, relying on mature 2D video coding techniques. However, the huge computational complexity of 2D video codec is inherited by V-PCC, thereby resulting in a notably time-consuming encoding process for the projection videos. To accelerate the compression, this paper proposes a low complexity coding unit decision algorithm for V-PCC intra coding. First, the 2D sequences (occupancy, geometry, and attribute sequences) are projected from same 3D point could frames in V-PCC. By exploring the strong correlations among them, the cross-projection information is creatively proposed for improving the predication performance of CU partition. Second, considering the disparate coding losses generated by incorrect partitioning decisions of different CUs, we develop a rate-distortion-oriented learning approach aimed at increasing the decision accuracy of the CUs, severely affecting coding performance. Third, to accommodate the particular coding architecture of V-PCC intra configuration, we further devise an overall framework, including targeted feature extraction and partitioning decision for intra and inter coding of geometry and attribute sequences. The final experimental results strongly demonstrate the effectiveness of our proposed algorithm. The time consumption of total projection sequence compression can be reduced by 57.80%, while the coding losses on Geom.BD-TotalRate (D1 and D2) and Attr.BD-TotalRate Luma component are only 0.08%, 0.33%, and 0.16%, respectively, which can be negligible. To the best of our knowledge, the proposed algorithm achieves state-of-the-art performance for accelerating the projection sequence compression in V-PCC All-Intra configuration.
Wei Gao 0003, Ge Li 0002, Zhu Li 0001, Hui Yuan 0001
IEEE Trans. Image Process.5
2024 Support Vector Regression-Based Reduced- Reference Perceptual Quality Model for Compressed Point Clouds
abstract
Video-based point cloud compression (V-PCC) is a state-of-the-art moving picture experts group (MPEG) standard for point cloud compression. V-PCC can be used to compress both static and dynamic point clouds in a lossless, near lossless, or lossy way. Many objective quality metrics have been proposed for distorted point clouds. Most of these metrics are full-reference metrics that require both the original point cloud and the distorted one. However, in some real-time applications, the original point cloud is not available, and no-reference or reduced-reference quality metrics are needed. Three main challenges in the design of a reduced-reference quality metric are how to build a set of features that characterize the visual quality of the distorted point cloud, how to select the most effective features from this set, and how to map the selected features to a perceptual quality score. We address the first challenge by proposing a comprehensive set of features consisting of compression, geometry, normal, curvature, and luminance features. To deal with the second challenge, we use the least absolute shrinkage and selection operator (LASSO) method, which is a variable selection method for regression problems. Finally, we map the selected features to the mean opinion score in a nonlinear space. Although we have used only 19 features in our current implementation, our metric is flexible enough to allow any number of features, including future more effective ones. Experimental results on the Waterloo point cloud dataset version 2 (WPC2.0) and the MPEG point cloud compression dataset (M-PCCD) show that our method, namely PCQAML, outperforms state-of-the-art full-reference and reduced-reference quality metrics in terms of Pearson linear correlation coefficient, Spearman rank order correlation coefficient, Kendall's rank-order correlation coefficient, and root mean squared error.
Honglei Su, Qi Liu 0029, Hui Yuan 0001, Qiang Shawn Cheng, Raouf Hamzaoui
IEEE Trans. Multim.3
2024 UIERL: Internal-External Representation Learning Network for Underwater Image Enhancement
abstract
Underwater image enhancement (UIE) is a meaningful but challenging task, and many learning-based UIE methods have been proposed in recent years. Although much progress has been made, these methods still have two issues: (1) There exists a significant region-wise quality difference in a single underwater image due to the underwater imaging process, especially in regions with different scene depths. However, existing methods neglect this internal characteristic of underwater images, resulting in inferior performance; (2) Due to the uniqueness of the acquisition approach, underwater image acquisition tools usually capture multiple images in the same or similar scenes. Thus, the underwater images to be enhanced in practical usage are highly correlated. However, when processing a single image, existing methods do not consider the rich external information provided by the related images. There is still room for improvement in their performance. Motivated by these two aspects, we propose a novel internal-external representation learning (UIERL) network to better perform UIE tasks with internal and external information, simultaneously. In the internal representation learning stage, a new depth-based region feature guidance network is designed, including a region segmentation module based on scene depth to sense regions with different quality levels, followed by a region-wise space encoder module. With performing region-wise feature learning for regions with different quality separately, the network provides an effective guidance for global features and thus guides intra-image differentiated enhancement. In the external representation learning stage, we first propose an external information extraction network to mine the rich external information in the related images. Then, internal and external features interact with each other via the proposed external-assist-internal module (external features are updated with the help of internal features) and internal-assist-external module (internal features are updated with the help of external features). In this way, our UIERL fully explores the rich internal and external information to better enhance a single image. All results show that our method can achieve state-of-the-art performance on five benchmarks.
Zhengyong Wang, Liquan Shen, Yihan Yu, Hui Yuan 0001
IEEE Trans. Multim.4
2024 3DTA: No-Reference 3D Point Cloud Quality Assessment With Twin Attention
abstract
Point clouds are rapidly gaining popularity in many practical applications, and point cloud quality assessment (PCQA) is an important research topic that helps us measure and improve the visual experience in applications using point clouds. Research on full-reference (FR) PCQAs has recently made impressive progress, and research on no-reference (NR) PCQAs has also gradually increased. However, the performance of the prior NR PCQA methods still suffers from weak generalization ability and lower accuracy than the FR metrics in general. In this work, we propose a two-stage sampling method that can reasonably represent a whole point cloud, making it possible to efficiently calculate the point cloud quality. For quality prediction, we designed a twin-attention-based transformer PCQA model (3DTA), which uses the data of the two-stage sampling method as input and directly outputs the predicted quality score. Our model is accurate and widely applicable, and it has a simple and flexible structure. Experimental results show that in most cases, the proposed 3DTA model substantially outperforms the benchmark NR methods. The accuracy of the proposed method is competitive even against that of the FR method, which makes 3DTA a strong candidate for the PCQA task, regardless of the reference availability. The code of the proposed model is publicly available athttps://github.com/philox12358/3DTA-PCQA.
Linxia Zhu, Xu Wang 0006, Honglei Su, Huan Yang 0001, Hui Yuan 0001, Jari Korhonen
IEEE Trans. Multim.6
2023 Fourier Series and Laplacian Noise-Based Quantization Error Compensation for End-to-End Learning-Based Image Compression
abstract
Quantization is a core operation in lossy image compression. In the end-to-end learning-based image compression framework, quantization is conducted by a rounding operation during test, while it is replaced by additive uniform noise during training, leading to a mismatched problem between train and test. To address this problem, we propose a quantization error compensation method for the end-to-end learning-based image compression framework. The method uses Fourier series to approximate the periodic changes of the quantization error, and adds Laplacian noise to the quantized latent during test. The proposed method can be flexibly combined with different end-to-end learning-based image compression methods. Experimental results show that higher coding efficiency can be achieved by adding the proposed method with the state-of-the-art methods.
Shiqi Jiang 0006, Hui Yuan 0001, Shuai Li 0005, Xiaolong Mao
ICIP2
2023 CAS-Net: Cascade Attention-Based Sampling Neural Network for Point Cloud Simplification
abstract
Point cloud sampling can reduce storage requirements and computation costs for various vision tasks. Traditional sampling methods, such as farthest point sampling, are not geared towards downstream tasks and may fail on such tasks. In this paper, we propose a cascade attention-based sampling network (CAS-Net), which is end-to-end trainable. Specifically, we propose an attention-based sampling module (ASM) to capture the semantic features and preserve the geometry of the original point cloud. Experimental results on the ModelNet40 dataset show that CAS-Net outperforms state-of-the-art methods in a sampling-based point cloud classification task, while preserving the geometric structure of the sampled point cloud.
Chen Chen 0063, Hui Yuan 0001, Hao Liu 0044, Junhui Hou, Raouf Hamzaoui
ICME2
2023 Global Structure-Aware Diffusion Process for Low-light Image Enhancement
abstract
This paper studies a diffusion-based framework to address the low-light image enhancement problem. To harness the capabilities of diffusion models, we delve into this intricate process and advocate for the regularization of its inherent ODE-trajectory. To be specific, inspired by the recent research that low curvature ODE-trajectory results in a stable and effective diffusion process, we formulate a curvature regularization term anchored in the intrinsic non-local structures of image data, i.e., global structure-aware regularization, which gradually facilitates the preservation of complicated details and the augmentation of contrast during the diffusion process. This incorporation mitigates the adverse effects of noise and artifacts resulting from the diffusion process, leading to a more precise and flexible enhancement. To additionally promote learning in challenging regions, we introduce an uncertainty-guided regularization technique, which wisely relaxes constraints on the most extreme regions of the image. Experimental evaluations reveal that the proposed diffusion-based framework, complemented by rank-informed regularization, attains distinguished performance in low-light enhancement. The outcomes indicate substantial advancements in image quality, noise suppression, and contrast amplification in comparison with state-of-the-art methods. We believe this innovative approach will stimulate further exploration and advancement in low-light image processing, with potential implications for other applications of diffusion models. The code is publicly available at https://github.com/jinnh/GSAD.
Jinhui Hou, Junhui Hou, Hui Liu 0032, Huanqiang Zeng, Hui Yuan 0001
NeurIPS6
2023 A novel context inconsistency elimination algorithm based on the optimized Dempster-Shafer evidence theory for context-awareness systems
Qiang Liu 0052, Hongji Xu, Hui Yuan 0001, Zhi Liu 0004, Shidi Fan, Tiankuo Li
Appl. Intell.4
2023 Cuboid-Net: A multi-branch convolutional neural network for joint space-time video super resolution
abstract
Abstract The demand for high‐resolution videos has been consistently rising across various domains, propelled by continuous advancements in societal. Nonetheless, limitations in imaging and economic factors often result in obtaining low‐resolution images. The currently available space‐time video super‐resolution methods often fail to fully exploit the information existing within the spatio‐temporal domain. To address this problem, the issue is tackled by conceptualizing the input low‐resolution video as a cuboid structure. An innovative methodology called “Cuboid‐Net”, which incorporates a multi‐branch convolutional neural network, is introduced. Cuboid‐Net is designed to collectively enhance the spatial and temporal resolutions of videos, enabling the extraction of rich and meaningful information across both spatial and temporal dimensions. Specifically, the input video is taken as a cuboid to generate different directional slices as input for different branches of the network. The proposed network contains four modules, that is, a multi‐branch‐based hybrid feature extraction module, a multi‐branch‐based reconstruction module, a first‐stage quality enhancement module, and a second‐stage cross frame quality enhancement module for interpolated frames only. Experimental results demonstrate that the proposed method is not only effective for spatial and temporal super‐resolution of video but also for spatial and angular super‐resolution of light field.
Congrui Fu, Hui Yuan 0001, Hongji Xu, Hao Zhang 0211, Liquan Shen
IET Image Process.2
2023 3D pedestrian localization fusing via monocular camera
Jiande Sun 0001, Shanxin Zhang, Hui Yuan 0001, Huaxiang Zhang 0001, Jia Zhang 0028
J. Vis. Commun. Image Represent.4
2023 TMSO-Net: Texture adaptive multi-scale observation for light field image depth estimation
Congrui Fu, Hui Yuan 0001, Hongji Xu, Hao Zhang 0211, Liquan Shen
J. Vis. Commun. Image Represent.2
2023 Adaptive Chroma Prediction Based on Luma Difference for H.266/VVC
abstract
Cross-component chroma prediction plays an important role in improving coding efficiency for H.266/VVC. We use the differences between reference samples and the predicted sample to design an attention model for chroma prediction, namely luma difference-based chroma prediction (LDCP). Specifically, the luma differences (LDs) between reference samples and the predicted sample are employed as the input of the attention model, which is designed as a softmax function to map LDs to chroma weights nonlinearly. Finally, a weighted chroma prediction is conducted based on the weights and chroma reference samples. To provide adaptive weights, the model parameter of the softmax function can be determined based on the template (T-LDCP) or offline learning (L-LDCP), respectively. Experimental results show that the T-LDCP achieves BD-rate reductions of 0.34%, 2.02%, and 2.34% for the Y, Cb, and Cr components, and the L-LDCP brings 0.32%, 2.06%, and 2.21% BD-rate savings for Y, Cb, and Cr components, respectively. The L-LDCP introduces slight encoding and decoding time increments, i.e., 2% and 1%, when integrated into the latest VVC test model version 18.0. Besides, the LDCP can be implemented by a pixel-level parallelization which is hardware-friendly.
Junyan Huo, Danni Wang, Hui Yuan 0001, Shuai Wan, Fuzheng Yang 0001
IEEE Trans. Image Process.3
2023 Bitstream-Based Perceptual Quality Assessment of Compressed 3D Point Clouds
abstract
With the increasing demand of compressing and streaming 3D point clouds under constrained bandwidth, it has become ever more important to accurately and efficiently determine the quality of compressed point clouds, so as to assess and optimize the quality-of-experience (QoE) of end users. Here we make one of the first attempts developing a bitstream-based no-reference (NR) model for perceptual quality assessment of point clouds without resorting to full decoding of the compressed data stream. Specifically, we first establish a relationship between texture complexity and the bitrate and texture quantization parameters based on an empirical rate-distortion model. We then construct a texture distortion assessment model upon texture complexity and quantization parameters. By combining this texture distortion model with a geometric distortion model derived from Trisoup geometry encoding parameters, we obtain an overall bitstream-based NR point cloud quality model named streamPCQ. Experimental results show that the proposed streamPCQ model demonstrates highly competitive performance when compared with existing classic full-reference (FR) and reduced-reference (RR) point cloud quality assessment methods with a fraction of computational cost.
Honglei Su, Qi Liu 0029, Hui Yuan 0001, Huan Yang 0001, Zhenkuan Pan 0001, Zhou Wang 0001
IEEE Trans. Image Process.4
2023 GQE-Net: A Graph-Based Quality Enhancement Network for Point Cloud Color Attribute
abstract
In recent years, point clouds have become increasingly popular for representing three-dimensional (3D) visual objects and scenes. To efficiently store and transmit point clouds, compression methods have been developed, but they often result in a degradation of quality. To reduce color distortion in point clouds, we propose a graph-based quality enhancement network (GQE-Net) that uses geometry information as an auxiliary input and graph convolution blocks to extract local features efficiently. Specifically, we use a parallel-serial graph attention module with a multi-head graph attention mechanism to focus on important points or features and help them fuse together. Additionally, we design a feature refinement module that takes into account the normals and geometry distance between points. To work within the limitations of GPU memory capacity, the distorted point cloud is divided into overlap-allowed 3D patches, which are sent to GQE-Net for quality enhancement. To account for differences in data distribution among different color components, three models are trained for the three color components. Experimental results show that our method achieves state-of-the-art performance. For example, when implementing GQE-Net on a recent test model of the geometry-based point cloud compression (G-PCC) standard, 0.43 dB, 0.25 dB and 0.36 dB Bjφntegaard delta (BD)-peak-signal-to-noise ratio (PSNR), corresponding to 14.0%, 9.3% and 14.5% BD-rate savings were achieved on dense point clouds for the Y, Cb, and Cr components, respectively. The source code of our method is available at https://github.com/xjr998/GQE-Net.
Jinrui Xing, Hui Yuan 0001, Raouf Hamzaoui, Hao Liu 0044, Junhui Hou
IEEE Trans. Image Process.2
2023 No-Reference Bitstream-Layer Model for Perceptual Quality Assessment of V-PCC Encoded Point Clouds
abstract
No-reference bitstream-layer models for point cloud quality assessment (PCQA) use the information extracted from a bitstream for real-time and nonintrusive quality monitoring. We propose a no-reference bitstream-layer model for the perceptual quality assessment of video-based point cloud compression (V-PCC) encoded point clouds. First, we study the relationship between the perceptual coding distortion and the texture quantization parameter (TQP) when geometry encoding is lossless. The results indicate that the perceptual coding distortion depends on the texture complexity (TC). Next, we estimate TC using TQP and the texture bitrate per pixel (TBPP), both of which are extracted from the compressed bitstream without resorting to complete decoding. This allows us to build a texture distortion model as a function of TQP and TBPP. By combining this texture distortion model with a geometry distortion model that depends on the geometry quantization parameter (GQP), we obtain an overall no-reference bitstream-layer PCQA model that we call bitstreamPCQ. Experimental results show that the proposed model markedly outperforms existing models in terms of widely used performance criteria, including the Pearson linear correlation coefficient (PLCC), the Spearman rank order correlation coefficient (SRCC) and the root mean square error (RMSE).
Qi Liu 0029, Honglei Su, Tianxin Chen, Hui Yuan 0001, Raouf Hamzaoui
IEEE Trans. Multim.4
2023 Exploiting Manifold Feature Representation for Efficient Classification of 3D Point Clouds
abstract
In this paper, we propose an efficient point cloud classification method via manifold learning based feature representation. Different from conventional methods, we use manifold learning algorithms to embed point cloud features for better considering the geometric continuity on the surface. Then, the nature of point cloud can be acquired in low dimensional space, and after being concatenated with features in the original three-dimensional (3D) space, both the capability of feature representation and the classification network performance can be improved. We explore three traditional manifold algorithms (i.e., Isomap, Locally-Linear Embedding, and Laplacian eigenmaps) in detail, and finally, we select the Locally-Linear Embedding (LLE) algorithm due to its low complexity and locality consistency preservation. Furthermore, we propose a neural network based manifold learning (NNML) method to implement manifold learning based non-linear projection. Experiments demonstrate that the proposed two manifold learning methods can obtain better performances than the state-of-the-art methods, and the obtained mean class accuracy (mA) and overall accuracy (oA) can reach 91.4% and 94.4%, respectively. Moreover, because of the improved feature learning capability, the proposed NNML method can also have better classification accuracy on models with prominent geometric shapes. To further demonstrate the advantages of PointManifold, we extend it as a plug and play method for point cloud classification task, which can be directly used with existing methods and gain a significant improvement.
Dinghao Yang, Wei Gao 0003, Ge Li 0002, Hui Yuan 0001, Junhui Hou, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.4
2022 PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling
abstract
We present PU-Refiner, a generative adversarial network for point cloud upsampling. The generator of our network includes a coarse feature expansion module to create coarse upsampled features, a geometry generation module to regress a coarse point cloud from the coarse upsampled features, and a progressive geometry refinement module to restore the dense point cloud in a coarse-to-fine fashion based on the coarse upsampled point cloud. The discriminator of our network helps the generator produce point clouds closer to the target distribution. It makes full use of multi-level features to improve its classification performance. Extensive experimental results show that PU-Refiner is superior to five state-of-the-art point cloud upsampling methods. Code: https://github.com/liuhaoyun/PU-Refiner.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Wei Gao 0003, Shuai Li 0005
ICASSP2
2022 JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration
abstract
Previous reinforcement learning (RL) based image restoration studies typically train RL agents to search for recovery tools from a constructed toolset and iteratively recover images. However, we argue that these agents rely on pre-trained RL models with fixed-length paths for restoration, which performs poorly in the case of unknown distortions. To address these issues, we propose a joint exploitation and exploration reinforcement learning network (JE2Net). Specifically, we propose a new deep classification network for image feature extraction and tool selection, which serves as a model prior. Second, we design a stochastic strategy to randomly select tools and a dynamic termination strategy to adaptively stop the recovery process. In this way, the model prior and exploration mechanism can be jointly used to expand the search space and obtain more quality gain. Experimental results show that our proposed method is more flexible compared to other state-of-the-art methods and achieves significant quality improvements in the presence of unknown distortions.
Xiaoyu Zhang 0002, Wei Gao 0003, Hui Yuan 0001, Ge Li 0002
ICASSP3
2022 Graph Filter-Based Fast Motion Matching for Inter Frame Coding of MPEG G-PCC
abstract
As a key technique for geometry-based point cloud inter frame coding, motion estimation (ME) has shown the effectiveness to improve the coding efficiency. But the high complexity limits its application. We propose a novel and effective method for fast global motion matching between consecutive frames. We first divide an input point cloud into several blocks, and construct a graph for each block. Then, a graph filter is applied to remove the highest and lowest frequency coefficients based on predefined thresholds. Final-ly, by quantitatively analyzing the geometry gradients of the output of the graph filter, only some key points are reserved. Experimental results show that an average 49% encoding time can be saved with only 0.5% loss of coding efficiency for lossy geometry-lossy attributes configuration while an average 45% encoding time can be saved with no loss of coding efficiency for lossless geometry-lossless attributes configuration.
Hui Yuan 0001, Wei Gao 0003
ICIP2
2022 An Attention-Based Network for Single Image HDR Reconstruction
abstract
High dynamic range (HDR) imaging can represent a great range of real-world luminosity. In contrast, the traditional low dynamic range (LDR) imaging fails to represent a wide range of luminance since most digital cameras can capture a limited range of light intensity in a natural scene. Recent advances in deep learning allow reconstructing an HDR image from a single LDR image and surpass conventional methods performance. In this work, we propose a novel CNN for HDR image reconstruction based on residual learning and attention mechanism. The proposed network adopts an autoencoder structure with residual blocks trained in a fully end-to-end manner. Residual learning boosts the performance by optimizing the network to converge faster. Moreover, the attention mechanism allows the network to select and enhance meaningful features that will contribute to the reconstruction of the HDR image. In addition, we employ a contextual attention module to perform patch replacement on deep feature maps to help recover information in over-exposed areas. Extensive quantitative and qualitative experiments on public HDR datasets demonstrate the ability of our proposed method to effectively reconstruct a visually pleasing HDR image from a single LDR image and outperform existing approaches.
Mohamed Dafaallah, Hui Yuan 0001, Shiqi Jiang 0006
ISCAS2
2022 APCCPA '22: 1st International Workshop on Advances in Point Cloud Compression, Processing and Analysis
abstract
Point clouds are attracting much attention from academia, industry and standardization organizations such as MPEG, JPEG, and AVS. 3D Point clouds consisting of thousands or even millions of points with attributes can represent real-world objects and scenes in a way that enables an improved immersive visual experience and facilitates complex 3D vision tasks. In addition to various point cloud analysis and processing tasks (e.g., segmentation, classification, 3D object detection, registration), efficient compression for these large-scale 3D visual data is essential to make point cloud applications more effective. This workshop focuses on point cloud processing, analy sis, and compression in challenging situations to further improve visual experience and machine vision performance. Both learning-based and non-learning-based perception-oriented optimization algorithms for compression and processing are solicited. Contributions that advance the state-of-the-art in analysis tasks, are also welcomed.
Wei Gao 0003, Ge Li 0002, Hui Yuan 0001, Raouf Hamzaoui, Zhu Li 0001, Shan Liu 0001
ACM Multimedia3
2022 CATFPN: Adaptive Feature Pyramid With Scale-Wise Concatenation and Self-Attention
abstract
It is a typical problem in the field of object detection to simultaneously detect objects with large scale variation in one image. Recently proposed state-of-the-art object detectors generally learn pyramidal feature representation to deal with the scale variation, which has been proved effective via various feature pyramid networks. However, the majority of the feature pyramid networks based on heuristic feature fusion strategies may be suboptimal, as excess human guidance will restrict the self-learning of deep neural networks. An adaptive feature pyramid is bound to provide a significant performance boost. In this paper, we propose a novel feature pyramid network named CATFPN that consists of Scale-Wise Feature Concatenation (SWFC) module and Global Context (GC) block. The SWFC module evenly distributes semantic features for each feature layer and the GC block introduces a self-attention mechanism. As a feature pyramid network, the CATFPN can be applied to any detector based on multi-scale features. We adopt the CATFPN in typical RetinaNet and Faster R-CNN detector models, without bells and whistles, achieving 1.1% AP and 0.7% AP improvements over FPN on the MS COCO benchmark, respectively. Our competitive performance reported on the test-dev subset of COCO achieves 42.3% AP.
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu, Hui Yuan 0001, Weikai He
IEEE Trans. Circuits Syst. Video Technol.5
2022 A Hybrid Compression Framework for Color Attributes of Static 3D Point Clouds
abstract
The emergence of 3D point clouds (3DPCs) is promoting the rapid development of immersive communication, autonomous driving, and so on. Due to the huge data volume, the compression of 3DPCs is becoming more and more attractive. We propose a novel and efficient color attribute compression method for static 3DPCs. First, a 3DPC is partitioned into several sub-point clouds by color distribution analysis. Each sub-point cloud is then decomposed into a lot of 3D blocks by an improved k-d tree-based decomposition algorithm. Afterwards, a novel virtual adaptive sampling-based sparse representation strategy is proposed for each 3D block to remove the redundancy among points, in which the bases of the graph transform (GT) and the discrete cosine transform (DCT) are used as candidates of the complete dictionary. Experimental results over 10 common 3DPCs demonstrate that the proposed method can achieve superior or comparable coding performance when compared with the current state-of-the-art methods.
Hao Liu 0044, Hui Yuan 0001, Qi Liu 0029, Junhui Hou, Huanqiang Zeng, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.2
2022 Unified Cross-Component Linear Model in VVC Based on a Subset of Neighboring Samples
abstract
To compress industrial video content efficiently, H.266/Versatile Video Coding (VVC) introduces cross-component linear model (CCLM) prediction as a new coding tool, in which chroma components are predicted from the luma component based on a linear model. In this article, we propose a subset-based CCLM (S-CCLM), in which the model parameters are derived based on a subset of neighboring samples. To choose the most proper subset, we build the relationship between the prediction error and the geometric distance and resolve the optimal subset construction problem by minimizing the geometric distance. With the well-designed subset, a weight-guided parameter derivation algorithm is further proposed to improve the accuracy of the model parameters. The experimental results show that the proposed S-CCLM can achieve Bjontegaard delta bitrate (BD-rate) reductions of 0.14%, 0.64%, and 0.75% for the Y, Cb, and Cr components, respectively, when the number of samples in the subset,$N$, is 4 and BD-rate reductions of 0.22%, 0.80%, and 0.95% when$N$is 8. Given a small fixed$N$, fewer memory access operations are needed during the CCLM calculation, and a unified CCLM process can be achieved for coding blocks with different sizes and different modes. Due to its hardware-friendly architecture, the S-CCLM has been partially adopted by H.266/VVC.
Junyan Huo, Hongqing Du, Shuai Wan, Hui Yuan 0001, Yanzhuo Ma, Fuzheng Yang 0001
IEEE Trans. Ind. Informatics5
2022 PUFA-GAN: A Frequency-Aware Generative Adversarial Network for 3D Point Cloud Upsampling
abstract
We propose a generative adversarial network for point cloud upsampling, which can not only make the upsampled points evenly distributed on the underlying surface but also efficiently generate clean high frequency regions. The generator of our network includes a dynamic graph hierarchical residual aggregation unit and a hierarchical residual aggregation unit for point feature extraction and upsampling, respectively. The former extracts multiscale point-wise descriptive features, while the latter captures rich feature details with hierarchical residuals. To generate neat edges, our discriminator uses a graph filter to extract and retain high frequency points. The generated high resolution point cloud and corresponding high frequency points help the discriminator learn the global and high frequency properties of the point cloud. We also propose an identity distribution loss function to make sure that the upsampled points remain on the underlying surface of the input low resolution point cloud. To assess the regularity of the upsampled points in high frequency regions, we introduce two evaluation metrics. Objective and subjective results demonstrate that the visual quality of the upsampled point clouds generated by our method is better than that of the state-of-the-art methods.
Hao Liu 0044, Hui Yuan 0001, Junhui Hou, Raouf Hamzaoui, Wei Gao 0003
IEEE Trans. Image Process.2
2022 A Hybrid Control Scheme for 360-Degree Dynamic Adaptive Video Streaming Over Mobile Devices
abstract
A 360-degree streaming system can provide immersive, interactive, and autonomous experiences surrounding the user by means of viewpoint changes to see different angles of a 360-degree video. However, due to the limited capacity and highly dynamic conditions of cellular networks, high-resolution 360-degree video playback over mobile devices often suffers from playback freezing, and bandwidth waste is inevitably incurred in delivering out-of-view video data. In this paper, a hybrid control scheme is presented for segment-level continuous bitrate selection and tile-level bitrate allocation for 360-degree streaming over mobile devices to increase users’ quality of experience. First, a deep reinforcement learning (RL) method is proposed to predict the segment bitrate and avoid playback freezing. Second, a viewpoint-prediction-map-based cooperative bargaining game theory is proposed for bitrate allocation optimization to choose a suitable bitrate for each tile to reduce unreasonable bandwidth waste. The proposed scheme is compared with state-of-the-art approaches under a wide variety of mobile network conditions with multiple viewpoint traces and 360-degree video contents. The experimental results indicate that the proposed method outperforms the compared state-of-the-art approaches in terms of various experimental objectives on mobile devices.
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Weijia Jia 0001
IEEE Trans. Mob. Comput.4
2021 CorrNet3D: Unsupervised End-to-End Learning of Dense Correspondence for 3D Point Clouds
abstract
Motivated by the intuition that one can transform two aligned point clouds to each other more easily and meaningfully than a misaligned pair, we propose CorrNet3D – the first unsupervised and end-to-end deep learning-based framework – to drive the learning of dense correspondence between 3D shapes by means of deformation-like reconstruction to overcome the need for annotated data. Specifically, CorrNet3D consists of a deep feature embedding module and two novel modules called correspondence indicator and symmetric deformer. Feeding a pair of raw point clouds, our model first learns the pointwise features and passes them into the indicator to generate a learnable correspondence matrix used to permute the input pair. The symmetric deformer, with an additional regularized loss, transforms the two permuted point clouds to each other to drive the unsupervised learning of the correspondence. The extensive experiments on both synthetic and real-world datasets of rigid and non-rigid 3D shapes show our CorrNet3D outperforms state-of-the-art methods to a large extent, including those taking meshes as input. CorrNet3D is a flexible framework in that it can be easily adapted to supervised learning if annotated data are available. The source code and pre-trained model will be available at https://github.com/ZENGYIMINGEAMON/CorrNet3D.git.
Yiming Zeng 0002, Junhui Hou, Hui Yuan 0001, Ying He 0001
CVPR5
2021 Hierarchical Bit-Wise Differential Coding (HBDC) of Point Cloud Attributes
abstract
Targeting both computing and coding efficiencies, we propose in this work a novel hierarchical bit-wise differential coding scheme to compress point cloud attributes. The encoder firstly quantizes and organizes the points into an octree structure and, for each internal node, picks its attribute(s) from a child named source child. Next, the encoder conducts a top-down scanning of the hierarchy. For each node with more than one child, it computes the bit-wise attribute difference between the current node and each non-source child by exclusive-OR and encodes the difference with an arithmetic coder. Further, a table look-up approach is proposed to accelerate the online source child identification. The proposed scheme produces superior computing and coding efficiencies for lossless point cloud attribute compression, outperforming the MPEG benchmark coders by large margins in our experiments.
Bin Wang 0040, C.-C. Jay Kuo, Hui Yuan 0001, Jingliang Peng
ICASSP4
2021 Joint Reinforcement Learning and Game Theory Bitrate Control Method for 360-Degree Dynamic Adaptive Streaming
abstract
A joint reinforcement learning (RL) and game theory method is presented for segment-level continuous bitrate selection and tile-level bitrate allocation in tile-based 360-degree streaming to increase users’ quality of experience (QoE). First, a viewpoint prediction method based on single-user (SU) viewpoint traces and the saliency map (SM) model is presented to model viewing behaviours. Second, an RL method is proposed to predict segment bitrate and a cooperative bargaining game theory is proposed for bitrate allocation optimization to choose a suitable bitrate for every tile with the help of the viewpoint prediction map. Performance evaluation results indicate that the proposed method can outperform the state-of-the-art methods in terms of different QoE objectives.
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Tao Xiang 0001
ICASSP4
2021 Model-Based Rate-Distortion Optimized Video-Based Point Cloud Compression with Differential Evolution
Hui Yuan 0001, Raouf Hamzaoui, Ferrante Neri, Shengxiang Yang
ICIG (1)1
2021 DRLFNet: A Dense-Connection Residual Learning Neural Network for Light Field Super Resolution
Congrui Fu, Junhui Hou, Hui Yuan 0001
ICIG (3)5
2021 Adaptive Quantization for Predicting Transform-Based Point Cloud Compression
Guoxia Sun, Hui Yuan 0001, Raouf Hamzaoui
ICIG (1)3
2021 QOE-Based Neural Live Streaming Method with Continuous Dynamic Adaptive Video Quality Control
abstract
In this paper, a quality of experience (QoE)-based neural live streaming method with dynamic adaptive video quality control is developed to improve streaming performance. First, the dynamic adaptive streaming issue is formulated as a Markov decision process (MDP) problem. Second, an reinforcement learning (RL)-based approach is proposed as an appropriate solution, where the client functions as an RL agent and the environment is made up of various networks. User QoE is the reward by mutual consideration of video quality and play-back state. Finally, to optimize the total reward, the RL algorithm chooses the required video quality for each video segment. Experimental results show that the proposed RL-based streaming algorithm outperforms state-of-the-art schemes in terms of both temporal and visual QoE metrics by a noticeable margin while guaranteeing application-level fairness when multiple clients share a bottlenecked network. The code is available on the following website: https://github.com/OpenCode007/ICME2021.
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Tao Xiang 0001
ICME4
2021 Monocular 3D Pedestrian Localization Fusing with Bird's Eye View
abstract
In recent years, 3D target detection and location methods in the field of autonomous driving have attracted increasing attention, but monocular 3D pedestrian localization research is still facing challenges. In this paper, a monocular pedestrian localization framework and a fine-grained location optimization method which is based on a bird's eye view are proposed. The monocular pedestrian localization framework is divided into three parts, which are coarse-grained localization, depth information reconstruction and fine-grained location optimization. In the stage of coarse-grained location, the human skeleton information is obtained from the original image by using the human skeleton point detection method, and then the pedestrian position is predicted through the method of a light-weight feed-forward neural network. In the stage of depth information reconstruction, the original image is used to reconstruct the corresponding bird's eye view with depth information through a parallel network. Finally, a fine-grained positioning optimization method makes it possible to get a more precise location with the help of the last two stages. The experimental results on the KITTI dataset show that our method has achieved better performance than the state-of-the-art methods.
Shanxin Zhang, Hui Yuan 0001, Xinghai Yang, Huaxiang Zhang 0001, Jiande Sun 0001
ISCAS3
2021 Global Rate-distortion Optimization of Video-based Point Cloud Compression with Differential Evolution
abstract
In video-based point cloud compression (V-PCC), one geometry video and one color video are generated from a dynamic point cloud. Then, the two videos are compressed independently using a state-of-the-art video coder. In the Moving Picture Experts Group (MPEG) V-PCC test model, the quantization parameters for a given group of frames are constrained according to a fixed offset rule. For example, for the low-delay configuration, the difference between the quantization parameters of the first frame and the quantization parameters of the following frames in the same group is zero by default. We show that the rate-distortion performance of the V-PCC test model can be improved by lifting this constraint and considering the ratedistortion optimization problem as a multi-variable constrained combinatorial optimization problem where the variables are the quantization parameters of all frames. To solve the optimization problem, we use a variant of the differential evolution algorithm. Experimental results for the low-delay configuration show that our method can achieve a Bjøntegaard delta bitrate of up to -43.04% and more accurate rate control (average bitrate error to the target bitrate of 0.45% vs. 10.75%) compared to the state-of- the-art method, which optimizes the rate-distortion performance subject to the test model default offset rule. We also show that our optimization strategy can be used to improve the rate-distortion performance of two-dimensional video coders.
Hui Yuan 0001, Raouf Hamzaoui, Ferrante Neri, Shengxiang Yang
MMSP1
2021 Kalman filter-based prediction refinement and quality enhancement for geometry-based point cloud compression
abstract
A point cloud is a set of points representing a three-dimensional (3D) object or scene. To compress a point cloud, the Motion Picture Experts Group (MPEG) geometry-based point cloud compression (G-PCC) scheme may use three attribute coding methods: region adaptive hierarchical transform (RAHT), predicting transform (PT), and lifting transform (LT). To improve the coding efficiency of PT, we propose to use a Kalman filter to refine the predicted attribute values. We also apply a Kalman filter to improve the quality of the reconstructed attribute values at the decoder side. Experimental results show that the combination of the two proposed methods can achieve an average Bjøntegaard delta bitrate of −0.48%, −5.18%, and −6.27% for the Luma, Chroma Cb, and Chroma Cr components, respectively, compared with a recent G-PCC reference software.
Jian Sun 0013, Hui Yuan 0001, Raouf Hamzaoui
VCIP3
2021 ConvLSTM-based Neural Network for Video Semantic Segmentation
abstract
We propose a convolutional long short-term memory(ConvLSTM)-based neural network for video semantic segmentation. The network can capture the timing information between frames through the ConvLSTM module to improve the prediction accuracy. The back-bone network uses dense connection, atrous convolution, and pooling pyramid structure to expand the receptive field. During the training, to avoid over fitting, data augmentation and learning rate attenuation strategies were used. The proposed method is end-to-end trainable and evaluated on the street scene benchmark Cityscapes dataset. Experimental results show that, benefit from the ConvLSTM module, the proposed network can extract temporal information between frames effectively, and thus improve the accuracy of video semantic segmentation, especially for dynamic objects and small obiects. such as truck. pedestrian and pole.
Hui Yuan 0001, Chuan Ge
VCIP2
2021 Reinforcement learning-based QoE-oriented dynamic adaptive streaming framework
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Shiqi Wang 0001, Guopu Zhu, Jingchao Cao
Inf. Sci.4
2021 PQA-Net: Deep No Reference Point Cloud Quality Assessment via Multi-View Projection
abstract
Recently, 3D point cloud is becoming popular due to its capability to represent the real world for advanced content modality in modern communication systems. In view of its wide applications, especially for immersive communication towards human perception, quality metrics for point clouds are essential. Existing point cloud quality evaluations rely on a full or certain portion of the original point cloud, which severely limits their applications. To overcome this problem, we propose a novel deep learning-based no reference point cloud quality assessment method, namely PQA-Net. Specifically, the PQA-Net consists of a multi-view-based joint feature extraction and fusion (MVFEF) module, a distortion type identification (DTI) module, and a quality vector prediction (QVP) module. The DTI and QVP modules share the feature generated from the MVFEF module. By using the distortion type labels, the DTI and the MVFEF modules are first pre-trained to initialize the network parameters, based on which the whole network is then jointly trained to finally evaluate the point cloud quality. Experimental results on the Waterloo Point Cloud dataset show that PQA-Net achieves better or equivalent performance comparing with the state-of-the-art quality assessment methods. The code of the proposed model will be made publicly available to facilitate reproducible researchhttps://github.com/qdushl/PQA-Net.
Qi Liu 0029, Hui Yuan 0001, Honglei Su, Hao Liu 0044, Yu Wang 0106, Huan Yang 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.2
2021 Reduced Reference Perceptual Quality Model With Application to Rate Control for Video-Based Point Cloud Compression
abstract
In rate-distortion optimization, the encoder settings are determined by maximizing a reconstruction quality measure subject to a constraint on the bitrate. One of the main challenges of this approach is to define a quality measure that can be computed with low computational cost and which correlates well with the perceptual quality. While several quality measures that fulfil these two criteria have been developed for images and videos, no such one exists for point clouds. We address this limitation for the video-based point cloud compression (V-PCC) standard by proposing a linear perceptual quality model whose variables are the V-PCC geometry and color quantization step sizes and whose coefficients can easily be computed from two features extracted from the original point cloud. Subjective quality tests with 400 compressed point clouds show that the proposed model correlates well with the mean opinion score, outperforming state-of-the-art full reference objective measures in terms of Spearman rank-order and Pearson linear correlation coefficient. Moreover, we show that for the same target bitrate, rate-distortion optimization based on the proposed model offers higher perceptual quality than rate-distortion optimization based on exhaustive search with a point-to-point objective quality metric. Our datasets are publicly available at https://github.com/qdushl/Waterloo-Point-Cloud-Database-2.0.
Qi Liu 0029, Hui Yuan 0001, Raouf Hamzaoui, Honglei Su, Junhui Hou, Huan Yang 0001
IEEE Trans. Image Process.2
2021 Model-Based Joint Bit Allocation Between Geometry and Color for Video-Based 3D Point Cloud Compression
abstract
In video-based 3D point cloud compression, the quality of the reconstructed 3D point cloud depends on both the geometry, and color distortions. Finding an optimal allocation of the total bitrate between the geometry coder, and the color coder is a challenging task due to the large number of possible solutions. To solve this bit allocation problem, we first propose analytical distortion, and rate models for the geometry, and color information. Using these models, we formulate the joint bit allocation problem as a constrained convex optimization problem, and solve it with an interior point method. Experimental results show that the rate-distortion performance of the proposed solution is close to that obtained with exhaustive search but at only 0.66$\%$of its time complexity.
Qi Liu 0029, Hui Yuan 0001, Junhui Hou, Raouf Hamzaoui, Honglei Su
IEEE Trans. Multim.2
2020 Learning Light Field Angular Super-Resolution via a Geometry-Aware Network
abstract
The acquisition of light field images with high angular resolution is costly. Although many methods have been proposed to improve the angular resolution of a sparsely-sampled light field, they always focus on the light field with a small baseline, which is captured by a consumer light field camera. By making full use of the intrinsic geometry information of light fields, in this paper we propose an end-to-end learning-based approach aiming at angularly super-resolving a sparsely-sampled light field with a large baseline. Our model consists of two learnable modules and a physically-based module. Specifically, it includes a depth estimation module for explicitly modeling the scene geometry, a physically-based warping for novel views synthesis, and a light field blending module specifically designed for light field reconstruction. Moreover, we introduce a novel loss function to promote the preservation of the light field parallax structure. Experimental results over various light field datasets including large baseline light field images demonstrate the significant superiority of our method when compared with state-of-the-art ones, i.e., our method improves the PSNR of the second best method up to 2 dB in average, while saves the execution time 48×. In addition, our method preserves the light field parallax structure better.
Jing Jin 0006, Junhui Hou, Hui Yuan 0001, Sam Kwong
AAAI3
2020 A Sampling-based 3D Point Cloud Compression Algorithm for Immersive Communication
Hui Yuan 0001, Dexiang Zhang
Mob. Networks Appl.1
2020 Single image-based head pose estimation with spherical parametrization and 3D morphing
Hui Yuan 0001, Junhui Hou, Jimin Xiao
Pattern Recognit.1
2020 3D Point Cloud Attribute Compression via Graph Prediction
abstract
3D point clouds associated with attributes are considered as a promising data representation for immersive communication. The large amount of data, however, poses great challenges to the subsequent transmission and storage processes. In this letter, we propose a new compression scheme for the color attribute of static voxelized 3D point clouds. Specifically, we first partition the colors of a 3D point cloud into clusters by applying k-d tree to the geometry information, which are then successively encoded. To eliminate the redundancy, we propose a novel prediction module, namely graph prediction, in which a small number of representative points selected from previously encoded clusters are used to predict the points to be encoded by exploring the underlying graph structure constructed from the geometry information. Furthermore, the prediction residuals are transformed with the graph transform, and the resulting transform coefficients are finally uniformly quantified and entropy encoded. Experimental results show that the proposed compression scheme is able to achieve better rate-distortion performance at a lower computational cost when compared with state-of-the-art methods.
Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001
IEEE Signal Process. Lett.4
2020 3D Point Cloud Attribute Compression Using Geometry-Guided Sparse Representation
abstract
3D point clouds associated with attributes are considered as a promising paradigm for immersive communication. However, the corresponding compression schemes for this media are still in the infant stage. Moreover, in contrast to conventional image/video compression, it is a more challenging task to compress 3D point cloud data, arising from the irregular structure. In this paper, we propose a novel and effective compression scheme for the attributes of voxelized 3D point clouds. In the first stage, an input voxelized 3D point cloud is divided into blocks of equal size. Then, to deal with the irregular structure of 3D point clouds, a geometry-guided sparse representation (GSR) is proposed to eliminate the redundancy within each block, which is formulated as an ℓ0-norm regularized optimization problem. Also, an inter-block prediction scheme is applied to remove the redundancy between blocks. Finally, by quantitatively analyzing the characteristics of the resulting transform coefficients by GSR, an effective entropy coding strategy that is tailored to our GSR is developed to generate the bitstream. Experimental results over various benchmark datasets show that the proposed compression scheme is able to achieve better rate-distortion performance and visual quality, compared with state-of-the-art methods.
Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001, Kai-Kuang Ma
IEEE Trans. Image Process.4
2020 Frame-level Bit Allocation Optimization Based on Video Content Characteristics for HEVC
abstract
Rate control plays an important role in high efficiency video coding (HEVC), and bit allocation is the foundation of rate control. The video content characteristics are significant for bit allocation, and modeling an accurate relationship between video content characteristics and bit allocation is essential for bit allocation optimization. Therefore, in this article, a video content characteristics–based frame-level optimal bit allocation algorithm is proposed for improving the rate distortion (RD) performance of HEVC. First, the number of search points of motion estimation is used to evaluate the motion activity of video content, and the relationship between the search points and bit allocation is modeled as the search-points model. Second, the grey level co-occurrence matrix and temporal perceptual information are used to evaluate the spatial and temporal texture complexity, and the relationship between the video content texture complexity and bit allocation is modeled as the texture-complexity model. Then, the search-points model and texture-complexity model are jointly employed to allocate the coding bits for the second and third layers of the HEVC hierarchical coding structure. Finally, the remaining coding bits of a group-of-pictures (GOP) are allocated to the first layer of HEVC coding structure. To evaluate the performance of the proposed algorithm, the RD performance and bitrate accuracy are used as evaluation criteria, and the experimental results show that when compared with the popularly used R-λ model–based bit allocation algorithm, the proposed algorithm achieves an average of -3.43% BDBR reduction and 0.13 dB BDPSNR gains with only 0.02% loss of bitrate accuracy.
Zhaoqing Pan, Xiaokai Yi, Yun Zhang 0002, Hui Yuan 0001, Fu Lee Wang, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.4
2019 A Fidelity-Assured Rate Distortion Optimization Method for Perceptual-Based Video Coding
abstract
Rate-distortion optimization (RDO) is one of the essential method to improve the video coding efficiency. The main target of RDO is to find the optimal tradeoff between the reconstructed video quality and encoding rate. In traditional video coding standards, e.g. the emerging H.266/Versatile Video Coding(VVC), the H.265/High Efficiency Video Coding (HEVC) and H.264/Advanced Video Coding (AVC), sum of squared error (SSE) is used as the distortion criterion because SSE can represent the image fidelity efficiently. Based on the existing research on human visual characteristic, the perceptual visual quality is not consistent with image fidelity. Accordingly, we propose a video coding method to improve the subjective quality while avoiding great fidelity degradation. Experimental results demonstrate that the proposed method is efficient in preserving subjective qualities with only a little fidelity degradation when comparing to existing subjective quality based rate distortion optimization methods.
Hui Yuan 0001, Junyan Huo
ICIP2
2018 A novel distortion criterion of rate-distortion optimization for depth map coding
Ziqi Zheng, Junyan Huo, Hui Yuan 0001, Weisi Lin
J. Vis. Commun. Image Represent.4
2018 Adaptive Lagrangian Multiplier derivation model for depth map coding
Ziqi Zheng, Junyan Huo, Hui Yuan 0001
Signal Process. Image Commun.4
2018 Fine Virtual View Distortion Estimation Method for Depth Map Coding
abstract
In three-dimensional (3-D) video coding systems, depth maps represent the geometric information of a 3-D scene. Since depth maps are not displayed to viewers but to generate virtual views, the quality of depth maps needs to be measured by its effect on the virtual view quality, which is indicated by the virtual view distortion (VVD) in depth map coding. In this letter, a fine VVD estimation method is proposed based on the analysis of the VVD. Specifically, the depth distortion of the current pixel, the texture gradient of the colocated color video and the depth distortions of adjacent pixels are all taken into consideration to estimate the VVD accurately. Experimental results demonstrate that the proposed method can improve 13.1% bitrate saving compared with the sum of squared differences based depth distortion calculation method and can improve 1.2% bitrate saving compared with the VVD estimation method in three dimensional high efficiency video coding (3D-HEVC) reference software.
Ziqi Zheng, Junyan Huo, Hui Yuan 0001
IEEE Signal Process. Lett.4
2018 Region Adaptive R-λ Model-Based Rate Control for Depth Maps Coding
abstract
In this paper, a novel rate-control algorithm based on the region adaptive R-λ model is proposed for depth maps coding. First, in order to obtain an accurate rate control for depth maps coding, a modified frame level bit allocation method based on coding bits statistical distribution of depth maps is proposed. Second, considering that different areas in a depth map have an imparity effect on virtual view rendering, the blocks of the depth map are divided into two types, namely, interested blocks for virtual view rending (IBV) and noninterested blocks for virtual view rending (NIBV). Then, two different R-λ models are derived for IBV and NIBV, respectively. The optimal bitrates for IBV and NIBV are determined by solving an optimization problem. After that, based on the regional R-λ models, the optimal Lagrange multipliers are calculated for both IBV and NIBV. Finally, the largest coding unit (LCU) level rate control is performed by adaptively adjusting the Lagrange multiplier to avoid blocking artifacts and smooth the quality of coding. Experimental results demonstrate that the proposed method can achieve considerable BD-PSNR gains compared with the unified rate-quantization model and conventional R-λ modelbased algorithms in terms of rendered virtual views quality.
Jianjun Lei 0001, Xiaoxu He, Hui Yuan 0001, Feng Wu 0001, Nam Ling, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.3
2018 Convolutional Neural Network-Based Synthesized View Quality Enhancement for 3D Video Coding
abstract
The quality of synthesized view plays an important role in the three dimensional (3D) video system. In this paper, to further improve the coding efficiency, a convolutional neural network (CNN) based synthesized view quality enhancement method for 3D High Efficiency Video Coding (HEVC) is proposed. Firstly, the distortion elimination in synthesized view is formulated as an image restoration task with the aim to reconstruct the latent distortion free synthesized image. Secondly, the learned CNN models are incorporated into 3D HEVC codec to improve the view synthesis performance for both view synthesis optimization (VSO) and the final synthesized view, where the geometric and compression distortions are considered according to the specific characteristics of synthesized view. Thirdly, a new Lagrange multiplier in the rate-distortion (RD) cost function is derived to adapt the CNN based VSO process to embrace a better 3D video coding performance. Extensive experimental results show that the proposed scheme can efficiently eliminate the artifacts in the synthesized image, and reduce 25.9% and 11.7% bit rate in terms of peak-signal-to-noise ratio (PSNR) and structural similarity (SSIM) index, which significantly outperforms the state-of-theart methods.
Linwei Zhu, Yun Zhang 0002, Shiqi Wang 0001, Hui Yuan 0001, Sam Kwong, Horace Ho-Shing Ip
IEEE Trans. Image Process.4
2018 Non-Cooperative Game Theory Based Rate Adaptation for Dynamic Video Streaming over HTTP
abstract
Dynamic Adaptive Streaming over HTTP (DASH) has demonstrated to be an emerging and promising multimedia streaming technique, owing to its capability of dealing with the variability of networks. Rate adaptation mechanism, a challenging and open issue, plays an important role in DASH based systems since it affects Quality of Experience (QoE) of users, network utilization, etc. In this paper, based on non-cooperative game theory, we propose a novel algorithm to optimally allocate the limited export bandwidth of the server to multi-users to maximize their QoE with fairness guaranteed. The proposed algorithm is proxy-free. Specifically, a novel user QoE model is derived by taking a variety of factors into account, like the received video quality, the reference buffer length, and user accumulated buffer lengths, etc. Then, the bandwidth competing problem is formulated as a non-cooperation game with the existence of Nash Equilibrium that is theoretically proven. Finally, a distributed iterative algorithm with stability analysis is proposed to find the Nash Equilibrium. Compared with state-of-the-art methods, extensive experimental results in terms of both simulated and realistic networking scenarios demonstrate that the proposed algorithm can produce higher QoE, and the actual buffer lengths of all users keep nearly optimal states, i.e., moving around the reference buffer all the time. Besides, the proposed algorithm produces no playback interruption.
Hui Yuan 0001, Huayong Fu, Junhui Hou, Sam Kwong
IEEE Trans. Mob. Comput.1
2018 Cooperative Bargaining Game-Based Multiuser Bandwidth Allocation for Dynamic Adaptive Streaming Over HTTP
abstract
Dynamic adaptive streaming over HTTP (DASH) has emerged as an efficient technology for video streaming. For a DASH system, a most common case is that a limited server bandwidth is competed by multiusers. In order to improve user quality of experience (QoE) and guarantee fairness, we propose to use the game theory in a proxy server to allocate the bandwidth collaboratively for multiusers. By taking user buffer length, received video bit rates, video qualities, etc., into account, the bandwidth allocation problem is formulated as a cooperative bargaining problem and the Nash bargaining solution (NBS) is obtained by convex optimization. The requested bit rate of users will be rewritten as the proxy calculated bit rate (i.e., NBS) when the user requested bit rate is larger. Experimental results demonstrate that user QoE and fairness can be improved significantly, i.e., the delay frequency and duration are smaller, and the received video qualities are higher and more stable, when comparing the proposed method with existing methods.
Hui Yuan 0001, Xuekai Wei, Fuzheng Yang 0001, Jimin Xiao, Sam Kwong
IEEE Trans. Multim.1
2017 Spatial/temporal motion consistency based MERGE mode early decision for HEVC
Junaid Tariq, Sam Kwong, Hui Yuan 0001
J. Vis. Commun. Image Represent.3
2017 Motion-Homogeneous-Based Fast Transcoding Method From H.264/AVC to HEVC
abstract
With the popularity of high-efficiency video coding (HEVC) standard, a video server usually transcodes a video stream to HEVC for its higher compression ratio. In this paper, a fast H.264/advanced video coding (AVC) to HEVC transcoding method is proposed. In the HEVC encoding procedure, a coding unit (CU), which is a motion-homogeneous block, is first checked based on the analysis of the decoded information from H.264/AVC bit stream. Then, for motion-homogeneous blocks, CU depth and the corresponding prediction unit (PU) mode's early termination strategies are proposed based on the CU size and corresponding prior statistical knowledge. For non-motion-homogeneous blocks, a corresponding PU mode's early termination strategy is also proposed. Experimental results demonstrate the effectiveness of the proposed method.
Hui Yuan 0001, Chenglin Guo, Xu Wang 0006, Sam Kwong
IEEE Trans. Multim.1
2016 HEVC intra mode selection based on Rate Distortion (RD) cost and Sum of Absolute Difference (SAD)
Junaid Tariq, Sam Kwong, Hui Yuan 0001
J. Vis. Commun. Image Represent.3
2016 DCT Coefficient Distribution Modeling and Quality Dependency Analysis Based Frame-Level Bit Allocation for HEVC
abstract
A frame-level bit allocation optimization method is proposed to improve the rate-distortion performance for High Efficiency Video Coding. First, to avoid the demerits of the mixture Laplacian distribution model on complexity, a new synthesized Laplacian distribution (SynLD) model is proposed to describe the discrete cosine transform transformed coefficients based on Kullback-Leibler-divergence analysis. Second, quality dependencies among frames are investigated, and a linear relationship between quality dependency factor (QDF) and skipmode percentage is proposed for QDF prediction. Based on the proposed SynLD model and QDF prediction method, a p-domain-based frame-level bit allocation method is proposed. Experimental results show that when compared with the state-of-the-art pixel-based unified rate-quantization (URQ) model and R-λ-model-based algorithms, 1.75- and 0.16-dB BD-peak signalto-noise ratio (PSNR) gains can be achieved by the proposed bit allocation method, respectively. For quality consistency, the average PSNR standard deviation shows 0.16 and 0.02 dB lower than URQ and R-λ-model-based algorithms, respectively. The proposed method also has a much more stable buffer control status and works well for scene change cases.
Wei Gao 0003, Sam Kwong, Hui Yuan 0001, Xu Wang 0006
IEEE Trans. Circuits Syst. Video Technol.3
2016 SSIM-Based Game Theory Approach for Rate-Distortion Optimized Intra Frame CTU-Level Bit Allocation
abstract
A structural similarity (SSIM)-based game theory (GT) approach is proposed for rate-distortion (R-D) optimized CTU-level bit allocation in high efficiency video coding (HEVC). First, a SSIM-based bargaining game is formulated and the Nash bargaining solution (NBS) is proposed, in which a SSIM-based initial minimum utility is defined. Second, we propose a two-stage remaining bit refinement-based bit allocation scheme. The optimization scheme of the SSIM-based bargaining game sufficiently considers the different R-D characteristics of coding tree units (CTUs), in which the feasible utility set is proved to be convex based on the proposed SSIM-based utility and R-SSIM model. Compared with the other state-of-the-art CTU-level bit allocation methods, the R-D performance improvements on Bjøntegaard delta bit-rate (BD-BR), Bjøntegaard delta peak-signal-to-noise-ratio (BD-PSNR), and BD-SSIM metrics of the proposed method can averagely achieve significant gains, respectively. The achieved R-D performance gains have been very close to the coding performance limits from the FixedQP method. Moreover, the proposed SSIM-GT method also maintains good performances on quality smoothness, bit rate accuracy, and encoding complexity.
Wei Gao 0003, Sam Kwong, Yu Zhou 0027, Hui Yuan 0001
IEEE Trans. Multim.4
2015 Smooth View Quality Oriented Bit Allocation Optimization for 3D Video Coding
abstract
View level bit allocation is an fundamental optimization problem in multiview video plus depth (MVD) based 3D video coding (3DVC). In this paper, we propose a smooth view quality oriented view level bit allocation framework for MVD based 3DVC. The Cauchy-density based rate-distortion model of the texture video and depth map are employed to represent the rate distortion properties. The relationship between the distortion of synthesized view and quantization step size of texture videos and depth maps is approximately fitted as linear model. Final, the bit allocation problem is solved by convex optimization algorithms. Experimental results demonstrated that our proposed algorithm can achieve good performance with acceptable computational complexity comparing to the full search scheme.
Xu Wang 0006, Sam Kwong, Wei Gao 0003, Yu Zhou 0027, Hui Yuan 0001, Yun Zhang 0002
SMC5
2015 View synthesis distortion model based frame level rate control optimization for multiview depth video coding
Xu Wang 0006, Sam Kwong, Hui Yuan 0001, Yun Zhang 0002, Zhaoqing Pan
Signal Process.3
2015 Low Complexity HEVC INTRA Coding for High-Quality Mobile Video Communication
abstract
INTRA video coding is essential for high quality mobile video communication and industrial video applications since it enhances video quality, prevents error propagation, and facilitates random access. The latest high-efficiency video coding (HEVC) standard has adopted flexible quad-tree-based block structure and complex angular INTRA prediction to improve the coding efficiency. However, these technologies increase the coding complexity significantly, which consumes large hardware resources, computing time and power cost, and is an obstacle for real-time video applications. To reduce the coding complexity and save power cost, we propose a fast INTRA coding unit (CU) depth decision method based on statistical modeling and correlation analyses. First, we analyze the spatial CU depth correlation with different textures and present effective strategies to predict the most probable depth range based on the spatial correlation among CUs. Since the spatial correlation may fail for image boundary and transitional areas between textural and smooth areas, we then present a statistical model-based CU decision approach in which adaptive early termination thresholds are determined and updated based on the rate-distortion (RD) cost distribution, video content, and quantization parameters (QPs). Experimental results show that the proposed method can reduce the complexity by about 56.76% and 55.61% on average for various sequences and configurations; meanwhile, the RD degradation is negligible.
Yun Zhang 0002, Sam Kwong, Zhaoqing Pan, Hui Yuan 0001, Gangyi Jiang
IEEE Trans. Ind. Informatics5
2015 Machine Learning-Based Coding Unit Depth Decisions for Flexible Complexity Allocation in High Efficiency Video Coding
abstract
In this paper, we propose a machine learning-based fast coding unit (CU) depth decision method for High Efficiency Video Coding (HEVC), which optimizes the complexity allocation at CU level with given rate-distortion (RD) cost constraints. First, we analyze quad-tree CU depth decision process in HEVC and model it as a three-level of hierarchical binary decision problem. Second, a flexible CU depth decision structure is presented, which allows the performances of each CU depth decision be smoothly transferred between the coding complexity and RD performance. Then, a three-output joint classifier consists of multiple binary classifiers with different parameters is designed to control the risk of false prediction. Finally, a sophisticated RD-complexity model is derived to determine the optimal parameters for the joint classifier, which is capable of minimizing the complexity in each CU depth at given RD degradation constraints. Comparative experiments over various sequences show that the proposed CU depth decision algorithm can reduce the computational complexity from 28.82% to 70.93%, and 51.45% on average when compared with the original HEVC test model. The Bjøntegaard delta peak signal-to-noise ratio and Bjøntegaard delta bit rate are -0.061 dB and 1.98% on average, which is negligible. The overall performance of the proposed algorithm outperforms those of the state-of-the-art schemes.
Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Hui Yuan 0001, Zhaoqing Pan, Long Xu 0001
IEEE Trans. Image Process.4
2015 Rate Distortion Optimized Inter-View Frame Level Bit Allocation Method for MV-HEVC
abstract
In multi-view video coding, since inter-view prediction has been adopted as an important coding tool which could improve coding efficiency greatly, inter-view dependency is inevitable, i.e., the distortion of the reference view (RV) picture could be propagated to the non-reference view (NRV) pictures . Therefore, in order to achieve higher coding efficiency , the inter-view dependency must be taken into account for inter-view bit allocation. In this paper, the inter-view dependency is analyzed in detail, and a rate-distortion (RD) model for NRVs is derived by taking the distortion of RV into account. Based on the derived RD model, the inter-view bit allocation is represented as a mathematical problem with an analytic form, and is solved by a convex optimization (Lagrangian Multiplier) method. Experimental results demonstrate that the RD performance and the inter-view quality consistency of the proposed method is better than existing methods, while the complexity of the proposed method is comparable with the existing methods.
Hui Yuan 0001, Sam Kwong, Xu Wang 0006, Wei Gao 0003, Yun Zhang 0002
IEEE Trans. Multim.1
2014 Fast Coding Tree Unit depth decision for high efficiency video coding
abstract
High Efficiency Video Coding (HEVC) is the latest video coding standard, which adapts quadtree structure based Coding Tree Unit (CTU) to improve the coding efficiency. In HEVC encoding process, the CTU is recursively partitioned into coding units according to the quadtree depth. This technique increases the coding efficiency of HEVC, however, the achieved coding efficiency comes at the cost of high computational complexity. In this paper, we propose a fast C-TU quadtree depth decision algorithm to reduce the computational complexity of HEVC. Firstly, based on the best C-TU depth correlation among spatial and temporal neighboring CTUs, an early quadtree depth 0 decision algorithm is proposed. Then, according to the correlation between the prediction unit mode and the best CTU depth selection, a quadtree depth 3 skipped decision algorithm is proposed. Experimental results show that the proposed algorithm can achieve 40% on average encoding time saving, while maintaining a comparable rate-distortion performance.
Zhaoqing Pan, Sam Kwong, Yun Zhang 0002, Jianjun Lei 0001, Hui Yuan 0001
ICIP5
2014 A Novel Distortion Model and Lagrangian Multiplier for Depth Maps Coding
abstract
In three-dimensional videos (3-DV) coding systems, depth maps are not used for viewing but for rendering virtual views. Therefore, the traditional rate distortion criterion (including distortion criterion, and Lagrangian multiplier) is not suitable for depth map coding. In order to design an effective rate distortion criterion for depth maps, the relationship between the distortion of synthesized virtual view and the coding error of depth maps is analyzed in detail. Through the analysis, a polynomial model revealing the relationship between the coding error of depth maps and the distortion of synthesized virtual view is derived. Model parameters are estimated by utilizing camera parameters and features of the texture video corresponding to the depth map. Based on the model, a virtual view-based Lagrangian multiplier for depth map coding is also proposed. Experimental results demonstrated the accuracy of the model. The squared correlation coefficients between the actual distortion of virtual view and the estimated distortion are all larger than 0.98 for all tested sequences. When incorporating the proposed model and Lagrangian multiplier into the mode decision procedure of joint model version 18.5 (JM18.5) of H.264/AVC, a maximum 0.470 dB BD PSNR and an average 0.251 dB BD PSNR can be achieved.
Hui Yuan 0001, Sam Kwong, Jiande Sun 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Unequally Weighted Video Hashing for Copy Detection
Jiande Sun 0001, Hui Yuan 0001, Xiaocui Liu
MMM (1)3
2012 Affine Model Based Motion Compensation Prediction for Zoom
abstract
Zoom motion is classified into two categories, i.e., global zoom motion and local zoom motion. A simple affine motion model with four parameters is utilized to describe zoom motion efficiently based on the analyses of camera imaging principles. Based on the motion model, a basic candidate motion vector (BCMV) of a block could be derived when model parameters are confirmed. Then a set of candidate motion vectors (CMVs) could be obtained by modifying the BCMV. Thereafter, template matching is used to choose the optimal CMV (OCMV). Finally, the block is coded with the optimal CMV as an independent mode, and a rate distortion (RD) criterion is used to determine whether to use the mode or not. Experimental results demonstrate that by implementing the proposed method into Key Technology Area test platform version 2.6r1 (KTA2.6r1), a maximum -21.99% and average -8.72% bit rate savings can be achieved for videos involving zoom motion, while maintaining the same quality (evaluated by PSNR) of reconstructed videos when IPPPP coding structure is used. When Hierarchical B coding structure is employed, the maximum and average bit rate savings are - 10.48% and - 5.575% when the qualities of reconstructed videos remain unchanged. Besides, for videos involving camera rotation, translation, etc., an average -2.04% bit rate savings could also be achieved; while for videos containing common motions, an average -1.19% bit rate saving could be achieved at the same quality of reconstructed videos.
Hui Yuan 0001, Jiande Sun 0001, Hechao Liu
IEEE Trans. Multim.1
2011 Model-Based Joint Bit Allocation Between Texture Videos and Depth Maps for 3-D Video Coding
abstract
In 3-D video coding, texture videos and depth maps need to be jointly coded. The distortion of texture videos and depth maps can be propagated to the synthesized virtual views. Besides coding efficiency of texture videos and depth maps, joint bit allocation between texture videos and depth maps is also an important research issue in 3-D video coding. First, we present comprehensive analyses on the impacts of the compression distortion of texture videos and depth maps on the quality of the virtual views, and then derive a concise distortion model for the synthesized virtual views. Based on this model, the joint bit allocation problem is formulated as a constrained optimization problem, and is solved by using the Lagrangian multiplier method. Experimental results demonstrate the high accuracy of the derived distortion model. Meanwhile, the rate-distortion (R-D) performance of the proposed algorithm is close to those of search-based algorithms which can give the best R-D performance, while the complexity of the proposed algorithm is lower than that of search-based algorithms. Moreover, compared with the bit allocation method using fixed texture and depth bits ratio (5:1), a maximum 1.2 dB gain can be achieved by the proposed algorithm.
Hui Yuan 0001, Yilin Chang, Junyan Huo, Fuzheng Yang 0001, Zhaoyang Lu
IEEE Trans. Circuits Syst. Video Technol.1
2010 Model Based Motion Vector Predictor for Zoom Motion
abstract
As zoom motion is common in video applications, a linear motion model is derived to describe zoom motion based on the analyses of camera imaging principles. Based on the motion model, a motion vector predictor for videos involving zoom motion is proposed. A rate distortion (RD) criterion is used to choose the optimal motion vector predictor between the one utilized in H.264/AVC and the one derived from the linear motion model. Experimental results demonstrate that by implementing the proposed method into Key Technology Area test platform version 2.2r1(KTA2.2r1), the maximum and average bit rate savings can be achieved as high as 7.66% and 4.90% respectively, while maintaining the same quality of reconstructed videos.
Hui Yuan 0001, Yilin Chang, Zhaoyang Lu, Yanzhuo Ma
IEEE Signal Process. Lett.1