Qi Yang 0003

dblp:22/2344-3 · DBLP profile ↗
← Back
42ranked-venue papers
10as first author
42since 2021 · last 2026
0000-0002-4274-3457ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 7 first-author · 39 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lightweight 3D Gaussian Splatting Compression via Video Codec
abstract
Current video-based GS compression methods rely on using Parallel Linear Assignment Sorting (PLAS) to convert 3D GS into smooth 2D maps, which are computationally expensive and time-consuming, limiting the application of GS on lightweight devices. In this paper, we propose a Lightweight 3D Gaussian Splatting (GS) Compression method based on Video codec (LGSCV). First, a two-stage Morton scan is proposed to generate blockwise 2D maps that are friendly for canonical video codecs in which the coding units (CU) are square blocks. A 3D Morton scan is used to permute GS primitives, followed by a 2D Morton scan to map the ordered GS primitives to 2D maps in a blockwise style. However, although the blockwise 2D maps report close performance to the PLAS map in high-bitrate regions, they show a quality collapse at medium-to-low bitrates. Therefore, a principal component analysis (PCA) is used to reduce the dimensionality of spherical harmonics (SH), and a MiniPLAS, which is flexible and fast, is designed to permute the primitives within certain block sizes. Incorporating SH PCA and MiniPLAS leads to a significant gain in rate-distortion (RD) performance, especially at medium and low bitrates. MiniPLAS can also guide the setting of the codec CU size configuration and significantly reduce encoding time. Experimental results on the MPEG dataset demonstrate that the proposed LGSCV achieves over 20% RD gain compared with state-of-the-art methods, while reducing 2D map generation time to approximately 1 second and cutting encoding time by 50%. The code is available at https://github.com/Qi-Yangsjtu/LGSCV.
Qi Yang 0003, Geert Van der Auwera, Zhu Li 0001
DCC1
2026 Not All Views Matter: A View-Selective Approach to Point Cloud Perceptual Quality Assessment
Yida Xiang, Bingyang Cui, Kaifa Yang, Qi Yang 0003, Yiling Xu
QoMEX5
2026 Light4GS: Lightweight Compact 4D Gaussian Splatting Generation via Context Model
Mufan Liu, Qi Yang 0003, Zhenlong Yuan, Zhu Li 0001, Yiling Xu, Yunfeng Guan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 From Images to Point Clouds: An Efficient Solution for Cross-Media Blind Quality Assessment Without Annotated Training
abstract
We present a novel quality assessment method which can predict the perceptual quality of point clouds from new scenes without available annotations by leveraging the rich prior knowledge in images, called the Distribution-Weighted Image-Transferred Point Cloud Quality Assessment (DWIT-PCQA). Recognizing the human visual system (HVS) as the decision-maker in quality assessment regardless of media types, we can emulate the evaluation criteria for human perception via neural networks and further transfer the capability of quality prediction from images to point clouds by leveraging the prior knowledge in the images. Specifically, domain adaptation (DA) can be leveraged to bridge the images and point clouds by aligning feature distributions of the two media in the same feature space. However, the different manifestations of distortions in images and point clouds make feature alignment a difficult task. To reduce the alignment difficulty and consider the different distortion distributions during alignment, we have derived formulas to decompose the optimization objective of the conventional DA into two suboptimization functions with distortion as a transition. Specifically, through network implementation, we propose the distortion-guided biased feature alignment which integrates existing/estimated distortion distribution into the adversarial DA framework, emphasizing common distortion patterns during feature alignment. Besides, we propose the quality-aware feature disentanglement to mitigate the destruction of the mapping from features to quality during alignment with biased distortions. Experimental results demonstrate that our proposed method exhibits reliable performance compared to general blind PCQA methods without needing point cloud annotations.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu, Le Yang 0001, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 On the Efficient Adaptive Streaming of 3D Gaussian Splatting Over Dynamic Networks
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising representation for immersive media. Its explicit splat-based structure offers high visual quality and real-time rendering, making it particularly suitable for six degrees of freedom streaming applications. However, its deployment in practical streaming scenarios is still limited due to several key challenges such as the large data volume, and insufficient support for dynamic bitrate adaptation under fluctuating network conditions. This paper presents an efficient 3DGS streaming framework that operates directly on pre-generated 3DGS models without retraining or fine-tuning. First, a training-free perceptual pruning method, which removes visually redundant Gaussians according to the human visual system metrics, is introduced. The resulting 3DGS is then encoded into a compact representation using the extended 3D codecs, exploiting its point-based structure. Next, we build a scene-specific bitrate ladder through analyzing the trade-off between resolution, bitrate, and perceptual quality. This enables efficient and fine-grained representation selection. Finally, a progressive streaming mechanism is developed. It is driven by a reinforcement learning scheduler that adaptively decides whether to download new content or enhance previously buffered content based on real-time network feedback. Experiments on real-world 3DGS datasets and bandwidth traces show that the proposed method evidently improves the quality of experience and streaming efficiency in various network scenarios.
Mufan Liu, Qi Yang 0003, Le Yang 0001, Yiling Xu
IEEE Trans. Circuits Syst. Video Technol.3
2026 Anchor-Driven Compact Gaussian Splatting for Dynamic Scene Reconstruction
abstract
Existing 4D Gaussian Splatting methods typically rely on per-Gaussian deformation from a canonical space to target frames, which overlooks the strong redundancy among spatially and temporally adjacent Gaussian primitives and leads to suboptimal efficiency. To address this limitation, we propose ADC-GS++, an anchor-driven compact Gaussian splatting framework for efficient and high-quality dynamic scene reconstruction. Specifically, ADC-GS++ organizes Gaussian primitives into an anchor-based canonical representation, enabling attribute sharing across local regions. To efficiently model dynamic scenes, we introduce a static-dynamic decomposition mechanism and further employ a coarse-to-fine deformation strategy driven by dynamic anchors at multiple granularities. In addition, a unified rate-distortion optimization is adopted to achieve a balanced trade-off between storage efficiency and reconstruction fidelity. Furthermore, a temporal significance-based anchor refinement strategy is employed to dynamically grow and prune anchors, allowing robust adaptation to complex and large-scale motions. Extensive experiments on multiple real-world dynamic scene datasets demonstrate that ADC-GS++ significantly improves rendering speed over deformation-based approaches by 300%-700%, while maintaining competitive rendering quality. Moreover, ADC-GS++ achieves a more favorable rate-distortion trade-off, resulting in substantially reduced storage consumption across different bitrate settings.
Qi Yang 0003, Mufan Liu, Yiling Xu, Zhu Li 0001
IEEE Trans. Vis. Comput. Graph.2
2026 Deformable 2D Gaussian Splatting for Efficient Wireless Radiance Field Rendering
abstract
Modeling the wireless radiance field (WRF) is fundamental to modern communication systems, enabling key tasks such as localization, sensing, and channel estimation. Traditional approaches, which rely on empirical formulas or physical simulations, often suffer from limited accuracy or require strong scene priors. Recent neural radiance field (NeRF)-based methods improve reconstruction fidelity through differentiable volumetric rendering, but their reliance on computationally expensive multilayer perceptron (MLP) queries hinders real-time deployment. To overcome these challenges, we introduce Gaussian splatting (GS) to the wireless domain, leveraging its efficiency in modeling optical radiance fields to enable compact and accurate WRF reconstruction. Specifically, we propose SwiftWRF, a deformable 2D Gaussian splatting framework that synthesizes WRF spectra at arbitrary positions under single-sided transceiver mobility. SwiftWRF employs CUDA-accelerated rasterization to render spectra at over 100 k FPS and uses the lightweight MLP to model the deformation of 2D Gaussians, effectively capturing mobility-induced WRF variations. In addition to novel spectrum synthesis, the efficacy of SwiftWRF is further underscored in its applications in angle-of-arrival (AoA) and received signal strength indicator (RSSI) prediction. Experiments conducted on both real-world and synthetic indoor scenes demonstrate that SwiftWRF can reconstruct WRF spectra up to 500x faster than existing state-of-the-art methods, while significantly enhancing its signal quality.
Mufan Liu, Cixiao Zhang, Qi Yang 0003, Yiling Xu, Yin Xu 0001, Shu Sun 0001, Mingzeng Dai, Yunfeng Guan 0001
IEEE Trans. Vis. Comput. Graph.3
2025 A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission
abstract
The large volume of data from the point cloud brings significant demands on network bandwidth. However, the current transmission framework only considers using lossy compression to control the size of data, while ignoring visually redundant information due to the setting of rendering devices. Based on the fact that point overlapping might occur for the case that a dense point cloud is rendered on a relatively low resolution 2D monitor, we propose a novel quality-aware sampling framework for point cloud transmission. When a target visual quality is determined, an optimal sampling module is designed to remove overlapped points with the help of a simple but effective quality model. By taking into account the impact of multiple factors (i.e., sampling, lossy compression, and client rendering resolution), this quality model can predict the final perceptual quality in the client. Based on a newly constructed dataset which consists of 420 samples, experiment results show that the proposed transmission framework can significantly reduce bandwidth cost (e.g., 6.10% to 84.43%) and processing time (e.g., 8.99% to 92.53%) without introducing noticeable distortion under certain rendering conditions, thus achieving higher bandwidth utilization and better real-time performance.
Puyue Hou, Qi Yang 0003, Yue Li 0015, Jianchao Yang, Yiling Xu, Tiejun Huang 0001
ICASSP2
2025 A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
abstract
3D Gaussian Splatting (GS) demonstrates excellent rendering quality and generation speed in novel view synthesis. However, substantial data size poses challenges for storage and transmission, making 3D GS compression an essential technology. Current 3D GS compression research primarily focuses on developing more compact scene representations, such as converting explicit 3D GS data into implicit forms. In contrast, compression of the GS data itself has hardly been explored. To address this gap, we propose a Hierarchical GS Compression (HGSC) technique. Initially, we prune unimportant Gaussians based on importance scores derived from both global and local significance, effectively reducing redundancy while maintaining visual quality. An Octree structure is used to compress 3D positions. Based on the 3D GS Octree, we implement a hierarchical attribute compression strategy by employing a KD-tree to partition the 3D GS into multiple blocks. We apply farthest point sampling to select anchor primitives within each block and others as non-anchor primitives with varying Levels of Details (LoDs). Anchor primitives serve as reference points for predicting non-anchor primitives across different LoDs to reduce spatial redundancy. For anchor primitives, we use the region adaptive hierarchical transform to achieve near-lossless compression of various attributes. For non-anchor primitives, each is predicted based on the k-nearest anchor primitives. To further minimize prediction errors, the reconstructed LoD and anchor primitives are combined to form new anchor primitives to predict the next LoD. Our method notably achieves superior compression quality and a significant data size reduction of over 4.5× compared to the state-of-the-art compression method on small scenes datasets.
Qi Yang 0003, Yiling Xu, Zhu Li 0001
ICASSP3
2025 LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression
abstract
Existing AI-based point cloud compression methods struggle with dependence on specific training data distributions, which limits their real-world deployment. Implicit Neural Representation (INR) methods solve the above problem by encoding overfitted network parameters to the bitstream, resulting in more distribution-agnostic results. However, due to the limitation of encoding time and decoder size, current INR based methods only consider lossy geometry compression. In this paper, we propose the first INR based lossless point cloud geometry compression method called Lossless Implicit Neural Representations for Point Cloud Geometry Compression (LINR-PCGC). To accelerate encoding speed, we design a group of point clouds level coding framework with an effective network initialization strategy, which can reduce around 60% encoding time. A lightweight coding network based on multiscale SparseConv, consisting of scale context extraction, child node prediction, and model compression modules, is proposed to realize fast inference and compact decoder size. Experimental results show that our method consistently outperforms traditional and AI-based methods: for example, with the convergence time in the MVUB dataset, our method reduces the bitstream by approximately 21.21% compared to G-PCC TMC13v23 and 21.95% compared to SparsePCGC. Our project can be seen on https://huangwenjie2023.github.io/LINR-PCGC/.
Qi Yang 0003, Shuting Xia, Yiling Xu, Zhu Li 0001
ICCV2
2025 Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-To-3D Generation
abstract
Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) Existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation dimensions. ii) Previous evaluation metrics only focus on a single aspect (e.g., text-3D alignment) and fail to perform multi-dimensional quality assessment. To address these problems, we first propose a comprehensive benchmark named MATE-3D. The benchmark contains eight well-designed prompt categories that cover single and multiple object generation, resulting in 1,280 generated textured meshes. We have conducted a large-scale subjective experiment from four different evaluation dimensions and collected 107,520 annotations, followed by detailed analyses of the results. Based on MATE-3D, we propose a novel quality evaluator named HyperScore. Utilizing hypernetwork to generate specified mapping functions for each evaluation dimension, our metric can effectively perform multi-dimensional quality assessment. HyperScore presents superior performance over existing metrics on MATE-3D, making it a promising metric for assessing and improving text-to-3D generation. The project is available at https://mate-3d.github.io/.
Bingyang Cui, Qi Yang 0003, Zhu Li 0001, Yiling Xu
ICCV3
2025 DPCD: A Quality Assessment Database for Dynamic Point Clouds
abstract
Recently, the advancements in Virtual/Augmented Reality (VR/AR) have driven the demand for Dynamic Point Clouds (DPC). Unlike static point clouds, DPCs are capable of capturing temporal changes within objects or scenes, offering a more accurate simulation of the real world. While significant progress has been made in the quality assessment research of static point cloud, little study has been done on Dynamic Point Cloud Quality Assessment (DPCQA), which hinders the development of quality-oriented applications, such as interframe compression and transmission in practical scenarios. In this paper, we introduce a large-scale DPCQA database, named DPCD, which includes 15 reference DPCs and 525 distorted DPCs from seven types of lossy compression and noise distortion. By rendering these samples to Processed Video Sequences (PVS), a comprehensive subjective experiment is conducted to obtain Mean Opinion Scores (MOS) from 21 viewers for analysis. The characteristic of contents, impact of various distortions, and accuracy of MOSs are presented to validate the heterogeneity and reliability of the proposed database. Furthermore, we evaluate the performance of several objective metrics on DPCD. The experiment results show that DPCQA is more challenge than that of static point cloud. The DPCD, which serves as a catalyst for new research endeavors on DPCQA, is publicly available at https://huggingface.co/datasets/Olivialyt/DPCD.
Qi Yang 0003, Yiling Xu, Zhu Li 0001, Ye-Kui Wang
ICME3
2025 HybridGS: High-Efficiency Gaussian Splatting Data Compression using Dual-Channel Sparse Representation and Point Cloud Encoder
abstract
Most existing 3D Gaussian Splatting (3DGS) compression schemes focus on producing compact 3DGS representation via implicit data embedding. They have long encoding and decoding times and highly customized data format, making it difficult for widespread deployment. This paper presents a new 3DGS compression framework called HybridGS, which takes advantage of both compact generation and standardized point cloud data encoding. HybridGS first generates compact and explicit 3DGS data. A dual-channel sparse representation is introduced to supervise the primitive position and feature bit depth. It then utilizes a canonical point cloud encoder to carry out further data compression and form standard output bitstreams. A simple and effective rate control scheme is proposed to pivot the interpretable data compression scheme. HybridGS does not include any modules aimed at improving 3DGS quality during generation. But experiment results show that it still provides comparable reconstruction performance against state-of-the-art methods, with evidently faster encoding and decoding speed. The code is publicly available at https://github.com/Qi-Yangsjtu/HybridGS .
Qi Yang 0003, Le Yang 0001, Geert Van der Auwera, Zhu Li 0001
ICML1
2025 ADC-GS: Anchor-Driven Deformable and Compressed Gaussian Splatting for Dynamic Scene Reconstruction
abstract
Existing 4D Gaussian Splatting methods rely on per-Gaussian deformation from a canonical space to target frames, which overlooks redundancy among adjacent Gaussian primitives and result in suboptimal performance. To address this limitation, we propose Anchor-Driven Deformable and Compressed Gaussian Splatting (ADC-GS), a compact and efficient representation for dynamic scene reconstruction. Specifically, ADC-GS organizes Gaussian primitives into an anchor-based structure within the canonical space, enhanced by a temporal significance-based anchor refinement strategy. To reduce deformation redundancy, ADC-GS introduces a hierarchical coarse-to-fine pipeline that captures motions at varying granularities. Moreover, a rate-distortion optimization is adopted to achieve an optimal balance between bitrate consumption and representation fidelity. Experimental results demonstrate that ADC-GS outperforms the per-Gaussian deformation approaches in rendering speed by 300%-800% while achieving state-of-the-art storage efficiency without compromising rendering quality. The code is released at https://github.com/H-Huang774/ADC-GS.git.
Qi Yang 0003, Mufan Liu, Yiling Xu, Zhu Li 0001
IJCAI2
2025 3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compression
abstract
3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual fidelity, but its significant storage demands limit practical deployment, prompting recent methods to integrate compression modules into 3DGS. However, these 3DGS generative compression techniques introduce unique distortions that lack systematic quality assessment research. To this end, we establish 3DGS-VBench, a large-scale Video Quality Assessment (VQA) dataset and benchmark with 660 compressed 3DGS models and video sequences generated from 11 scenes across 6 representative 3DGS compression algorithms with systematically designed parameter levels. With annotations from 50 participants, we obtain MOS scores with outlier removal and validate dataset reliability. We benchmark 6 3DGS compression algorithms on storage efficiency and visual quality, and evaluate 15 quality assessment metrics. Our dataset enables specialized VQA model training for 3DGS. The dataset is available at https://github.com/YukeXing/3DGS-VBench.
Yuke Xing, William Gordon, Qi Yang 0003, Kaifa Yang, Yiling Xu
VCIP3
2025 Asynchronous Feedback Network for Perceptual Point Cloud Quality Assessment
abstract
Recent years have witnessed the success of the deep learning-based technique in research of no-reference point cloud quality assessment (NR-PCQA). For a more accurate quality prediction, many previous studies have attempted to capture global and local features in a bottom-up manner, but ignored the interaction and promotion between them. To solve this problem, we propose a novel asynchronous feedback quality prediction network (AFQ-Net). Motivated by human visual perception mechanisms, AFQ-Net employs a dual-branch structure to deal with global and local features, simulating the left and right hemispheres of the human brain, and constructs a feedback module between them. Specifically, the input point clouds are first fed into a transformer-based global encoder to generate the attention maps that highlight these semantically rich regions, followed by being merged into the global feature. Then, we utilize the generated attention maps to perform dynamic convolution for different semantic regions and obtain the local feature. Finally, a coarse-to-fine strategy is adopted to merge the two features into the final quality score. We conduct comprehensive experiments on three datasets and achieve superior performance over the state-of-the-art approaches on all of these datasets. The code will be available athttps://github.com/zhangyujie-1998/AFQ-Net
Qi Yang 0003, Ziyu Shan, Yiling Xu
IEEE Trans. Circuits Syst. Video Technol.2
2025 GeodesicPSIM: Predicting the Quality of Static Mesh With Texture Map via Geodesic Patch Similarity
abstract
Static meshes with texture maps have attracted considerable attention in both industrial manufacturing and academic research, leading to an urgent requirement for effective and robust objective quality evaluation. However, current model-based static mesh quality metrics (i.e., metrics that directly use the raw data of the static mesh to extract features and predict the quality) have obvious limitations: most of them only consider geometry information, while color information is ignored, and they have strict constraints for the meshes' geometrical topology. Other metrics, such as image-based and point-based metrics, are easily influenced by the prepossessing algorithms, e.g., projection and sampling, hampering their ability to perform at their best. In this paper, we propose Geodesic Patch Similarity (GeodesicPSIM), a novel model-based metric to accurately predict human perception quality for static meshes. After selecting a group keypoints, 1-hop geodesic patches are constructed based on both the reference and distorted meshes cleaned by an effective mesh cleaning algorithm. A two-step patch cropping algorithm and a patch texture mapping module refine the size of 1-hop geodesic patches and build the relationship between the mesh geometry and color information, resulting in the generation of 1-hop textured geodesic patches. Three types of features are extracted to quantify the distortion: patch color smoothness, patch discrete mean curvature, and patch pixel color average and variance. To the best of our knowledge, GeodesicPSIM is the first model-based metric especially designed for static meshes with texture maps. GeodesicPSIM provides state-of-the-art performance in comparison with image-based, point-based, and video-based metrics on a newly created and challenging database. We also prove the robustness of GeodesicPSIM by introducing different settings of hyperparameters. Ablation studies also exhibit the effectiveness of three proposed features and the patch cropping algorithm. The code is available at https://multimedia.tencent.com/resources/GeodesicPSIM.
Qi Yang 0003, Joël Jung, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Image Process.1
2025 Textured Mesh Quality Assessment Using Geometry and Color Field Similarity
abstract
Textured mesh quality assessment (TMQA) is critical for various 3D mesh applications. However, existing TMQA methods often struggle to provide accurate and robust evaluations. Motivated by the effectiveness of fields in representing both 3D geometry and color information, we propose a novel point-based TMQA method called field mesh quality metric (FMQM). FMQM utilizes signed distance fields and a newly proposed color field named nearest surface point color field to realize effective mesh feature description. Four features related to visual perception are extracted from the geometry and color fields: geometry similarity, geometry gradient similarity, space color distribution similarity, and space color gradient similarity. Experimental results on three benchmark datasets demonstrate that FMQM outperforms state-of-the-art (SOTA) TMQA metrics. Furthermore, FMQM exhibits low computational complexity, making it a practical and efficient solution for real-world applications in 3D graphics and visualization.
Kaifa Yang, Qi Yang 0003, Yiling Xu, Zhu Li 0001
IEEE Trans. Vis. Comput. Graph.2
2025 TDMD: A Database for Dynamic Color Mesh Quality Assessment Study
abstract
Dynamic colored meshes (DCM) are widely used in various applications. However, this kind of meshes may undergo different processes, such as compression or transmission, which can distort them and degrade their quality. To facilitate the development of objective metrics for DCMs and study the influence of typical distortions on their perception, we create the Tencent - Dynamic colored Mesh Database (TDMD) containing eight reference DCM objects with six typical distortions. Using processed video sequences (PVS) derived from the DCM, we conduct a large-scale subjective experiment that resulted in 303 distorted DCM samples with mean opinion scores, making the TDMD the largest available DCM database to our knowledge. This database enables us to study the impact of different types of distortion on human perception and offers recommendations for DCM compression and related tasks. Additionally, we have evaluated three types of state-of-the-art objective metrics on the TDMD, including image-based, point-based, and video-based metrics, on the TDMD. Our experimental results highlight the strengths and weaknesses of each metric, and we provide suggestions about the selection of metrics in practical DCM applications.
Qi Yang 0003, Joël Jung, Timon Deschamps, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
abstract
No-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference, which have achieved tremendous improvements due to the utilization of deep neural networks. However, learning-based NR-PCQA methods suffer from the scarcity of labeled data and usually perform suboptimally in terms of generalization. To solve the problem, we propose a novel contrastive pre-training framework tailored for PCQA (CoPA), which enables the pre-trained model to learn quality-aware representations from unlabeled data. To obtain anchors in the representation space, we project point clouds with different distortions into images and randomly mix their local patches to form mixed images with multiple distortions. Utilizing the generated anchors, we constrain the pretraining process via a quality-aware contrastive loss following the philosophy that perceptual quality is closely related to both content and distortion. Furthermore, in the model fine-tuning stage, we propose a semantic-guided multi-view fusion module to effectively integrate the features of projected images from multiple perspectives. Extensive experiments show that our method outperforms the state-of-the-art PCQA methods on popular benchmarks. Further investigations demonstrate that CoPA can also benefit existing learning-based PCQA models.
Ziyu Shan, Qi Yang 0003, Haichen Yang, Yiling Xu, Jenq-Neng Hwang, Xiaozhong Xu, Shan Liu 0001
CVPR3
2024 SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map
abstract
In recent years, static meshes with texture maps have become one of the most prevalent digital representations of 3D shapes in various applications, such as animation, gaming, medical imaging, and cultural heritage applications. However, little research has been done on the quality assessment of textured meshes, which hinders the development of quality-oriented applications, such as mesh compression and enhancement. In this paper, we create a large-scale textured mesh quality assessment database, namely SJTU-TMQA, which includes 21 reference meshes and 945 distorted samples. The meshes are rendered into processed video sequences and then conduct subjective experiments to obtain mean opinion scores (MOS). The diversity of content and accuracy of MOS has been shown to validate its heterogeneity and reliability. The impact of various types of distortion on human perception is demonstrated. 13 state-of-the-art objective metrics are evaluated on SJTU-TMQA. The results report the highest correlation is around 0.6, indicating the need for more effective objective metrics. The SJTU-TMQA is available at https://ccccby.github.io
Bingyang Cui, Qi Yang 0003, Kaifa Yang, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
ICASSP2
2024 MS-GeodesicPSIM: Predicting the Quality of Static Mesh with Texture Map via multi-scale Geodesic Patch Similarity
abstract
To address the mesh quality assessment (MQA) problem, GeodesicP-SIM was proposed by jointly considering geometry and color features, demonstrating compelling performance in multiple benchmarks.However, GeodesicPSIM does not consider the multi-scale characteristics of human perception.To better mimic human subjective perception, we proposed a multi-scale MQA model called multi-scale Geodesic Patch Similarity (MS-GeodesicPSIM).Firstly, inspired by the multi-scale processing methods used in image and point cloud analysis, we propose a novel multi-scale representation of textured meshes based on mesh simplification techniques.Secondly, we extend GeodesicPSIM into a multi-scale version leveraging the proposed multi-scale representation.Specifically, we construct a multi-scale representation for the reference and distorted meshes, followed by fusing the results of GeodesicPSIM at different scales to obtain an overall quality score.Experimental results demonstrate the superior performance of the proposed MS-GeodesicPSIM compared to the single-scale GeodesicPSIM and other MQA metrics on three large and independent databases.Ablation studies further confirm that MS-GeodesicPSIM is robust to different model hyperparameter settings.The code for MS-GeodesicPSIM is available at https://github.com/ccccby/MS-GeodesicPSIM
Bingyang Cui, Qi Yang 0003, Yiling Xu
MMAsia3
2024 A Benchmark for Gaussian Splatting Compression and Quality Assessment Study
Qi Yang 0003, Kaifa Yang, Yuke Xing, Yiling Xu, Zhu Li 0001
MMAsia1
2024 Cross-Modal Distortion Approximation for Fast Bit Allocation of Video-Based Point Cloud Compression
abstract
In video-based point cloud compression (V-PCC), the optimal allocation of the total bitrate between geometry and color is a challenging but rewarding problem. Existing bit allocation approaches leverage statistical models to describe the rate and distortion of geometry and color information as functions of V-PCC quantization steps. However, to obtain the parameters of statistical models, these methods need to perform pre-coding for input point clouds for multiple times, resulting in high computational complexity. Consequently, the capability of these methods for practical application is limited. To address this problem, we derive the rate and distortion models based on projected images produced in the V-PCC encoding process, transforming the expensive point cloud pre-coding process into an efficient image pre-coding process. By utilizing the image-based distortion and rate models, the bit allocation problem is further formulated as a constrained convex optimization problem. Experimental results demonstrate that the proposed method exhibits significantly lower time complexity and higher rate-distortion performance compared to the existing methods.
Haichen Yang, Qi Yang 0003, Ziyu Shan, Yiling Xu, Yunfeng Guan 0001
MMSP3
2024 Inter-Frame Coding for Dynamic Meshes via Coarse-to-Fine Anchor Mesh Generation
abstract
In the current Video-based Dynamic Mesh Coding (V-DMC) standard, inter-frame coding is restricted to mesh frames with constant topology. Consequently, temporal redundancy is not fully leveraged, resulting in suboptimal compression efficacy. To address this limitation, this paper introduces a novel coarse-to-fine scheme to generate anchor meshes for frames with time-varying topology. Initially, we generate a coarse anchor mesh using an octree-based nearest neighbor search. Motion estimation compensates for regions with significant motion changes during this process. However, the quality of the coarse mesh is low due to its suboptimal vertices. To enhance details, the fine anchor mesh is further optimized using the Quadric Error Metrics (QEM) algorithm to calculate more precise anchor points. The inter-frame anchor mesh generated herein retains the connectivity of the reference base mesh, while concurrently preserving superior quality. Experimental results show that our method achieves 7.2% ∼ 10.3% BD-rate gain compared to the existing V-DMC test model version 7.
Lizhi Hou, Qi Yang 0003, Yiling Xu
VCIP3
2024 Differentiable Low-computation Global Correlation Loss for Monotonicity Evaluation in Quality Assessment
abstract
In this paper, we propose a global monotonicity consistency training strategy for quality assessment, which includes a differentiable, low-computation monotonicity evaluation loss function and a global perception training mechanism. Specifically, unlike conventional ranking loss and linear programming approaches that indirectly implement the Spearman rank-order correlation coefficient (SROCC) function, our method directly converts SROCC into a loss function by making the sorting operation within SROCC differentiable and functional. Furthermore, to mitigate the discrepancies between batch optimization during network training and global evaluation of SROCC, we introduce a memory bank mechanism. This mechanism stores gradient-free predicted results from previous batches and uses them in the current batch’s training to prevent abrupt gradient changes. We evaluate the performance of the proposed method on both images and point clouds quality assessment tasks, demonstrating performance gains in both cases.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu
VCIP2
2024 Explicit-NeRF-QA: A Quality Assessment Database for Explicit NeRF Model Compression
abstract
In recent years, Neural Radiance Fields (NeRF) have demonstrated significant advantages in representing and synthesizing 3D scenes. Explicit NeRF models facilitate the practical NeRF applications with faster rendering speed, and also attract considerable attention in NeRF compression due to its huge storage cost. To address the challenge of the NeRF compression study, in this paper, we construct a new dataset, called Explicit-NeRF-QA. We use 22 3D objects with diverse geometries, textures, and material complexities to train four typical explicit NeRF models across five parameter levels. Lossy compression is introduced during the model generation, pivoting the selection of key parameters such as hash table size for InstantNGP and voxel grid resolution for Plenoxels. By rendering NeRF samples to processed video sequences (PVS), a large scale subjective experiment with lab environment is conducted to collect subjective scores from 21 viewers. The diversity of content, accuracy of mean opinion scores (MOS), and characteristics of NeRF distortion are comprehensively presented, establishing the heterogeneity of the proposed dataset. The state-of-the-art objective metrics are tested in the new dataset. Best Pearson correlation, which is around 0.85, is collected from the full-reference objective metric. All tested no-reference metrics report very poor results with 0.4 to 0.6 correlations, demonstrating the need for further development of more robust no-reference metrics. The dataset, including NeRF samples, source 3D objects, multiview images for NeRF generation, PVSs, MOS, is made publicly available at the following location: https://github.com/YukeXing/Explicit-NeRF-QA.
Yuke Xing, Qi Yang 0003, Kaifa Yang, Yiling Xu, Zhu Li 0001
VCIP2
2024 Perception-Guided Quality Metric of 3D Point Clouds Using Hybrid Strategy
abstract
Full-reference point cloud quality assessment (FR-PCQA) aims to infer the quality of distorted point clouds with available references. Most of the existing FR-PCQA metrics ignore the fact that the human visual system (HVS) dynamically tackles visual information according to different distortion levels (i.e., distortion detection for high-quality samples and appearance perception for low-quality samples) and measure point cloud quality using unified features. To bridge the gap, in this paper, we propose a perception-guided hybrid metric (PHM) that adaptively leverages two visual strategies with respect to distortion degree to predict point cloud quality: to measure visible difference in high-quality samples, PHM takes into account the masking effect and employs texture complexity as an effective compensatory factor for absolute difference; on the other hand, PHM leverages spectral graph theory to evaluate appearance degradation in low-quality samples. Variations in geometric signals on graphs and changes in the spectral graph wavelet coefficients are utilized to characterize geometry and texture appearance degradation, respectively. Finally, the results obtained from the two components are combined in a non-linear method to produce an overall quality score of the tested point cloud. The results of the experiment on five independent databases show that PHM achieves state-of-the-art (SOTA) performance and offers significant performance improvement in multiple distortion environments. The code is publicly available at https://github.com/zhangyujie-1998/PHM.
Qi Yang 0003, Yiling Xu, Shan Liu 0001
IEEE Trans. Image Process.2
2024 Blind Quality Assessment of Dense 3D Point Clouds with Structure Guided Resampling
abstract
Objective quality assessment of three-dimensional (3D) point clouds is essential for the development of immersive multimedia systems in real-world applications. Despite the success of perceptual quality evaluation for 2D images and videos, blind/no-reference metrics are still scarce for 3D point clouds with large-scale irregularly distributed 3D points. Therefore, in this article, we propose an objective point cloud quality index with Structure Guided Resampling (SGR) to automatically evaluate the perceptually visual quality of dense 3D point clouds. The proposed SGR is a general-purpose blind quality assessment method without the assistance of any reference information. Specifically, considering that the human visual system is highly sensitive to structure information, we first exploit the unique normal vectors of point clouds to execute regional pre-processing that consists of keypoint resampling and local region construction. Then, we extract three groups of quality-related features, including (1) geometry density features, (2) color naturalness features, and (3) angular consistency features. Both the cognitive peculiarities of the human brain and naturalness regularity are involved in the designed quality-aware features that can capture the most vital aspects of distorted 3D point clouds. Extensive experiments on several publicly available subjective point cloud quality databases validate that our proposed SGR can compete with state-of-the-art full-reference, reduced-reference, and no-reference quality assessment algorithms.
Wei Zhou 0021, Qi Yang 0003, Qiuping Jiang, Guangtao Zhai, Weisi Lin
ACM Trans. Multim. Comput. Commun. Appl.2
2024 GPA-Net:No-Reference Point Cloud Quality Assessment With Multi-Task Graph Convolutional Network
abstract
With the rapid development of 3D vision, point cloud has become an increasingly popular 3D visual media content. Due to the irregular structure, point cloud has posed novel challenges to the related research, such as compression, transmission, rendering and quality assessment. In these latest researches, point cloud quality assessment (PCQA) has attracted wide attention due to its significant role in guiding practical applications, especially in many cases where the reference point cloud is unavailable. However, current no-reference metrics which based on prevalent deep neural network have apparent disadvantages. For example, to adapt to the irregular structure of point cloud, they require preprocessing such as voxelization and projection that introduce extra distortions, and the applied grid-kernel networks, such as Convolutional Neural Networks, fail to extract effective distortion-related features. Besides, they rarely consider the various distortion patterns and the philosophy that PCQA should exhibit shift, scaling, and rotation invariance. In this paper, we propose a novel no-reference PCQA metric named the Graph convolutional PCQA network (GPA-Net). To extract effective features for PCQA, we propose a new graph convolution kernel, i.e., GPAConv, which attentively captures the perturbation of structure and texture. Then, we propose the multi-task framework consisting of one main task (quality regression) and two auxiliary tasks (distortion type and degree predictions). Finally, we propose a coordinate normalization module to stabilize the results of GPAConv under shift, scale and rotation transformations. Experimental results on two independent databases show that GPA-Net achieves the best performance compared to the state-of-the-art no-reference PCQA metrics, even better than some full-reference metrics in some cases.
Ziyu Shan, Qi Yang 0003, Rui Ye 0001, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Vis. Comput. Graph.2
2024 TCDM: Transformational Complexity Based Distortion Metric for Perceptual Point Cloud Quality Assessment
abstract
The goal of objective point cloud quality assessment (PCQA) research is to develop quantitative metrics that measure point cloud quality in a perceptually consistent manner. Merging the research of cognitive science and intuition of the human visual system (HVS), in this article, we evaluate the point cloud quality by measuring the complexity of transforming the distorted point cloud back to its reference, which in practice can be approximated by the code length of one point cloud when the other is given. For this purpose, we first make space segmentation for the reference and distorted point clouds based on a 3D Voronoi diagram to obtain a series of local patch pairs. Next, inspired by the predictive coding theory, we utilize a space-aware vector autoregressive (SA-VAR) model to encode the geometry and color channels of each reference patch with and without the distorted patch, respectively. Assuming that the residual errors follow the multi-variate Gaussian distributions, the self-complexity of the reference and transformational complexity between the reference and distorted samples are computed using covariance matrices. Additionally, the prediction terms generated by SA-VAR are introduced as one auxiliary feature to promote the final quality prediction. The effectiveness of the proposed transformational complexity based distortion metric (TCDM) is evaluated through extensive experiments conducted on five public point cloud quality assessment databases. The results demonstrate that TCDM achieves state-of-the-art (SOTA) performance, and further analysis confirms its robustness in various scenarios.
Qi Yang 0003, Xiaozhong Xu, Le Yang 0001, Yiling Xu
IEEE Trans. Vis. Comput. Graph.2
2023 Improving Point Cloud Quality Metrics with Noticeable Possibility Maps
abstract
Point cloud quality assessment (PCQA) plays a vital role in the quality of experience (QoE) oriented data processing. To reflect the visual degradation introduced by various distortions, many PCQA metrics have been proposed in recent years. However, these metrics often take all distortions indiscriminately into account, ignoring the fact that some distortions are below the noticeable threshold and thus do not affect subjective perception. To solve this problem, we involve the characteristic of just noticeable difference (JND) into PCQA. Specifically, we first rotate the reference and distorted point cloud repeatedly to obtain multiple perspectives, then utilize the 2D JND models to derive the 3D noticeable possibility maps (NPM) to infer the possibility that the distortion is perceivable at each point. Utilizing the generated NPM, we modify current point-wise and structure-wise quality metrics to help them correlate better with subjective perception. Extensive experiments show the universal effectiveness of the proposed NPM in improving PCQA metrics. Code will be available at https://github.com/NekoooooOoi/NPM.
Qi Yang 0003, Yiling Xu, Jun Sun 0005, Shan Liu 0001
ICME3
2023 Exploring the Influence of View and Camera Path Selection for Dynamic Mesh Quality Assessment
abstract
With the development of 3D mesh processing and applications, the quality assessment of dynamic mesh sequences attracts more and more attention. One prevalent strategy for performing the subjective experiment and designing objective quality metrics is to convert the 3D dynamic mesh into 2D images or videos via projection and to collect subjective scores or calculate objective indexes based on these images or videos. In this paper, we study the influence of the view, or camera path selection for the projection, for both subjective and objective dynamic mesh quality assessment, and compare the performance of image-based metrics and point-based metrics corresponding to the collected subjective scores. First, we use the dynamic mesh sequences proposed by MPEG as anchors and generate videos corresponding to different coding configurations and different camera paths. Then, we conduct subjective experiments to collect the ground truth of mean opinion scores. Besides, we calculate the state-of-the-art objective metric scores for each sequence. We analyze the differences between subjective scores with respect to different camera paths and the correlation between subjective scores and objective metrics. The results show that different camera paths tend to generate close subjective perceptions and that the selection of views can influence some objective metrics.
Kaifa Yang, Qi Yang 0003, Joël Jung, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
ICME2
2023 Point Cloud Quality Assessment using 3D Saliency Maps
abstract
Point cloud quality assessment (PCQA) has become an appealing research field in recent days. Considering the importance of saliency detection in quality assessment, we propose an effective full-reference PCQA metric which makes an attempt to utilize the saliency information to facilitate quality prediction, called point cloud quality assessment using 3D saliency maps (PQSM). Specifically, we first propose a projectionbased point cloud saliency map generation method, in which depth information is introduced to better reflect the geometric characteristics of point clouds. Then, we construct point cloud local neighborhoods to derive three structural descriptors to indicate the geometry, color and saliency discrepancies. Finally, a saliency-based pooling strategy is proposed to generate the final quality score. Extensive experiments are performed on four independent PCQA databases. The results demonstrate that the proposed PQSM shows competitive performances compared to multiple state-of-the-art PCQA metrics.
Qi Yang 0003, Yiling Xu, Jun Sun 0005, Shan Liu 0001
VCIP3
2023 TSMD: A Database for Static Color Mesh Quality Assessment Study
abstract
Static meshes with texture map are widely used in modern industrial and manufacturing sectors, attracting considerable attention in the mesh compression community due to its huge amount of data. To facilitate the study of static mesh compression algorithm and objective quality metric, we create the Tencent – Static Mesh Dataset (TSMD) containing 42 reference meshes with rich visual characteristics. 210 distorted samples are generated by the lossy compression scheme developed for the Call for Proposals on polygonal static mesh coding, released on June 23 by the Alliance for Open Media Volumetric Visual Media group. Using processed video sequences, a large-scale, crowdsourcing-based, subjective experiment was conducted to collect subjective scores from 74 viewers. The dataset undergoes analysis to validate its sample diversity and Mean Opinion Scores (MOS) accuracy, establishing its heterogeneous nature and reliability. State-of-the-art objective metrics are evaluated on the new dataset. Pearson and Spearman correlations around 0.75 are reported, deviating from results typically observed on less heterogeneous datasets, demonstrating the need for further development of more robust metrics. The TSMD, including meshes, PVSs, bitstreams, and MOS, is made publicly available at the following location: https://multimedia.tencent.com/resources/tsmd.
Qi Yang 0003, Joël Jung, Haiqiang Wang, Xiaozhong Xu, Shan Liu 0001
VCIP1
2023 MPED: Quantifying Point Cloud Distortion Based on Multiscale Potential Energy Discrepancy
abstract
In this article, we propose a new distortion quantification method for point clouds, the multiscale potential energy discrepancy (MPED). Currently, there is a lack of effective distortion quantification for a variety of point cloud perception tasks. Specifically, in human vision tasks, a distortion quantification method is used to predict human subjective scores and optimize the selection of human perception task parameters, such as dense point cloud compression and enhancement. In machine vision tasks, a distortion quantification method usually serves as loss function to guide the training of deep neural networks for unsupervised learning tasks (e.g., sparse point cloud reconstruction, completion, and upsampling). Therefore, an effective distortion quantification should be differentiable, distortion discriminable, and have low computational complexity. However, current distortion quantification cannot satisfy all three conditions. To fill this gap, we propose a new point cloud feature description method, the point potential energy (PPE), inspired by classical physics. We regard the point clouds are systems that have potential energy and the distortion can change the total potential energy. By evaluating various neighborhood sizes, the proposed MPED achieves global-local tradeoffs, capturing distortion in a multiscale fashion. We further theoretically show that classical Chamfer distance is a special case of our MPED. Extensive experiments show that the proposed MPED is superior to current methods on both human and machine perception tasks. Our code is available at https://github.com/Qi-Yangsjtu/MPED.
Qi Yang 0003, Siheng Chen, Yiling Xu, Jun Sun 0005, Zhan Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Point Cloud Quality Assessment: Dataset Construction and Learning-based No-reference Metric
abstract
Full-reference (FR) point cloud quality assessment (PCQA) has achieved impressive progress in recent years. However, in many cases, obtaining the reference point clouds is difficult, so no-reference (NR) metrics have become a research hotspot. Few researches about NR-PCQA are carried out due to the lack of a large-scale PCQA dataset. In this article, we first build a large-scale PCQA dataset named LS-PCQA, which includes 104 reference point clouds and more than 22,000 distorted samples. In the dataset, each reference point cloud is augmented with 31 types of impairments (e.g., Gaussian noise, contrast distortion, local missing, and compression loss) at 7 distortion levels. Besides, each distorted point cloud is assigned with a pseudo-quality score as its substitute of Mean Opinion Score. Inspired by the hierarchical perception system and considering the intrinsic attributes of point clouds, we propose a NR metric ResSCNN based on sparse convolutional neural network (CNN) to accurately estimate the subjective quality of point clouds. We conduct several experiments to evaluate the performance of the proposed NR metric. The results demonstrate that ResSCNN exhibits the state-of-the-art performance among all the existing NR-PCQA metrics and even outperforms some FR metrics. The dataset presented in this work will be made publicly accessible at https://smt.sjtu.edu.cn . The source code for the proposed ResSCNN can be found at https://github.com/lyp22/ResSCNN .
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu, Le Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2022 No-Reference Point Cloud Quality Assessment via Domain Adaptation
abstract
We present a novel no-reference quality assessment metric, the image transferred point cloud quality assessment (IT-PCQA), for 3D point clouds. For quality assessment, deep neural network (DNN) has shown compelling performance on no-reference metric design. However, the most challenging issue for no-reference PCQA is that we lack large-scale subjective databases to drive robust networks. Our motivation is that the human visual system (HVS) is the decision-maker regardless of the type of media for quality assessment. Leveraging the rich subjective scores of the natural images, we can quest the evaluation criteria of human perception via DNN and transfer the capability of prediction to 3D point clouds. In particular, we treat natural images as the source domain and point clouds as the target domain, and infer point cloud quality via unsupervised adversarial domain adaptation. To extract effective latent features and minimize the domain discrepancy, we propose a hierarchical feature encoder and a conditional-discriminative network. Considering that the ultimate pur-pose is regressing objective score, we introduce a novel con-ditional cross entropy loss in the conditional-discriminative network to penalize the negative samples which hinder the convergence of the quality regression network. Experi-mental results show that the proposed method can achieve higher performance than traditional no-reference metrics, even comparable results with full-reference metrics. The proposed method also suggests the feasibility of assessing the quality of specific media content without the expensive and cumbersome subjective evaluations. Code is available at https://github.com/Qi-Yangsjtu/IT-PCQA.
Qi Yang 0003, Yipeng Liu 0003, Siheng Chen, Yiling Xu, Jun Sun 0005
CVPR1
2022 Reduced Reference Quality Assessment for Point Cloud Compression
abstract
In this paper, we propose a reduced reference (RR) point cloud quality assessment (PCQA) model named R-PCQA to quantify the distortions introduced by the lossy compression. Specifically, we use the attribute and geometry quantization steps of different compression methods (i.e., V-PCC, G-PCC and AVS) to infer the point cloud quality, assuming that the point clouds have no other distortions before compression. First, we analyze the compression distortion of point clouds under separate attribute compression and geometry compression to avoid their mutual masking, for which we consider 5 point clouds as references to generate a compression dataset (PCCQA) containing independent attribute compression and geometry compression samples. Then, we develop the proposed R-PCQA via fitting the relationship between the quantization steps and the perceptual quality. We evaluate the performance of R-PCQA on both the established dataset and another independent dataset. The results demonstrate that the proposed R-PCQA can exhibit reliable performance and high generalization ability.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu
VCIP2
2022 Inferring Point Cloud Quality via Graph Similarity
abstract
Objective quality estimation of media content plays a vital role in a wide range of applications. Though numerous metrics exist for 2D images and videos, similar metrics are missing for 3D point clouds with unstructured and non-uniformly distributed points. In this paper, we propose [Formula: see text]-a metric to accurately and quantitatively predict the human perception of point cloud with superimposed geometry and color impairments. Human vision system is more sensitive to the high spatial-frequency components (e.g., contours and edges), and weighs local structural variations more than individual point intensities. Motivated by this fact, we use graph signal gradient as a quality index to evaluate point cloud distortions. Specifically, we first extract geometric keypoints by resampling the reference point cloud geometry information to form an object skeleton. Then, we construct local graphs centered at these keypoints for both reference and distorted point clouds. Next, we compute three moments of color gradients between centered keypoint and all other points in the same local graph for local significance similarity feature. Finally, we obtain similarity index by pooling the local graph significance across all color channels and averaging across all graphs. We evaluate [Formula: see text] on two large and independent point cloud assessment datasets that involve a wide range of impairments (e.g., re-sampling, compression, and additive noise). [Formula: see text] provides state-of-the-art performance for all distortions with noticeable gains in predicting the subjective mean opinion score (MOS) in comparison with point-wise distance-based metrics adopted in standardized reference software. Ablation studies further show that [Formula: see text] can be generalized to various scenarios with consistent performance by adjusting its key modules and parameters. Models and associated materials will be made available at https://njuvision.github.io/GraphSIM or http://smt.sjtu.edu.cn/papers/GraphSIM.
Qi Yang 0003, Zhan Ma 0001, Yiling Xu, Zhu Li 0001, Jun Sun 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 MS-GraphSIM: Inferring Point Cloud Quality via Multiscale Graph Similarity
abstract
To address the point cloud quality assessment (PCQA) problem, GraphSIM was proposed via jointly considering geometrical and color features, which shows compelling performance in multiple distortion detection. However, GraphSIM does not take into account the mutiscale characteristics of human perception. In this paper, we propose a multiscale PCQA model, called Multiscale Graph Similarity (MS-GraphSIM), that can better predict human subjective perception. First, exploring the multiscale processing method used in image processing, we introduce a multiscale representation of point clouds based on graph signal processing. Second, we extend GraphSIM into multiscale version based on the proposed multiscale representation. Specifically, MS-GraphSIM constructs a multiscale representation for each local patch extracted from the reference point cloud or the distorted point cloud, and then fuses GraphSIM at different scales to obtain an overall quality score. Experiment results demonstrate that the proposed MS-GraphSIM outperforms the state-of-the-art PCQA metrics over two fairly large and independent databases. Ablation studies further prove the proposed MS-GraphSIM is robust to different model hyperparameter settings. The code is available at https://github.com/zyj1318053/MS_GraphSIM.
Qi Yang 0003, Yiling Xu
ACM Multimedia2
2021 Predicting the Perceptual Quality of Point Cloud: A 3D-to-2D Projection-Based Exploration
abstract
Point cloud is emerged as a promising media format to represent realistic 3D objects or scenes in applications, such as virtual reality, teleportation, etc. How to accurately quantify the subjective point cloud quality for application-driven optimization, however, is still a challenging and open problem. In this paper, we attempt to tackle this problem in a systematic means. First, we produce a fairly large point cloud dataset where ten popular point clouds are augmented with seven types of impairments (e.g., compression, photometry/color noise, geometry noise, scaling) at six different distortion levels, and organize a formal subjective assessment with tens of subjects to collect mean opinion scores (MOS) for all 420 processed point cloud samples (PPCS). We then try to develop an objective metric that can accurately estimate the subjective quality. Towards this goal, we choose to project the 3D point cloud onto six perpendicular image planes of a cube for the color texture image and corresponding depth image, and aggregate image-based global (e.g., Jensen-Shannon (JS) divergence) and local features (e.g., edge, depth, pixel-wise similarity, complexity) among all projected planes for a final objective index. Model parameters are fixed constants after performing the regression using a small and independent dataset previously published. The proposed metric has demonstrated the state-of-the-art performance for predicting the subjective point cloud quality compared with multiple full-reference and no-reference models, e.g., the weighted peak signal-to-noise ratio (PSNR), structural similarity (SSIM), feature similarity (FSIM) and natural image quality evaluator (NIQE). The dataset is made publicly accessible athttp://smt.sjtu.edu.cnorhttp://vision.nju.edu.cnfor all interested audiences.
Qi Yang 0003, Hao Chen 0036, Zhan Ma 0001, Yiling Xu, Rongjun Tang, Jun Sun 0005
IEEE Trans. Multim.1