VLDB 2026 Research / reviewers in the wild / expert
Ziyu Shan
dblp:332/0929
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-3346-4261ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality AssessmentabstractIn recent years, No-Reference Point Cloud Quality Assessment (NR-PCQA) research has achieved significant progress. However, existing methods mostly seek a direct mapping function from visual data to the Mean Opinion Score (MOS), which is contradictory to the mechanism of practical subjective evaluation. To address this, we propose a novel language-driven PCQA method named CLIP-PCQA. Considering that human beings prefer to describe visual quality using discrete quality descriptions (e.g., "excellent" and "poor") rather than specific scores, we adopt a retrieval-based mapping strategy to simulate the process of subjective assessment. More specifically, based on the philosophy of CLIP, we calculate the cosine similarity between the visual features and multiple textual features corresponding to different quality descriptions, in which process an effective contrastive loss and learnable prompts are introduced to enhance the feature extraction. Meanwhile, given the personal limitations and bias in subjective experiments, we further covert the feature similarities into probabilities and consider the Opinion Score Distribution (OSD) rather than a single MOS as the final target. Experimental results show that our CLIP-PCQA outperforms other State-Of-The-Art (SOTA) approaches. Ziyu Shan, Yiling Xu |
AAAI | 3 |
| 2025 | Asynchronous Feedback Network for Perceptual Point Cloud Quality AssessmentabstractRecent years have witnessed the success of the deep learning-based technique in research of no-reference point cloud quality assessment (NR-PCQA). For a more accurate quality prediction, many previous studies have attempted to capture global and local features in a bottom-up manner, but ignored the interaction and promotion between them. To solve this problem, we propose a novel asynchronous feedback quality prediction network (AFQ-Net). Motivated by human visual perception mechanisms, AFQ-Net employs a dual-branch structure to deal with global and local features, simulating the left and right hemispheres of the human brain, and constructs a feedback module between them. Specifically, the input point clouds are first fed into a transformer-based global encoder to generate the attention maps that highlight these semantically rich regions, followed by being merged into the global feature. Then, we utilize the generated attention maps to perform dynamic convolution for different semantic regions and obtain the local feature. Finally, a coarse-to-fine strategy is adopted to merge the two features into the final quality score. We conduct comprehensive experiments on three datasets and achieve superior performance over the state-of-the-art approaches on all of these datasets. The code will be available athttps://github.com/zhangyujie-1998/AFQ-Net Qi Yang 0003, Ziyu Shan, Yiling Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality AssessmentabstractNo-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference, which have achieved tremendous improvements due to the utilization of deep neural networks. However, learning-based NR-PCQA methods suffer from the scarcity of labeled data and usually perform suboptimally in terms of generalization. To solve the problem, we propose a novel contrastive pre-training framework tailored for PCQA (CoPA), which enables the pre-trained model to learn quality-aware representations from unlabeled data. To obtain anchors in the representation space, we project point clouds with different distortions into images and randomly mix their local patches to form mixed images with multiple distortions. Utilizing the generated anchors, we constrain the pretraining process via a quality-aware contrastive loss following the philosophy that perceptual quality is closely related to both content and distortion. Furthermore, in the model fine-tuning stage, we propose a semantic-guided multi-view fusion module to effectively integrate the features of projected images from multiple perspectives. Extensive experiments show that our method outperforms the state-of-the-art PCQA methods on popular benchmarks. Further investigations demonstrate that CoPA can also benefit existing learning-based PCQA models. Ziyu Shan, Qi Yang 0003, Haichen Yang, Yiling Xu, Jenq-Neng Hwang, Xiaozhong Xu, Shan Liu 0001 |
CVPR | 1 |
| 2024 | MFT-PCQA: Multi-Modal Fusion Transformer for No-Reference Point Cloud Quality AssessmentabstractThe multi-modal information fusion for point cloud quality assessment (PCQA) is still understudied in existing work. Previous methods mostly adopt a late-fusion strategy without fully exploiting the advantages of different modalities and integrating them effectively. Considering that there exist both segregated processing and intertwined fusion when the human visual system (HVS) tackles different types of information, we propose a novel fusion transformer module for PCQA (MFT-PCQA). Specifically, we block the attention between point cloud features and image features to protect the self-attention of each modality, and utilize a mediate-fusion strategy to promote the cross-attention between modalities, encouraging the network to extract crucial interactions. Experimental results show that the proposed method outperforms state-of-the-art PCQA approaches. Ziyu Shan, Yiling Xu |
ICASSP | 2 |
| 2024 | Cross-Modal Distortion Approximation for Fast Bit Allocation of Video-Based Point Cloud CompressionabstractIn video-based point cloud compression (V-PCC), the optimal allocation of the total bitrate between geometry and color is a challenging but rewarding problem. Existing bit allocation approaches leverage statistical models to describe the rate and distortion of geometry and color information as functions of V-PCC quantization steps. However, to obtain the parameters of statistical models, these methods need to perform pre-coding for input point clouds for multiple times, resulting in high computational complexity. Consequently, the capability of these methods for practical application is limited. To address this problem, we derive the rate and distortion models based on projected images produced in the V-PCC encoding process, transforming the expensive point cloud pre-coding process into an efficient image pre-coding process. By utilizing the image-based distortion and rate models, the bit allocation problem is further formulated as a constrained convex optimization problem. Experimental results demonstrate that the proposed method exhibits significantly lower time complexity and higher rate-distortion performance compared to the existing methods. Haichen Yang, Qi Yang 0003, Ziyu Shan, Yiling Xu, Yunfeng Guan 0001 |
MMSP | 4 |
| 2024 | Learning Disentangled Representations for Perceptual Point Cloud Quality Assessment via Mutual Information MinimizationabstractNo-Reference Point Cloud Quality Assessment (NR-PCQA) aims to objectively assess the human perceptual quality of point clouds without relying on pristine-quality point clouds for reference. It is becoming increasingly significant with the rapid advancement of immersive media applications such as virtual reality (VR) and augmented reality (AR). However, current NR-PCQA models attempt to indiscriminately learn point cloud content and distortion representations within a single network, overlooking their distinct contributions to quality information. To address this issue, we propose DisPA, a novel disentangled representation learning framework for NR-PCQA. The framework trains a dual-branch disentanglement network to minimize mutual information (MI) between representations of point cloud content and distortion. Specifically, to fully disentangle representations, the two branches adopt different philosophies: the content-aware encoder is pretrained by a masked auto-encoding strategy, which can allow the encoder to capture semantic information from rendered images of distorted point clouds; the distortion-aware encoder takes a mini-patch map as input, which forces the encoder to focus on low-level distortion patterns. Furthermore, we utilize an MI estimator to estimate the tight upper bound of the actual MI and further minimize it to achieve explicit representation disentanglement. Extensive experimental results demonstrate that DisPA outperforms state-of-the-art methods on multiple PCQA datasets. Ziyu Shan, Yipeng Liu 0003, Yiling Xu |
NeurIPS | 1 |
| 2024 | GPA-Net:No-Reference Point Cloud Quality Assessment With Multi-Task Graph Convolutional NetworkabstractWith the rapid development of 3D vision, point cloud has become an increasingly popular 3D visual media content. Due to the irregular structure, point cloud has posed novel challenges to the related research, such as compression, transmission, rendering and quality assessment. In these latest researches, point cloud quality assessment (PCQA) has attracted wide attention due to its significant role in guiding practical applications, especially in many cases where the reference point cloud is unavailable. However, current no-reference metrics which based on prevalent deep neural network have apparent disadvantages. For example, to adapt to the irregular structure of point cloud, they require preprocessing such as voxelization and projection that introduce extra distortions, and the applied grid-kernel networks, such as Convolutional Neural Networks, fail to extract effective distortion-related features. Besides, they rarely consider the various distortion patterns and the philosophy that PCQA should exhibit shift, scaling, and rotation invariance. In this paper, we propose a novel no-reference PCQA metric named the Graph convolutional PCQA network (GPA-Net). To extract effective features for PCQA, we propose a new graph convolution kernel, i.e., GPAConv, which attentively captures the perturbation of structure and texture. Then, we propose the multi-task framework consisting of one main task (quality regression) and two auxiliary tasks (distortion type and degree predictions). Finally, we propose a coordinate normalization module to stabilize the results of GPAConv under shift, scale and rotation transformations. Experimental results on two independent databases show that GPA-Net achieves the best performance compared to the state-of-the-art no-reference PCQA metrics, even better than some full-reference metrics in some cases. Ziyu Shan, Qi Yang 0003, Rui Ye 0001, Yiling Xu, Xiaozhong Xu, Shan Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |