Yiling Xu

dblp:164/8950 · DBLP profile ↗
← Back
91ranked-venue papers
2as first author
67since 2021 · last 2026
0000-0002-4942-5232ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 2 first-author · 56 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Computer networks · 8 · 4 since 2021Systems, architecture and hardware · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Not All Views Matter: A View-Selective Approach to Point Cloud Perceptual Quality Assessment
Yida Xiang, Bingyang Cui, Kaifa Yang, Qi Yang 0003, Yiling Xu
QoMEX6
2026 Light4GS: Lightweight Compact 4D Gaussian Splatting Generation via Context Model
Mufan Liu, Qi Yang 0003, Zhenlong Yuan, Zhu Li 0001, Yiling Xu, Yunfeng Guan 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 From Images to Point Clouds: An Efficient Solution for Cross-Media Blind Quality Assessment Without Annotated Training
abstract
We present a novel quality assessment method which can predict the perceptual quality of point clouds from new scenes without available annotations by leveraging the rich prior knowledge in images, called the Distribution-Weighted Image-Transferred Point Cloud Quality Assessment (DWIT-PCQA). Recognizing the human visual system (HVS) as the decision-maker in quality assessment regardless of media types, we can emulate the evaluation criteria for human perception via neural networks and further transfer the capability of quality prediction from images to point clouds by leveraging the prior knowledge in the images. Specifically, domain adaptation (DA) can be leveraged to bridge the images and point clouds by aligning feature distributions of the two media in the same feature space. However, the different manifestations of distortions in images and point clouds make feature alignment a difficult task. To reduce the alignment difficulty and consider the different distortion distributions during alignment, we have derived formulas to decompose the optimization objective of the conventional DA into two suboptimization functions with distortion as a transition. Specifically, through network implementation, we propose the distortion-guided biased feature alignment which integrates existing/estimated distortion distribution into the adversarial DA framework, emphasizing common distortion patterns during feature alignment. Besides, we propose the quality-aware feature disentanglement to mitigate the destruction of the mapping from features to quality during alignment with biased distortions. Experimental results demonstrate that our proposed method exhibits reliable performance compared to general blind PCQA methods without needing point cloud annotations.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu, Le Yang 0001, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 On the Efficient Adaptive Streaming of 3D Gaussian Splatting Over Dynamic Networks
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising representation for immersive media. Its explicit splat-based structure offers high visual quality and real-time rendering, making it particularly suitable for six degrees of freedom streaming applications. However, its deployment in practical streaming scenarios is still limited due to several key challenges such as the large data volume, and insufficient support for dynamic bitrate adaptation under fluctuating network conditions. This paper presents an efficient 3DGS streaming framework that operates directly on pre-generated 3DGS models without retraining or fine-tuning. First, a training-free perceptual pruning method, which removes visually redundant Gaussians according to the human visual system metrics, is introduced. The resulting 3DGS is then encoded into a compact representation using the extended 3D codecs, exploiting its point-based structure. Next, we build a scene-specific bitrate ladder through analyzing the trade-off between resolution, bitrate, and perceptual quality. This enables efficient and fine-grained representation selection. Finally, a progressive streaming mechanism is developed. It is driven by a reinforcement learning scheduler that adaptively decides whether to download new content or enhance previously buffered content based on real-time network feedback. Experiments on real-world 3DGS datasets and bandwidth traces show that the proposed method evidently improves the quality of experience and streaming efficiency in various network scenarios.
Mufan Liu, Qi Yang 0003, Le Yang 0001, Yiling Xu
IEEE Trans. Circuits Syst. Video Technol.6
2026 Standardizing Generative Face Video Compression Using Supplemental Enhancement Information
abstract
This paper proposes a Generative Face Video Compression (GFVC) approach using Supplemental Enhancement Information (SEI), where a series of compact spatial and temporal representations of a face video signal (e.g., 2D/3D key-points, facial semantics and compact features) can be coded using SEI messages and inserted into the coded video bitstream. At the time of writing, the proposed GFVC approach using SEI messages has been included into a draft amendment of the Versatile Supplemental Enhancement Information (VSEI) standard by the Joint Video Experts Team (JVET) of ISO/IEC JTC 1/SC 29 and ITU-T SG21, which will be standardized as a new version of ITU-T H.274$|$ISO/IEC 23002-7. To the best of the authors' knowledge, the JVET work on the proposed SEI-based GFVC approach is the first standardization activity for generative video compression. The proposed SEI approach has not only advanced the reconstruction quality of early-day Model-Based Coding (MBC) via the state-of-the-art generative technique, but also established a new SEI definition for future GFVC applications and deployment. Experimental results illustrate that the proposed SEI-based GFVC approach can achieve remarkable rate-distortion performance compared with the latest Versatile Video Coding (VVC) standard, whilst also potentially enabling a wide variety of functionalities including user-specified animation/filtering and metaverse-related applications.
Yan Ye 0003, Jie Chen 0006, Ru-Ling Liao, Shanzhi Yin, Shiqi Wang 0001, Kaifa Yang, Yue Li 0015, Yiling Xu, Ye-Kui Wang, Shiv Gehlot, Guan-Ming Su, Peng Yin 0002, Sean McCarthy, Gary J. Sullivan
IEEE Trans. Multim.9
2026 Anchor-Driven Compact Gaussian Splatting for Dynamic Scene Reconstruction
abstract
Existing 4D Gaussian Splatting methods typically rely on per-Gaussian deformation from a canonical space to target frames, which overlooks the strong redundancy among spatially and temporally adjacent Gaussian primitives and leads to suboptimal efficiency. To address this limitation, we propose ADC-GS++, an anchor-driven compact Gaussian splatting framework for efficient and high-quality dynamic scene reconstruction. Specifically, ADC-GS++ organizes Gaussian primitives into an anchor-based canonical representation, enabling attribute sharing across local regions. To efficiently model dynamic scenes, we introduce a static-dynamic decomposition mechanism and further employ a coarse-to-fine deformation strategy driven by dynamic anchors at multiple granularities. In addition, a unified rate-distortion optimization is adopted to achieve a balanced trade-off between storage efficiency and reconstruction fidelity. Furthermore, a temporal significance-based anchor refinement strategy is employed to dynamically grow and prune anchors, allowing robust adaptation to complex and large-scale motions. Extensive experiments on multiple real-world dynamic scene datasets demonstrate that ADC-GS++ significantly improves rendering speed over deformation-based approaches by 300%-700%, while maintaining competitive rendering quality. Moreover, ADC-GS++ achieves a more favorable rate-distortion trade-off, resulting in substantially reduced storage consumption across different bitrate settings.
Qi Yang 0003, Mufan Liu, Yiling Xu, Zhu Li 0001
IEEE Trans. Vis. Comput. Graph.5
2026 Deformable 2D Gaussian Splatting for Efficient Wireless Radiance Field Rendering
abstract
Modeling the wireless radiance field (WRF) is fundamental to modern communication systems, enabling key tasks such as localization, sensing, and channel estimation. Traditional approaches, which rely on empirical formulas or physical simulations, often suffer from limited accuracy or require strong scene priors. Recent neural radiance field (NeRF)-based methods improve reconstruction fidelity through differentiable volumetric rendering, but their reliance on computationally expensive multilayer perceptron (MLP) queries hinders real-time deployment. To overcome these challenges, we introduce Gaussian splatting (GS) to the wireless domain, leveraging its efficiency in modeling optical radiance fields to enable compact and accurate WRF reconstruction. Specifically, we propose SwiftWRF, a deformable 2D Gaussian splatting framework that synthesizes WRF spectra at arbitrary positions under single-sided transceiver mobility. SwiftWRF employs CUDA-accelerated rasterization to render spectra at over 100 k FPS and uses the lightweight MLP to model the deformation of 2D Gaussians, effectively capturing mobility-induced WRF variations. In addition to novel spectrum synthesis, the efficacy of SwiftWRF is further underscored in its applications in angle-of-arrival (AoA) and received signal strength indicator (RSSI) prediction. Experiments conducted on both real-world and synthetic indoor scenes demonstrate that SwiftWRF can reconstruct WRF spectra up to 500x faster than existing state-of-the-art methods, while significantly enhancing its signal quality.
Mufan Liu, Cixiao Zhang, Qi Yang 0003, Yiling Xu, Yin Xu 0001, Shu Sun 0001, Mingzeng Dai, Yunfeng Guan 0001
IEEE Trans. Vis. Comput. Graph.5
2025 CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
abstract
In recent years, No-Reference Point Cloud Quality Assessment (NR-PCQA) research has achieved significant progress. However, existing methods mostly seek a direct mapping function from visual data to the Mean Opinion Score (MOS), which is contradictory to the mechanism of practical subjective evaluation. To address this, we propose a novel language-driven PCQA method named CLIP-PCQA. Considering that human beings prefer to describe visual quality using discrete quality descriptions (e.g., "excellent" and "poor") rather than specific scores, we adopt a retrieval-based mapping strategy to simulate the process of subjective assessment. More specifically, based on the philosophy of CLIP, we calculate the cosine similarity between the visual features and multiple textual features corresponding to different quality descriptions, in which process an effective contrastive loss and learnable prompts are introduced to enhance the feature extraction. Meanwhile, given the personal limitations and bias in subjective experiments, we further covert the feature similarities into probabilities and consider the Opinion Score Distribution (OSD) rather than a single MOS as the final target. Experimental results show that our CLIP-PCQA outperforms other State-Of-The-Art (SOTA) approaches.
Ziyu Shan, Yiling Xu
AAAI4
2025 A Quality-Aware Sampling Framework for Efficient 3D Point Cloud Transmission
abstract
The large volume of data from the point cloud brings significant demands on network bandwidth. However, the current transmission framework only considers using lossy compression to control the size of data, while ignoring visually redundant information due to the setting of rendering devices. Based on the fact that point overlapping might occur for the case that a dense point cloud is rendered on a relatively low resolution 2D monitor, we propose a novel quality-aware sampling framework for point cloud transmission. When a target visual quality is determined, an optimal sampling module is designed to remove overlapped points with the help of a simple but effective quality model. By taking into account the impact of multiple factors (i.e., sampling, lossy compression, and client rendering resolution), this quality model can predict the final perceptual quality in the client. Based on a newly constructed dataset which consists of 420 samples, experiment results show that the proposed transmission framework can significantly reduce bandwidth cost (e.g., 6.10% to 84.43%) and processing time (e.g., 8.99% to 92.53%) without introducing noticeable distortion under certain rendering conditions, thus achieving higher bandwidth utilization and better real-time performance.
Puyue Hou, Qi Yang 0003, Yue Li 0015, Jianchao Yang, Yiling Xu, Tiejun Huang 0001
ICASSP6
2025 A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
abstract
3D Gaussian Splatting (GS) demonstrates excellent rendering quality and generation speed in novel view synthesis. However, substantial data size poses challenges for storage and transmission, making 3D GS compression an essential technology. Current 3D GS compression research primarily focuses on developing more compact scene representations, such as converting explicit 3D GS data into implicit forms. In contrast, compression of the GS data itself has hardly been explored. To address this gap, we propose a Hierarchical GS Compression (HGSC) technique. Initially, we prune unimportant Gaussians based on importance scores derived from both global and local significance, effectively reducing redundancy while maintaining visual quality. An Octree structure is used to compress 3D positions. Based on the 3D GS Octree, we implement a hierarchical attribute compression strategy by employing a KD-tree to partition the 3D GS into multiple blocks. We apply farthest point sampling to select anchor primitives within each block and others as non-anchor primitives with varying Levels of Details (LoDs). Anchor primitives serve as reference points for predicting non-anchor primitives across different LoDs to reduce spatial redundancy. For anchor primitives, we use the region adaptive hierarchical transform to achieve near-lossless compression of various attributes. For non-anchor primitives, each is predicted based on the k-nearest anchor primitives. To further minimize prediction errors, the reconstructed LoD and anchor primitives are combined to form new anchor primitives to predict the next LoD. Our method notably achieves superior compression quality and a significant data size reduction of over 4.5× compared to the state-of-the-art compression method on small scenes datasets.
Qi Yang 0003, Yiling Xu, Zhu Li 0001
ICASSP4
2025 Neural Adaptive Contextual Video Streaming
abstract
Video streaming services typically employ traditional codecs, such as H.264, to encode videos into multiple bitrate representations. These codecs are tightly limited by discrete quantization parameters (QPs), resulting in encoded rates that do not align with the target bitrate. Additionally, the subpar video quality produced by conventional codecs does not meet the demands of high-resolution communication. Considering the limitations of traditional codecs, we take a fresh new approach to video streaming by leveraging advanced deep learning-based video codecs. Specifically, we develop a neural adaptive contextual video streaming framework that incorporates: 1) an ensemble deep reinforcement learning based adaptive bitrate algorithm named TSAC that enables continuous bitrate adjustment to varying network conditions 2) a two-stage proportional-integral-derivative-based rate control module that dynamically fine-tunes QPs to ensure the encoded bitrate aligning with the target bitrate. Furthermore, we implement intra-GoP and inter-GoP techniques to accelerate the inference process of the contextual video codec for real-time processing needs. Our experiments demonstrate that the average relative error in bitrate remains below 2%, the quality of experience provided by our TSAC agents surpasses that of existing discrete algorithms by 13%-20%. Our optimization techniques enable real-time decoding at approximately 24 frames per second for quad high definition videos.
Jianchao Yang, Mufan Liu, Puyue Hou, Yiling Xu, Jun Sun 0005
ICASSP4
2025 Deep Joint Source-Channel Coding for Wireless Point Cloud Transmission
abstract
The growing demand for high-quality point cloud transmission over wireless networks presents significant challenges, primarily due to the large data sizes and the need for efficient encoding techniques. In response to these challenges, we introduce a novel system named Deep Point Cloud Semantic Transmission (PCST), designed for end-to-end wireless point cloud transmission. Our approach employs a progressive resampling framework using sparse convolution to project point cloud data into a semantic latent space. These semantic features are subsequently encoded through a deep joint source-channel (JSCC) encoder, generating the channel-input sequence. To enhance transmission efficiency, we use an adaptive entropy-based approach to assess the importance of each semantic feature, allowing transmission lengths to vary according to their predicted entropy. PCST is robust across diverse Signal-to-Noise Ratio (SNR) levels and supports an adjustable rate-distortion (RD) trade-off, ensuring flexible and efficient transmission. Experimental results indicate that PCST significantly outperforms traditional separate source-channel coding (SSCC) schemes, delivering superior reconstruction quality while achieving over a 50% reduction in bandwidth usage.
Cixiao Zhang, Mufan Liu, Yin Xu 0001, Yiling Xu, Dazhi He
ICASSP5
2025 LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression
abstract
Existing AI-based point cloud compression methods struggle with dependence on specific training data distributions, which limits their real-world deployment. Implicit Neural Representation (INR) methods solve the above problem by encoding overfitted network parameters to the bitstream, resulting in more distribution-agnostic results. However, due to the limitation of encoding time and decoder size, current INR based methods only consider lossy geometry compression. In this paper, we propose the first INR based lossless point cloud geometry compression method called Lossless Implicit Neural Representations for Point Cloud Geometry Compression (LINR-PCGC). To accelerate encoding speed, we design a group of point clouds level coding framework with an effective network initialization strategy, which can reduce around 60% encoding time. A lightweight coding network based on multiscale SparseConv, consisting of scale context extraction, child node prediction, and model compression modules, is proposed to realize fast inference and compact decoder size. Experimental results show that our method consistently outperforms traditional and AI-based methods: for example, with the convergence time in the MVUB dataset, our method reduces the bitstream by approximately 21.21% compared to G-PCC TMC13v23 and 21.95% compared to SparsePCGC. Our project can be seen on https://huangwenjie2023.github.io/LINR-PCGC/.
Qi Yang 0003, Shuting Xia, Yiling Xu, Zhu Li 0001
ICCV5
2025 Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-To-3D Generation
abstract
Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) Existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation dimensions. ii) Previous evaluation metrics only focus on a single aspect (e.g., text-3D alignment) and fail to perform multi-dimensional quality assessment. To address these problems, we first propose a comprehensive benchmark named MATE-3D. The benchmark contains eight well-designed prompt categories that cover single and multiple object generation, resulting in 1,280 generated textured meshes. We have conducted a large-scale subjective experiment from four different evaluation dimensions and collected 107,520 annotations, followed by detailed analyses of the results. Based on MATE-3D, we propose a novel quality evaluator named HyperScore. Utilizing hypernetwork to generate specified mapping functions for each evaluation dimension, our metric can effectively perform multi-dimensional quality assessment. HyperScore presents superior performance over existing metrics on MATE-3D, making it a promising metric for assessing and improving text-to-3D generation. The project is available at https://mate-3d.github.io/.
Bingyang Cui, Qi Yang 0003, Zhu Li 0001, Yiling Xu
ICCV5
2025 DPCD: A Quality Assessment Database for Dynamic Point Clouds
abstract
Recently, the advancements in Virtual/Augmented Reality (VR/AR) have driven the demand for Dynamic Point Clouds (DPC). Unlike static point clouds, DPCs are capable of capturing temporal changes within objects or scenes, offering a more accurate simulation of the real world. While significant progress has been made in the quality assessment research of static point cloud, little study has been done on Dynamic Point Cloud Quality Assessment (DPCQA), which hinders the development of quality-oriented applications, such as interframe compression and transmission in practical scenarios. In this paper, we introduce a large-scale DPCQA database, named DPCD, which includes 15 reference DPCs and 525 distorted DPCs from seven types of lossy compression and noise distortion. By rendering these samples to Processed Video Sequences (PVS), a comprehensive subjective experiment is conducted to obtain Mean Opinion Scores (MOS) from 21 viewers for analysis. The characteristic of contents, impact of various distortions, and accuracy of MOSs are presented to validate the heterogeneity and reliability of the proposed database. Furthermore, we evaluate the performance of several objective metrics on DPCD. The experiment results show that DPCQA is more challenge than that of static point cloud. The DPCD, which serves as a catalyst for new research endeavors on DPCQA, is publicly available at https://huggingface.co/datasets/Olivialyt/DPCD.
Qi Yang 0003, Yiling Xu, Zhu Li 0001, Ye-Kui Wang
ICME4
2025 ADC-GS: Anchor-Driven Deformable and Compressed Gaussian Splatting for Dynamic Scene Reconstruction
abstract
Existing 4D Gaussian Splatting methods rely on per-Gaussian deformation from a canonical space to target frames, which overlooks redundancy among adjacent Gaussian primitives and result in suboptimal performance. To address this limitation, we propose Anchor-Driven Deformable and Compressed Gaussian Splatting (ADC-GS), a compact and efficient representation for dynamic scene reconstruction. Specifically, ADC-GS organizes Gaussian primitives into an anchor-based structure within the canonical space, enhanced by a temporal significance-based anchor refinement strategy. To reduce deformation redundancy, ADC-GS introduces a hierarchical coarse-to-fine pipeline that captures motions at varying granularities. Moreover, a rate-distortion optimization is adopted to achieve an optimal balance between bitrate consumption and representation fidelity. Experimental results demonstrate that ADC-GS outperforms the per-Gaussian deformation approaches in rendering speed by 300%-800% while achieving state-of-the-art storage efficiency without compromising rendering quality. The code is released at https://github.com/H-Huang774/ADC-GS.git.
Qi Yang 0003, Mufan Liu, Yiling Xu, Zhu Li 0001
IJCAI4
2025 Adaptive Local Reconstruction for Geometry-based Point Cloud Compression
abstract
In recent years, point cloud compression has emerged as a focal research topic. Geometry-based Point Cloud Compression (G-PCC) developed by the Moving Picture Experts Group (MPEG) has demonstrated to achieve remarkable compression efficiency across a diverse range of datasets. Triangle soup (Trisoup) is an efficient lossy geometry compression tool in G-PCC. However, its sampling rate in the geometry reconstruction stage is globally fixed, resulting in a relatively uniform distribution of the reconstructed point cloud, which cannot accurately represent the variations in density across different regions. In this paper, we propose an adaptive local reconstruction method to enhance the local quality of the point cloud reconstructed by Trisoup. On the encoding side, we first analyze the distribution density of the input point cloud in each leaf node. Based on this information, we adjust the local sampling rate of ray tracing during the geometry reconstruction process, thus generating region-specific local representations that closely resemble the ground truth. Experimental results show that the proposed method achieves an average geometry coding gain of 10.7% under D2 quality metric, along with a reduction in time complexity on both the encoder and decoder. Benefiting from the accurately reconstructed geometry, the attribute coding performance is improved by 2.1%, 2.0%, 2.5% in Luma, Chroma Cb, Chroma Cr channels, respectively.
Lizhi Hou, Yiling Xu
ISCAS3
2025 A Learning Framework for Predicting CT-Based PRM Biomarker from MRI Sequences in COPD
Yiling Xu, Simon M. F. Triphan, Julian Grolig, Hanyi Zhang, Jürgen Biederer, Craig J. Galbán, Hans-Ulrich Kauczor, Mark Oliver Wielpütz, Oliver Weinheimer
MICCAI (8)1
2025 3DGS-IEval-15K: A Large-scale Image Quality Evaluation Database for 3D Gaussian-Splatting
Yuke Xing, Peizhi Niu, Guangtao Zhai, Yiling Xu
ACM Multimedia6
2025 CompBench: Benchmarking and Comparing Image Generation with Large Multimodal Models
abstract
Recent advancements in large multimodal models (LMMs) have significantly enhanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, critical challenges in perceptual quality and text-image correspondence remain hindering the practicality of AI-generated images (AIGIs). Therefore, a reliable benchmark and automatic model for AIGI evaluation is desirable, which heavily relies on the scale and quality of human annotations. To this end, we present CompBench, the largest dataset for benchmarking and comparing image generation, which features: (i) the largest AIGI pair comparison dataset, comprising 616,346 carefully curated image pairs generated by 24 state-of-the-art AIGI models annotated with 1.6M+ human annotations, enabling robust relative quality assessment through pairwise comparison, (ii) multi-dimensional pairwise comparison from perceptual and text-image correspondence perspectives across three difficulty levels, and (iii) bidirectional benchmarking and evaluating for both T2I generation models and AIGI comparison models. Based on CompBench, we propose LMM4Comp, a LMM-based evaluation metric that learns nuanced quality distinctions from multiple dimensions for pairwise comparison at both instance level and model level. Experiments demonstrate that LMM4Comp achieves state-of-the-art performance, highly aligning to human preference. Both of the CompBench dataset and LMM4Comp metric will be released at https://github.com/IntMeGroup/CompBench.
Huiyu Duan, Yuke Xing, Yiling Xu, Guangtao Zhai, Xiongkuo Min
MMSP4
2025 Video Streaming with Kairos: An MPC-Based ABR with Streaming-Aware Throughput Prediction
abstract
Throughput prediction in current adaptive bitrate (ABR) schemes often neglects streaming-aware characteristics, such as sequence irregularity and prediction smoothness, resulting in inaccurate predictions and suboptimal performance. To address these challenges, we propose Kairos, an MPC-based ABR scheme that integrates an attention-based throughput predictor with buffer-aware uncertainty control to enhancing both prediction accuracy and adaptability to dynamic network conditions. Specifically, Kairos employs a multi-time attention network (mTAN) to process irregularly sampled streaming data, producing uniformly spaced latent representations. Based on these, we introduce a percentile prediction network to estimate future throughput percentiles, along with a buffer-aware uncertainty control module that selects the optimal percentile based on the current buffer status. As smoothness is another key component of QoE, we incorporate a smoothness regularizer to ensure consistent throughput predictions, thereby facilitating smoother ABR decisions. Our Kairos design integrates sampling irregularity, prediction uncertainty, and smoothness into the throughput prediction, significantly enhancing bitrate decision making within the MPC framework. Extensive trace-driven and real-world experiments demonstrate that Kairos outperforms state-of-the-art ABR schemes, achieving a QoE improvement ranging from 6.42% to 29.45% across diverse network conditions.
Ziyu Zhong, Mufan Liu, Le Yang 0001, Yiling Xu, Jenq-Neng Hwang
NOSSDAV5
2025 3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compression
abstract
3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual fidelity, but its significant storage demands limit practical deployment, prompting recent methods to integrate compression modules into 3DGS. However, these 3DGS generative compression techniques introduce unique distortions that lack systematic quality assessment research. To this end, we establish 3DGS-VBench, a large-scale Video Quality Assessment (VQA) dataset and benchmark with 660 compressed 3DGS models and video sequences generated from 11 scenes across 6 representative 3DGS compression algorithms with systematically designed parameter levels. With annotations from 50 participants, we obtain MOS scores with outlier removal and validate dataset reliability. We benchmark 6 3DGS compression algorithms on storage efficiency and visual quality, and evaluate 15 quality assessment metrics. Our dataset enables specialized VQA model training for 3DGS. The dataset is available at https://github.com/YukeXing/3DGS-VBench.
Yuke Xing, William Gordon, Qi Yang 0003, Kaifa Yang, Yiling Xu
VCIP6
2025 Asynchronous Feedback Network for Perceptual Point Cloud Quality Assessment
abstract
Recent years have witnessed the success of the deep learning-based technique in research of no-reference point cloud quality assessment (NR-PCQA). For a more accurate quality prediction, many previous studies have attempted to capture global and local features in a bottom-up manner, but ignored the interaction and promotion between them. To solve this problem, we propose a novel asynchronous feedback quality prediction network (AFQ-Net). Motivated by human visual perception mechanisms, AFQ-Net employs a dual-branch structure to deal with global and local features, simulating the left and right hemispheres of the human brain, and constructs a feedback module between them. Specifically, the input point clouds are first fed into a transformer-based global encoder to generate the attention maps that highlight these semantically rich regions, followed by being merged into the global feature. Then, we utilize the generated attention maps to perform dynamic convolution for different semantic regions and obtain the local feature. Finally, a coarse-to-fine strategy is adopted to merge the two features into the final quality score. We conduct comprehensive experiments on three datasets and achieve superior performance over the state-of-the-art approaches on all of these datasets. The code will be available athttps://github.com/zhangyujie-1998/AFQ-Net
Qi Yang 0003, Ziyu Shan, Yiling Xu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Lossless LiDAR Point Cloud Reflectance Compression With a Deep Hierarchical KNN Context Model
abstract
Recently, numerous learning-based point cloud compression methods with outstanding performance have been developed. The majority of them concentrate on point cloud geometry compression, and several works have demonstrated advances in the color attribute compression for dense point clouds. However, compression of the reflectance attribute attached to the point captured by the light detection and ranging (LiDAR) sensors remains a major challenge. In this article, we present a lossless reflectance compression method for LiDAR point clouds (LPCs) that learns reflectance probability distributions with a deep hierarchical k-nearest-neighbors (KNN) context model, namely, the HK-PCRC. We first represent the original LPC with a series of hierarchical layers. Relying on the hierarchical structure, points in the same layer are coded in parallel by referencing the points in the previously coded layers. The approach balances the coding efficiency and time complexity while also supporting the progressive coding functionality. By introducing the KNN context, the context size is significantly reduced, which eases the computational burden while maintaining the coding performance. To enrich the context information, we further search for enhanced neighbors for each point in the context window. For each enhanced neighbor, in addition to its reflectance value, the relative distance, elevation angle, and local density are further collected. Then, a transformer-style sequential model is applied to construct an accurate deep context model. Furthermore, to efficiently fuse context features from different sources, a cross-feature fusion attention mechanism is designed for the transformer network. The comprehensive experimental results on SemanticKITTI, a large scale LiDAR benchmark, and Ford, an MPEG-specified dataset, demonstrate that our proposed framework achieves a state-of-the-art reflectance lossless compression performance, with average bit savings of 11.3% and 9.6% when compared to the state-of-the-art hand-crafted methods.
Lizhi Hou, Tingyu Fan, Yiling Xu, Zhu Li 0001
IEEE Trans. Multim.3
2025 Textured Mesh Quality Assessment Using Geometry and Color Field Similarity
abstract
Textured mesh quality assessment (TMQA) is critical for various 3D mesh applications. However, existing TMQA methods often struggle to provide accurate and robust evaluations. Motivated by the effectiveness of fields in representing both 3D geometry and color information, we propose a novel point-based TMQA method called field mesh quality metric (FMQM). FMQM utilizes signed distance fields and a newly proposed color field named nearest surface point color field to realize effective mesh feature description. Four features related to visual perception are extracted from the geometry and color fields: geometry similarity, geometry gradient similarity, space color distribution similarity, and space color gradient similarity. Experimental results on three benchmark datasets demonstrate that FMQM outperforms state-of-the-art (SOTA) TMQA metrics. Furthermore, FMQM exhibits low computational complexity, making it a practical and efficient solution for real-world applications in 3D graphics and visualization.
Kaifa Yang, Qi Yang 0003, Yiling Xu, Zhu Li 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
abstract
No-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference, which have achieved tremendous improvements due to the utilization of deep neural networks. However, learning-based NR-PCQA methods suffer from the scarcity of labeled data and usually perform suboptimally in terms of generalization. To solve the problem, we propose a novel contrastive pre-training framework tailored for PCQA (CoPA), which enables the pre-trained model to learn quality-aware representations from unlabeled data. To obtain anchors in the representation space, we project point clouds with different distortions into images and randomly mix their local patches to form mixed images with multiple distortions. Utilizing the generated anchors, we constrain the pretraining process via a quality-aware contrastive loss following the philosophy that perceptual quality is closely related to both content and distortion. Furthermore, in the model fine-tuning stage, we propose a semantic-guided multi-view fusion module to effectively integrate the features of projected images from multiple perspectives. Extensive experiments show that our method outperforms the state-of-the-art PCQA methods on popular benchmarks. Further investigations demonstrate that CoPA can also benefit existing learning-based PCQA models.
Ziyu Shan, Qi Yang 0003, Haichen Yang, Yiling Xu, Jenq-Neng Hwang, Xiaozhong Xu, Shan Liu 0001
CVPR5
2024 Mesla: Neural Adaptive Layered Point Cloud Streaming with Enhanced Buffer Management
abstract
Point cloud video provides an immersive experience with six degrees of freedom (6DoF). However, streaming video imposes significant challenges due to the huge data volume involved, which increases the transmission burden and makes it difficult to maintain a high quality of experience (QoE) in bandwidth-constrained networks. Current methods either employ pre-specified control rules or make irreversible decisions to optimize the QoE, hindering adaptation to the varying network conditions and not employing buffer management. In this work, we present Mesla1, an adaptive layered point cloud streaming framework, which is empowered by deep reinforcement learning (DRL) and hierarchical buffer management. By maximizing the utility-driven QoE via proximal policy optimization (PPO)-based DRL, Mesla makes bitrate decisions based on the obtained network observations and viewing behaviors of users. The viewport-oriented utility is incorporated into the QoE objective for assigning point cloud tiles with different priorities. To achieve fine-grained rate adaptation, we partition point clouds into layers with different levels of details (LoD) and develop a buffering heuristic (i.e., the hierarchical buffer management) that enhances the quality of unconsumed content. Extensive simulations demonstrate the advantages of Mesla compared to the benchmarking schemes, corroborating the effectiveness of the proposed system.
Mufan Liu, Puyue Hou, Le Yang 0001, Yiling Xu
GLOBECOM5
2024 SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map
abstract
In recent years, static meshes with texture maps have become one of the most prevalent digital representations of 3D shapes in various applications, such as animation, gaming, medical imaging, and cultural heritage applications. However, little research has been done on the quality assessment of textured meshes, which hinders the development of quality-oriented applications, such as mesh compression and enhancement. In this paper, we create a large-scale textured mesh quality assessment database, namely SJTU-TMQA, which includes 21 reference meshes and 945 distorted samples. The meshes are rendered into processed video sequences and then conduct subjective experiments to obtain mean opinion scores (MOS). The diversity of content and accuracy of MOS has been shown to validate its heterogeneity and reliability. The impact of various types of distortion on human perception is demonstrated. 13 state-of-the-art objective metrics are evaluated on SJTU-TMQA. The results report the highest correlation is around 0.6, indicating the need for more effective objective metrics. The SJTU-TMQA is available at https://ccccby.github.io
Bingyang Cui, Qi Yang 0003, Kaifa Yang, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
ICASSP4
2024 MFT-PCQA: Multi-Modal Fusion Transformer for No-Reference Point Cloud Quality Assessment
abstract
The multi-modal information fusion for point cloud quality assessment (PCQA) is still understudied in existing work. Previous methods mostly adopt a late-fusion strategy without fully exploiting the advantages of different modalities and integrating them effectively. Considering that there exist both segregated processing and intertwined fusion when the human visual system (HVS) tackles different types of information, we propose a novel fusion transformer module for PCQA (MFT-PCQA). Specifically, we block the attention between point cloud features and image features to protect the self-attention of each modality, and utilize a mediate-fusion strategy to promote the cross-attention between modalities, encouraging the network to extract crucial interactions. Experimental results show that the proposed method outperforms state-of-the-art PCQA approaches.
Ziyu Shan, Yiling Xu
ICASSP4
2024 EVAN: Evolutional Video Streaming Adaptation via Neural Representation
abstract
Adaptive bitrate (ABR) using conventional codecs cannot further modify the bitrate once a decision has been made, exhibiting limited adaptation capability. This may result in either overly conservative or overly aggressive bitrate selection, which could cause either inefficient utilization of the network bandwidth or frequent re-buffering, respectively. Neural representation for video (NeRV), which embeds the video content into neural network weights, allows video reconstruction with incomplete models. Specifically, the recovery of one frame can be achieved without relying on the decoding of adjacent frames. NeRV has the potential to provide high video reconstruction quality and, more importantly, pave the way for developing more flexible ABR strategies for video transmission. In this work, a new framework, named Evolutional Video streaming Adaptation via Neural representation (EVAN), which can adaptively transmit NeRV models based on soft actor-critic (SAC) reinforcement learning, is proposed. EVAN is trained with a more exploitative strategy and utilizes progressive playback to avoid re-buffering. Experiments showed that EVAN can outperform existing ABRs with 50% reduction in re-buffering and achieve nearly 20% improvement in users’ quality of experience (QoE).
Mufan Liu, Le Yang 0001, Yiling Xu, Ye-Kui Wang, Jenq-Neng Hwang
ICME3
2024 Enhancing Real-Time Video Streaming with Joint Frame Size and Rate Adaptation
abstract
With advancing network technologies, real-time communication (RTC) scenarios like cloud gaming and video conferencing have gained more attention. However, when addressing the challenge of meeting users’ high-quality demands while dealing with network fluctuations, existing adaptive bitrate (ABR) or adaptive framerate (AFR) algorithms encounter limitations in enhancing Quality of Experience (QoE). This paper introduces Adaptive Frame Size and Rate (AFSR), a joint adaption algorithm based on DRL. AFSR dynamically adjusts the frame rate and frame size (magnitude of bits), and improves QoE in RTC by accurately assessing inter-frame quality and its impact on latency and stall. AFSR uses non-linear bitrate-quality relationships and precise end-to-end latency measurements. Comparative evaluations confirm AFSR’s superior performance in RTC video transmission.
Hengchao Wang, Ziyu Zhong, Jiaoyang Yin, Yiling Xu, Le Yang 0001
ISCAS4
2024 MS-GeodesicPSIM: Predicting the Quality of Static Mesh with Texture Map via multi-scale Geodesic Patch Similarity
abstract
To address the mesh quality assessment (MQA) problem, GeodesicP-SIM was proposed by jointly considering geometry and color features, demonstrating compelling performance in multiple benchmarks.However, GeodesicPSIM does not consider the multi-scale characteristics of human perception.To better mimic human subjective perception, we proposed a multi-scale MQA model called multi-scale Geodesic Patch Similarity (MS-GeodesicPSIM).Firstly, inspired by the multi-scale processing methods used in image and point cloud analysis, we propose a novel multi-scale representation of textured meshes based on mesh simplification techniques.Secondly, we extend GeodesicPSIM into a multi-scale version leveraging the proposed multi-scale representation.Specifically, we construct a multi-scale representation for the reference and distorted meshes, followed by fusing the results of GeodesicPSIM at different scales to obtain an overall quality score.Experimental results demonstrate the superior performance of the proposed MS-GeodesicPSIM compared to the single-scale GeodesicPSIM and other MQA metrics on three large and independent databases.Ablation studies further confirm that MS-GeodesicPSIM is robust to different model hyperparameter settings.The code for MS-GeodesicPSIM is available at https://github.com/ccccby/MS-GeodesicPSIM
Bingyang Cui, Qi Yang 0003, Yiling Xu
MMAsia4
2024 A Benchmark for Gaussian Splatting Compression and Quality Assessment Study
Qi Yang 0003, Kaifa Yang, Yuke Xing, Yiling Xu, Zhu Li 0001
MMAsia4
2024 Cross-Modal Distortion Approximation for Fast Bit Allocation of Video-Based Point Cloud Compression
abstract
In video-based point cloud compression (V-PCC), the optimal allocation of the total bitrate between geometry and color is a challenging but rewarding problem. Existing bit allocation approaches leverage statistical models to describe the rate and distortion of geometry and color information as functions of V-PCC quantization steps. However, to obtain the parameters of statistical models, these methods need to perform pre-coding for input point clouds for multiple times, resulting in high computational complexity. Consequently, the capability of these methods for practical application is limited. To address this problem, we derive the rate and distortion models based on projected images produced in the V-PCC encoding process, transforming the expensive point cloud pre-coding process into an efficient image pre-coding process. By utilizing the image-based distortion and rate models, the bit allocation problem is further formulated as a constrained convex optimization problem. Experimental results demonstrate that the proposed method exhibits significantly lower time complexity and higher rate-distortion performance compared to the existing methods.
Haichen Yang, Qi Yang 0003, Ziyu Shan, Yiling Xu, Yunfeng Guan 0001
MMSP5
2024 Learning Disentangled Representations for Perceptual Point Cloud Quality Assessment via Mutual Information Minimization
abstract
No-Reference Point Cloud Quality Assessment (NR-PCQA) aims to objectively assess the human perceptual quality of point clouds without relying on pristine-quality point clouds for reference. It is becoming increasingly significant with the rapid advancement of immersive media applications such as virtual reality (VR) and augmented reality (AR). However, current NR-PCQA models attempt to indiscriminately learn point cloud content and distortion representations within a single network, overlooking their distinct contributions to quality information. To address this issue, we propose DisPA, a novel disentangled representation learning framework for NR-PCQA. The framework trains a dual-branch disentanglement network to minimize mutual information (MI) between representations of point cloud content and distortion. Specifically, to fully disentangle representations, the two branches adopt different philosophies: the content-aware encoder is pretrained by a masked auto-encoding strategy, which can allow the encoder to capture semantic information from rendered images of distorted point clouds; the distortion-aware encoder takes a mini-patch map as input, which forces the encoder to focus on low-level distortion patterns. Furthermore, we utilize an MI estimator to estimate the tight upper bound of the actual MI and further minimize it to achieve explicit representation disentanglement. Extensive experimental results demonstrate that DisPA outperforms state-of-the-art methods on multiple PCQA datasets.
Ziyu Shan, Yipeng Liu 0003, Yiling Xu
NeurIPS4
2024 Sample Adaptive Offset for Geometry-based Point Cloud Attribute Compression
abstract
Geometry-based Point Cloud Compression (G-PCC) has demonstrated remarkable coding efficiency across a wide range of datasets. However, due to the irreversible quantization loss, the G-PCC reconstructed attribute is inevitably distorted. In this paper, we propose a Sample Adaptive Offset (SAO) method for G-PCC to reduce attribute distortion. Specifically, after the regular attribute coding of G-PCC, the whole point cloud is first classified according to their distorted attribute values to maintain a low-complexity design. The cross-channel correlation is also utilized to enhance the classification process. Then, one optimal offset is derived for each class by a Rate Distortion Optimization (RDO). A second stage RDO is performed to ensure the maximum positive gain. Finally, the selected classes and their corresponding offsets are applied to the reconstructed attribute to reduce the distortion. We integrate the proposed method on top of the G-PCC solid point cloud test model GeSTM_v6.0. The simulation results show that the proposed method can achieve up to {0.4%, 6.3%, 6.1%} attribute Bjøntegaard-Delta (BD)-rate gain on different channels with negligible additional complexity on the encoding and decoding side.
Lizhi Hou, Yiling Xu
VCIP3
2024 Inter-Frame Coding for Dynamic Meshes via Coarse-to-Fine Anchor Mesh Generation
abstract
In the current Video-based Dynamic Mesh Coding (V-DMC) standard, inter-frame coding is restricted to mesh frames with constant topology. Consequently, temporal redundancy is not fully leveraged, resulting in suboptimal compression efficacy. To address this limitation, this paper introduces a novel coarse-to-fine scheme to generate anchor meshes for frames with time-varying topology. Initially, we generate a coarse anchor mesh using an octree-based nearest neighbor search. Motion estimation compensates for regions with significant motion changes during this process. However, the quality of the coarse mesh is low due to its suboptimal vertices. To enhance details, the fine anchor mesh is further optimized using the Quadric Error Metrics (QEM) algorithm to calculate more precise anchor points. The inter-frame anchor mesh generated herein retains the connectivity of the reference base mesh, while concurrently preserving superior quality. Experimental results show that our method achieves 7.2% ∼ 10.3% BD-rate gain compared to the existing V-DMC test model version 7.
Lizhi Hou, Qi Yang 0003, Yiling Xu
VCIP4
2024 Differentiable Low-computation Global Correlation Loss for Monotonicity Evaluation in Quality Assessment
abstract
In this paper, we propose a global monotonicity consistency training strategy for quality assessment, which includes a differentiable, low-computation monotonicity evaluation loss function and a global perception training mechanism. Specifically, unlike conventional ranking loss and linear programming approaches that indirectly implement the Spearman rank-order correlation coefficient (SROCC) function, our method directly converts SROCC into a loss function by making the sorting operation within SROCC differentiable and functional. Furthermore, to mitigate the discrepancies between batch optimization during network training and global evaluation of SROCC, we introduce a memory bank mechanism. This mechanism stores gradient-free predicted results from previous batches and uses them in the current batch’s training to prevent abrupt gradient changes. We evaluate the performance of the proposed method on both images and point clouds quality assessment tasks, demonstrating performance gains in both cases.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu
VCIP3
2024 Explicit-NeRF-QA: A Quality Assessment Database for Explicit NeRF Model Compression
abstract
In recent years, Neural Radiance Fields (NeRF) have demonstrated significant advantages in representing and synthesizing 3D scenes. Explicit NeRF models facilitate the practical NeRF applications with faster rendering speed, and also attract considerable attention in NeRF compression due to its huge storage cost. To address the challenge of the NeRF compression study, in this paper, we construct a new dataset, called Explicit-NeRF-QA. We use 22 3D objects with diverse geometries, textures, and material complexities to train four typical explicit NeRF models across five parameter levels. Lossy compression is introduced during the model generation, pivoting the selection of key parameters such as hash table size for InstantNGP and voxel grid resolution for Plenoxels. By rendering NeRF samples to processed video sequences (PVS), a large scale subjective experiment with lab environment is conducted to collect subjective scores from 21 viewers. The diversity of content, accuracy of mean opinion scores (MOS), and characteristics of NeRF distortion are comprehensively presented, establishing the heterogeneity of the proposed dataset. The state-of-the-art objective metrics are tested in the new dataset. Best Pearson correlation, which is around 0.85, is collected from the full-reference objective metric. All tested no-reference metrics report very poor results with 0.4 to 0.6 correlations, demonstrating the need for further development of more robust no-reference metrics. The dataset, including NeRF samples, source 3D objects, multiview images for NeRF generation, PVSs, MOS, is made publicly available at the following location: https://github.com/YukeXing/Explicit-NeRF-QA.
Yuke Xing, Qi Yang 0003, Kaifa Yang, Yiling Xu, Zhu Li 0001
VCIP4
2024 Perception-Guided Quality Metric of 3D Point Clouds Using Hybrid Strategy
abstract
Full-reference point cloud quality assessment (FR-PCQA) aims to infer the quality of distorted point clouds with available references. Most of the existing FR-PCQA metrics ignore the fact that the human visual system (HVS) dynamically tackles visual information according to different distortion levels (i.e., distortion detection for high-quality samples and appearance perception for low-quality samples) and measure point cloud quality using unified features. To bridge the gap, in this paper, we propose a perception-guided hybrid metric (PHM) that adaptively leverages two visual strategies with respect to distortion degree to predict point cloud quality: to measure visible difference in high-quality samples, PHM takes into account the masking effect and employs texture complexity as an effective compensatory factor for absolute difference; on the other hand, PHM leverages spectral graph theory to evaluate appearance degradation in low-quality samples. Variations in geometric signals on graphs and changes in the spectral graph wavelet coefficients are utilized to characterize geometry and texture appearance degradation, respectively. Finally, the results obtained from the two components are combined in a non-linear method to produce an overall quality score of the tested point cloud. The results of the experiment on five independent databases show that PHM achieves state-of-the-art (SOTA) performance and offers significant performance improvement in multiple distortion environments. The code is publicly available at https://github.com/zhangyujie-1998/PHM.
Qi Yang 0003, Yiling Xu, Shan Liu 0001
IEEE Trans. Image Process.3
2024 GPA-Net:No-Reference Point Cloud Quality Assessment With Multi-Task Graph Convolutional Network
abstract
With the rapid development of 3D vision, point cloud has become an increasingly popular 3D visual media content. Due to the irregular structure, point cloud has posed novel challenges to the related research, such as compression, transmission, rendering and quality assessment. In these latest researches, point cloud quality assessment (PCQA) has attracted wide attention due to its significant role in guiding practical applications, especially in many cases where the reference point cloud is unavailable. However, current no-reference metrics which based on prevalent deep neural network have apparent disadvantages. For example, to adapt to the irregular structure of point cloud, they require preprocessing such as voxelization and projection that introduce extra distortions, and the applied grid-kernel networks, such as Convolutional Neural Networks, fail to extract effective distortion-related features. Besides, they rarely consider the various distortion patterns and the philosophy that PCQA should exhibit shift, scaling, and rotation invariance. In this paper, we propose a novel no-reference PCQA metric named the Graph convolutional PCQA network (GPA-Net). To extract effective features for PCQA, we propose a new graph convolution kernel, i.e., GPAConv, which attentively captures the perturbation of structure and texture. Then, we propose the multi-task framework consisting of one main task (quality regression) and two auxiliary tasks (distortion type and degree predictions). Finally, we propose a coordinate normalization module to stabilize the results of GPAConv under shift, scale and rotation transformations. Experimental results on two independent databases show that GPA-Net achieves the best performance compared to the state-of-the-art no-reference PCQA metrics, even better than some full-reference metrics in some cases.
Ziyu Shan, Qi Yang 0003, Rui Ye 0001, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
IEEE Trans. Vis. Comput. Graph.5
2024 TCDM: Transformational Complexity Based Distortion Metric for Perceptual Point Cloud Quality Assessment
abstract
The goal of objective point cloud quality assessment (PCQA) research is to develop quantitative metrics that measure point cloud quality in a perceptually consistent manner. Merging the research of cognitive science and intuition of the human visual system (HVS), in this article, we evaluate the point cloud quality by measuring the complexity of transforming the distorted point cloud back to its reference, which in practice can be approximated by the code length of one point cloud when the other is given. For this purpose, we first make space segmentation for the reference and distorted point clouds based on a 3D Voronoi diagram to obtain a series of local patch pairs. Next, inspired by the predictive coding theory, we utilize a space-aware vector autoregressive (SA-VAR) model to encode the geometry and color channels of each reference patch with and without the distorted patch, respectively. Assuming that the residual errors follow the multi-variate Gaussian distributions, the self-complexity of the reference and transformational complexity between the reference and distorted samples are computed using covariance matrices. Additionally, the prediction terms generated by SA-VAR are introduced as one auxiliary feature to promote the final quality prediction. The effectiveness of the proposed transformational complexity based distortion metric (TCDM) is evaluated through extensive experiments conducted on five public point cloud quality assessment databases. The results demonstrate that TCDM achieves state-of-the-art (SOTA) performance, and further analysis confirms its robustness in various scenarios.
Qi Yang 0003, Xiaozhong Xu, Le Yang 0001, Yiling Xu
IEEE Trans. Vis. Comput. Graph.6
2023 Improving Point Cloud Quality Metrics with Noticeable Possibility Maps
abstract
Point cloud quality assessment (PCQA) plays a vital role in the quality of experience (QoE) oriented data processing. To reflect the visual degradation introduced by various distortions, many PCQA metrics have been proposed in recent years. However, these metrics often take all distortions indiscriminately into account, ignoring the fact that some distortions are below the noticeable threshold and thus do not affect subjective perception. To solve this problem, we involve the characteristic of just noticeable difference (JND) into PCQA. Specifically, we first rotate the reference and distorted point cloud repeatedly to obtain multiple perspectives, then utilize the 2D JND models to derive the 3D noticeable possibility maps (NPM) to infer the possibility that the distortion is perceivable at each point. Utilizing the generated NPM, we modify current point-wise and structure-wise quality metrics to help them correlate better with subjective perception. Extensive experiments show the universal effectiveness of the proposed NPM in improving PCQA metrics. Code will be available at https://github.com/NekoooooOoi/NPM.
Qi Yang 0003, Yiling Xu, Jun Sun 0005, Shan Liu 0001
ICME4
2023 Exploring the Influence of View and Camera Path Selection for Dynamic Mesh Quality Assessment
abstract
With the development of 3D mesh processing and applications, the quality assessment of dynamic mesh sequences attracts more and more attention. One prevalent strategy for performing the subjective experiment and designing objective quality metrics is to convert the 3D dynamic mesh into 2D images or videos via projection and to collect subjective scores or calculate objective indexes based on these images or videos. In this paper, we study the influence of the view, or camera path selection for the projection, for both subjective and objective dynamic mesh quality assessment, and compare the performance of image-based metrics and point-based metrics corresponding to the collected subjective scores. First, we use the dynamic mesh sequences proposed by MPEG as anchors and generate videos corresponding to different coding configurations and different camera paths. Then, we conduct subjective experiments to collect the ground truth of mean opinion scores. Besides, we calculate the state-of-the-art objective metric scores for each sequence. We analyze the differences between subjective scores with respect to different camera paths and the correlation between subjective scores and objective metrics. The results show that different camera paths tend to generate close subjective perceptions and that the selection of views can influence some objective metrics.
Kaifa Yang, Qi Yang 0003, Joël Jung, Yiling Xu, Xiaozhong Xu, Shan Liu 0001
ICME4
2023 Learning Dynamic Point Cloud Compression via Hierarchical Inter-frame Block Matching
abstract
3D dynamic point cloud (DPC) compression relies on mining its temporal context, which faces significant challenges due to DPC's sparsity and non-uniform structure. Existing methods are limited in capturing sufficient temporal dependencies. Therefore, this paper proposes a learning-based DPC compression framework via hierarchical block-matching-based inter-prediction module to compensate and compress the DPC geometry in latent space. Specifically, we propose a hierarchical motion estimation and motion compensation (Hie-ME/MC) framework for flexible inter-prediction, which dynamically selects the granularity of optical flow to encapsulate the motion information accurately. To improve the motion estimation efficiency of the proposed inter-prediction module, we further design a KNN-attention block matching (KABM) network that determines the impact of potential corresponding points based on the geometry and feature correlation. Finally, we compress the residual and the multi-scale optical flow with a fully-factorized deep entropy model. The experiment result on the MPEG-specified Owlii Dynamic Human Dynamic Point Cloud (Owlii) dataset shows that our framework outperforms the previous state-of-the-art methods and the MPEG standard V-PCC v18 in inter-frame low-delay mode.
Shuting Xia, Tingyu Fan, Yiling Xu, Jenq-Neng Hwang, Zhu Li 0001
ACM Multimedia3
2023 DMGC: Deep Triangle Mesh Geometry Compression via Connectivity Prediction
abstract
We propose a novel deep lossless geometry compression algorithm for triangle mesh. Typical traditional triangle mesh compression algorithms are connectivity-driven, which first codes connectivity and then codes vertices according to the encoded connectivity. However, vertex compression is inefficient since these algorithms lack the exploitation of vertices' intra-component redundancy, which point cloud compression (PCC) excels at. Therefore, our approach first compresses the vertices with a lossless PCC algorithm since the vertices could be considered as a point cloud. Moreover, based on the encoded vertices, the bitrate of connectivity could be decreased further by exploiting the cross-component redundancy between vertex and connectivity. Specifically, we divide the connectivity into KNN connectivity and isolated connectivity. We design a deep entropy model for KNN connectivity compression. This model extracts the spatial feature of encoded vertices first, then predicts the vertex-pairs' connection probabilities using the spatial feature. In order to exploit the intra-component redundancy, an auto-regressive strategy is also employed when predicting. The predicted probabilities are finally fed into a binary arithmetic coder to code the KNN connectivity into a compact bitstream. The isolated connectivity is encoded in direct coding mode (DCM). To our knowledge, this paper is the first work to utilize the deep neural network on mesh compression. We validate the effectiveness of our method on simplified MPEG V-DMC dataset. Experimental results demonstrate that the proposed method achieves average 7.33% and 40.75% bpv gains on connectivity and vertex separately, resulting average 30.26% bpv gains on total mesh compression compared with MPEG SC3DMC.
Xinyao Zeng, Linyao Gao, Yiling Xu, Yanfeng Wang 0003
NOSSDAV4
2023 Point Cloud Quality Assessment using 3D Saliency Maps
abstract
Point cloud quality assessment (PCQA) has become an appealing research field in recent days. Considering the importance of saliency detection in quality assessment, we propose an effective full-reference PCQA metric which makes an attempt to utilize the saliency information to facilitate quality prediction, called point cloud quality assessment using 3D saliency maps (PQSM). Specifically, we first propose a projectionbased point cloud saliency map generation method, in which depth information is introduced to better reflect the geometric characteristics of point clouds. Then, we construct point cloud local neighborhoods to derive three structural descriptors to indicate the geometry, color and saliency discrepancies. Finally, a saliency-based pooling strategy is proposed to generate the final quality score. Extensive experiments are performed on four independent PCQA databases. The results demonstrate that the proposed PQSM shows competitive performances compared to multiple state-of-the-art PCQA metrics.
Qi Yang 0003, Yiling Xu, Jun Sun 0005, Shan Liu 0001
VCIP4
2023 MPED: Quantifying Point Cloud Distortion Based on Multiscale Potential Energy Discrepancy
abstract
In this article, we propose a new distortion quantification method for point clouds, the multiscale potential energy discrepancy (MPED). Currently, there is a lack of effective distortion quantification for a variety of point cloud perception tasks. Specifically, in human vision tasks, a distortion quantification method is used to predict human subjective scores and optimize the selection of human perception task parameters, such as dense point cloud compression and enhancement. In machine vision tasks, a distortion quantification method usually serves as loss function to guide the training of deep neural networks for unsupervised learning tasks (e.g., sparse point cloud reconstruction, completion, and upsampling). Therefore, an effective distortion quantification should be differentiable, distortion discriminable, and have low computational complexity. However, current distortion quantification cannot satisfy all three conditions. To fill this gap, we propose a new point cloud feature description method, the point potential energy (PPE), inspired by classical physics. We regard the point clouds are systems that have potential energy and the distortion can change the total potential energy. By evaluating various neighborhood sizes, the proposed MPED achieves global-local tradeoffs, capturing distortion in a multiscale fashion. We further theoretically show that classical Chamfer distance is a special case of our MPED. Extensive experiments show that the proposed MPED is superior to current methods on both human and machine perception tasks. Our code is available at https://github.com/Qi-Yangsjtu/MPED.
Qi Yang 0003, Siheng Chen, Yiling Xu, Jun Sun 0005, Zhan Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Multiscale Latent-Guided Entropy Model for LiDAR Point Cloud Compression
abstract
The non-uniform distribution and extremely sparse nature of the LiDAR point cloud (LPC) bring significant challenges to its high-efficient compression. This paper proposes a novel end-to-end, fully-factorized deep framework that represents the original LiDAR point cloud into an octree structure and hierarchically constructs the octree entropy model in layers. The proposed framework utilizes a hierarchical latent variable as side information to encapsulate the sibling and ancestor dependence, which provides sufficient context information for the modeling of point cloud distribution while enabling the parallel encoding and decoding of octree nodes in the same layer. Besides, we propose a residual coding framework for the compression of the latent variable, which explores the spatial correlation of each layer by progressive downsampling, and model the corresponding residual with a fully-factorized entropy model. Furthermore, we propose soft addition and subtraction for residual coding to improve network flexibility. The comprehensive experiment results on the LiDAR benchmark SemanticKITTI and MPEG-specified dataset Ford demonstrate that our proposed framework achieves state-of-the-art performance among all the previous LPC frameworks. Besides, our end-to-end, fully-factorized framework is proved by experiment to be high-parallelized and time-efficient, which saves more than 99.8% of decoding time compared to previous state-of-the-art methods on LPC compression.
Tingyu Fan, Linyao Gao, Yiling Xu, Dong Wang 0004, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Point Cloud Quality Assessment: Dataset Construction and Learning-based No-reference Metric
abstract
Full-reference (FR) point cloud quality assessment (PCQA) has achieved impressive progress in recent years. However, in many cases, obtaining the reference point clouds is difficult, so no-reference (NR) metrics have become a research hotspot. Few researches about NR-PCQA are carried out due to the lack of a large-scale PCQA dataset. In this article, we first build a large-scale PCQA dataset named LS-PCQA, which includes 104 reference point clouds and more than 22,000 distorted samples. In the dataset, each reference point cloud is augmented with 31 types of impairments (e.g., Gaussian noise, contrast distortion, local missing, and compression loss) at 7 distortion levels. Besides, each distorted point cloud is assigned with a pseudo-quality score as its substitute of Mean Opinion Score. Inspired by the hierarchical perception system and considering the intrinsic attributes of point clouds, we propose a NR metric ResSCNN based on sparse convolutional neural network (CNN) to accurately estimate the subjective quality of point clouds. We conduct several experiments to evaluate the performance of the proposed NR metric. The results demonstrate that ResSCNN exhibits the state-of-the-art performance among all the existing NR-PCQA metrics and even outperforms some FR metrics. The dataset presented in this work will be made publicly accessible at https://smt.sjtu.edu.cn . The source code for the proposed ResSCNN can be found at https://github.com/lyp22/ResSCNN .
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu, Le Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 3DAC: Learning Attribute Compression for Point Clouds
abstract
We study the problem of attribute compression for large-scale unstructured 3D point clouds. Through an in-depth exploration of the relationships between different encoding steps and different attribute channels, we introduce a deep compression network, termed 3DAC, to explicitly compress the attributes of 3D point clouds and reduce storage usage in this paper. Specifically, the point cloud attributes such as color and reflectance are firstly converted to transform coefficients. We then propose a deep entropy model to model the probabilities of these coefficients by considering information hidden in attribute transforms and previous encoded attributes. Finally, the estimated probabilities are used to further compress these transform coefficients to a final attributes bitstream. Extensive experiments conducted on both indoor and outdoor large-scale open point cloud datasets, including ScanNet and SemanticKITTI, demonstrated the superior compression rates and reconstruction quality of the proposed method.
Guangchi Fang, Qingyong Hu, Hanyun Wang, Yiling Xu, Yulan Guo
CVPR4
2022 No-Reference Point Cloud Quality Assessment via Domain Adaptation
abstract
We present a novel no-reference quality assessment metric, the image transferred point cloud quality assessment (IT-PCQA), for 3D point clouds. For quality assessment, deep neural network (DNN) has shown compelling performance on no-reference metric design. However, the most challenging issue for no-reference PCQA is that we lack large-scale subjective databases to drive robust networks. Our motivation is that the human visual system (HVS) is the decision-maker regardless of the type of media for quality assessment. Leveraging the rich subjective scores of the natural images, we can quest the evaluation criteria of human perception via DNN and transfer the capability of prediction to 3D point clouds. In particular, we treat natural images as the source domain and point clouds as the target domain, and infer point cloud quality via unsupervised adversarial domain adaptation. To extract effective latent features and minimize the domain discrepancy, we propose a hierarchical feature encoder and a conditional-discriminative network. Considering that the ultimate pur-pose is regressing objective score, we introduce a novel con-ditional cross entropy loss in the conditional-discriminative network to penalize the negative samples which hinder the convergence of the quality regression network. Experi-mental results show that the proposed method can achieve higher performance than traditional no-reference metrics, even comparable results with full-reference metrics. The proposed method also suggests the feasibility of assessing the quality of specific media content without the expensive and cumbersome subjective evaluations. Code is available at https://github.com/Qi-Yangsjtu/IT-PCQA.
Qi Yang 0003, Yipeng Liu 0003, Siheng Chen, Yiling Xu, Jun Sun 0005
CVPR4
2022 D-DPCC: Deep Dynamic Point Cloud Compression via 3D Motion Prediction
abstract
The non-uniformly distributed nature of the 3D Dynamic Point Cloud (DPC) brings significant challenges to its high-efficient inter-frame compression. This paper proposes a novel 3D sparse convolution-based Deep Dynamic Point Cloud Compression (D-DPCC) network to compensate and compress the DPC geometry with 3D motion estimation and motion compensation in the feature space. In the proposed D-DPCC network, we design a Multi-scale Motion Fusion (MMF) module to accurately estimate the 3D optical flow between the feature representations of adjacent point cloud frames. Specifically, we utilize a 3D sparse convolution-based encoder to obtain the latent representation for motion estimation in the feature space and introduce the proposed MMF module for fused 3D motion embedding. Besides, for motion compensation, we propose a 3D Adaptively Weighted Interpolation (3DAWI) algorithm with a penalty coefficient to adaptively decrease the impact of distant neighbours. We compress the motion embedding and the residual with a lossy autoencoder-based network. To our knowledge, this paper is the first work proposing an end-to-end deep dynamic point cloud compression framework. The experimental result shows that the proposed D-DPCC framework achieves an average 76% BD-Rate (Bjontegaard Delta Rate) gains against state-of-the-art Video-based Point Cloud Compression (V-PCC) v13 in inter mode.
Tingyu Fan, Linyao Gao, Yiling Xu, Zhu Li 0001, Dong Wang 0004
IJCAI3
2022 Learning-based Intra-Prediction For Point Cloud Attribute Transform Coding
abstract
The attached attributes of each point in point cloud are fairly valuable but aggravate the burden for storage and transmission. In this paper, we propose a learning-based intra-prediction method for region adaptive hierarchical transform (RAHT) to compress point cloud attributes efficiently. First, we design an adaptive neighbor selection (ANS) module to produce the most correlated neighbors for child nodes. Then, the correlated neighbors obtained by ANS, the corresponding distance weights, and some additional auxiliary information are concatenated and fed to a multi-layer perception (MLP) based network to estimate the child node attributes precisely. Besides, residual learning is introduced to accelerate the network convergence. The predicted child node attributes are finally transformed by RAHT, and the residual of transform coefficients are then quantized and entropy coded. Experimental results demonstrate that our proposed methods can significantly improve attribute coding efficiency with average 10.2% BD-Rate gains compared with MPEG G-PCC reference software TMC13v14.0 on MPEG PCC dataset.
Lizhi Hou, Linyao Gao, Yiling Xu, Zhu Li 0001, Xiaozhong Xu, Shan Liu 0001
MMSP3
2022 Reduced Reference Quality Assessment for Point Cloud Compression
abstract
In this paper, we propose a reduced reference (RR) point cloud quality assessment (PCQA) model named R-PCQA to quantify the distortions introduced by the lossy compression. Specifically, we use the attribute and geometry quantization steps of different compression methods (i.e., V-PCC, G-PCC and AVS) to infer the point cloud quality, assuming that the point clouds have no other distortions before compression. First, we analyze the compression distortion of point clouds under separate attribute compression and geometry compression to avoid their mutual masking, for which we consider 5 point clouds as references to generate a compression dataset (PCCQA) containing independent attribute compression and geometry compression samples. Then, we develop the proposed R-PCQA via fitting the relationship between the quantization steps and the perceptual quality. We evaluate the performance of R-PCQA on both the established dataset and another independent dataset. The results demonstrate that the proposed R-PCQA can exhibit reliable performance and high generalization ability.
Yipeng Liu 0003, Qi Yang 0003, Yiling Xu
VCIP3
2022 Inferring Point Cloud Quality via Graph Similarity
abstract
Objective quality estimation of media content plays a vital role in a wide range of applications. Though numerous metrics exist for 2D images and videos, similar metrics are missing for 3D point clouds with unstructured and non-uniformly distributed points. In this paper, we propose [Formula: see text]-a metric to accurately and quantitatively predict the human perception of point cloud with superimposed geometry and color impairments. Human vision system is more sensitive to the high spatial-frequency components (e.g., contours and edges), and weighs local structural variations more than individual point intensities. Motivated by this fact, we use graph signal gradient as a quality index to evaluate point cloud distortions. Specifically, we first extract geometric keypoints by resampling the reference point cloud geometry information to form an object skeleton. Then, we construct local graphs centered at these keypoints for both reference and distorted point clouds. Next, we compute three moments of color gradients between centered keypoint and all other points in the same local graph for local significance similarity feature. Finally, we obtain similarity index by pooling the local graph significance across all color channels and averaging across all graphs. We evaluate [Formula: see text] on two large and independent point cloud assessment datasets that involve a wide range of impairments (e.g., re-sampling, compression, and additive noise). [Formula: see text] provides state-of-the-art performance for all distortions with noticeable gains in predicting the subjective mean opinion score (MOS) in comparison with point-wise distance-based metrics adopted in standardized reference software. Ablation studies further show that [Formula: see text] can be generalized to various scenarios with consistent performance by adjusting its key modules and parameters. Models and associated materials will be made available at https://njuvision.github.io/GraphSIM or http://smt.sjtu.edu.cn/papers/GraphSIM.
Qi Yang 0003, Zhan Ma 0001, Yiling Xu, Zhu Li 0001, Jun Sun 0005
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Point Cloud Geometry Compression Via Neural Graph Sampling
abstract
Compressing point cloud geometry (PCG) efficiently is of great interests for enabling abundant networked applications, because PCG is a promising representation to precisely illustrate arbitrary-shaped 3D objects and relevant physical scenes. To well exploit the unconstrained geometric correlation of input PCG, a three-step neural graph sampling (NGS) is developed. First, we construct the local graph of each point using its K nearest neighbors according to the Euclidean distance metric; Second, for each local graph, its graph center point expands associated feature attribute by aggregating neighbor weights via point-wise dynamic filter; We then perform attention-based sampling to select a subset of points to well represent input points. The proposed NGS is embedded into an end-to-end analysis/synthesis-based variational autoencoder (VAE), with which the encoder applies multiscale NGS to extract latent keypoints that are augmented with neighbor structures and compressed at bottleneck leveraging the hyperpriors for accurate entropy modeling, and the decoder directly uses layered convolutions to refine progressively for the reconstruction of final point cloud. Note that all computations are fulfilled using point-wise convolution, making our solution an attractive approach in practice. Experimental results demonstrate that the proposed method using NGS mechanism outperforms the state-of-the-art point-based PCG compression methods by more than $2\times \mathrm{B}\mathrm{D}$-Rate (Bjûntegaard Delta Rate) gains, and several orders of magnitude gains over the MPEG G-PCC across all testing categories on ShapeNetCorev2 dataset.
Linyao Gao, Tingyu Fan, Jianqiang Wan, Yiling Xu, Jun Sun 0005, Zhan Ma 0001
ICIP4
2021 Visual Quality Optimization for View-Dependent Point Cloud Compression
abstract
The video-based point cloud compression (V-PCC) is the state-of-the-art dynamic point cloud compression technique. V-PCC projects the 3D point cloud data patch by patch to its bounding box and organizes projected patches into a video frame, making full use of the well-developed video coding tools. Despite its high efficiency, cracks easily exist in the reconstructed point cloud in various viewing angles, which seriously degrades the visual quality. In this paper, we propose an efficient method to improve the visual quality of dynamic point cloud, especially for the main view from the content provider. The relationship between patches and views is exploited, and an algorithm intelligently reserving points that may be discarded in V-PCC is proposed. According to our subjective and perceptual objective evaluation experiments, compared with V-PCC, the overall visual quality of the reconstructed point could is evidently improved. In particular, cracks are mended with our proposed method. The Bjontegaard delta bit-rate reduction of up to 3.1% is achieved with respect to Point Cloud Quality Metric (PCQM), which partially verifies the improvement of subjective quality when adopting the proposed method.
Danying Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001
ISCAS4
2021 MS-GraphSIM: Inferring Point Cloud Quality via Multiscale Graph Similarity
abstract
To address the point cloud quality assessment (PCQA) problem, GraphSIM was proposed via jointly considering geometrical and color features, which shows compelling performance in multiple distortion detection. However, GraphSIM does not take into account the mutiscale characteristics of human perception. In this paper, we propose a multiscale PCQA model, called Multiscale Graph Similarity (MS-GraphSIM), that can better predict human subjective perception. First, exploring the multiscale processing method used in image processing, we introduce a multiscale representation of point clouds based on graph signal processing. Second, we extend GraphSIM into multiscale version based on the proposed multiscale representation. Specifically, MS-GraphSIM constructs a multiscale representation for each local patch extracted from the reference point cloud or the distorted point cloud, and then fuses GraphSIM at different scales to obtain an overall quality score. Experiment results demonstrate that the proposed MS-GraphSIM outperforms the state-of-the-art PCQA metrics over two fairly large and independent databases. Ablation studies further prove the proposed MS-GraphSIM is robust to different model hyperparameter settings. The code is available at https://github.com/zyj1318053/MS_GraphSIM.
Qi Yang 0003, Yiling Xu
ACM Multimedia3
2021 Point-Voting based Point Cloud Geometry Compression
abstract
The Geometry-based Point Cloud Compression (G-PCC) proposed by the Moving Picture Experts Group (MPEG) is the state-of-art point cloud compression algorithm. It provides an efficient lossy geometry compression technique called triangle soup (Trisoup) for static point clouds. Based on the pruned octree structure, Trisoup provides a local surface model consisting of multiple triangles and compresses vertices of the triangles instead of directly compressing the positions of the original points. Accordingly, we propose a point-voting based method to improve the triangle-construction within each leaf node. This new method leverages the node-based points distribution for more precise vertices determination, which better fits the local surface. Experimental results demonstrate the effectiveness of our point-voting based method for both objective and subjective quality evaluation.
Chaofei Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001
MMSP4
2021 Part-level attention networks for cross-domain person re-identification
abstract
Abstract Person re‐identification (Re‐ID) is in significant demand for intelligent security and single or multiple‐target tracking. However, there are issues in the person Re‐ID tasks, such as sharp decline in cross‐data sets detection accuracy, poor generalization and cross‐domain ability of the model. This work mainly studies the generalization and adaptation of cross‐domain person Re‐ID models. Different from most existing methods for cross‐domain Re‐ID tasks, the authors use diversified spatial semantic feature in pixel‐level learning in the target domain to improve the generality and adaptability of the model. In the case that no information of the target domain is used during the model training, the trained model is directly tested on the data set of the target domain. It has proven effective to add the attention cascade module into the backbone network combining with the part‐level branch. The authors conducted extensive experiments based on the three data sets of Market‐1501, DukeMTMC‐ReID and MSMT17, resulting in both single‐domain and cross‐domain tests with an average improvement of Rank1 and mAP values of about 10% compared with Baseline through the authors' proposed method named Part‐Level Attention Network.
Nisuo Du, Zhi Ouyang, Ning Kang 0012, Qing He 0007, Yiling Xu, Shichun Ge, Jingkuan Song
IET Image Process.8
2021 Secure boot, trusted boot and remote attestation for ARM TrustZone-based IoT Nodes
Zhen Ling 0001, Huaiyu Yan, Xinhui Shao, Junzhou Luo, Yiling Xu, Bryan Pearson, Xinwen Fu
J. Syst. Archit.5
2021 A Dual Camera System for High Spatiotemporal Resolution Video Acquisition
abstract
This paper presents a dual camera system for high spatiotemporal resolution (HSTR) video acquisition, where one camera shoots a video with high spatial resolution and low frame rate (HSR-LFR) and another one captures a low spatial resolution and high frame rate (LSR-HFR) video. Our main goal is to combine videos from LSR-HFR and HSR-LFR cameras to create an HSTR video. We propose an end-to-end learning framework, AWnet, mainly consisting of a FlowNet and a FusionNet that learn an adaptive weighting function in pixel domain to combine inputs in a frame recurrent fashion. To improve the reconstruction quality for cameras used in reality, we also introduce noise regularization under the same framework. Our method has demonstrated noticeable performance gains in terms of both objective PSNR measurement in simulation with different publicly available video and light-field datasets and subjective evaluation with real data captured by dual iPhone 7 and Grasshopper3 cameras. Ablation studies are further conducted to investigate and explore various aspects, such as reference structure, camera parallax, exposure time, etc) of our system to fully understand its capability for potential applications.
Zhan Ma 0001, Muhammad Salman Asif, Yiling Xu, Wenbo Bao, Jun Sun 0005
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 View-Dependent Dynamic Point Cloud Compression
abstract
Dynamic point cloud (DPC) captures 3D object and scene with realistic appearance to mimic the natural reality. It inherently offers the six-degree-of-freedom (6DoF) for content consumption, which motivates us to facilitate the network-friendly view-dependent DPC streaming by leveraging the limited field of view of the human visual system (HVS) at specific moments, for significant network bandwidth reduction without the quality of experience (QoE) loss. Therefore, in this work, we propose the view-dependent DPC compression, noted as View-PCC, by which we can enable instantaneous and complete view reconstruction from partial streams, and 6DoF navigation by stream adaptation. These functionalities are offered by the hybrid global and local projection to effectively map 3D points to five perpendicular 2D image planes of a cube. Multi-view and multi-layer based global projection is first applied progressively and followed by the patch-based local projection with both intra-frame displaced arrangement and inter-frame temporal alignment, so as to maximize the number of effectively projected points and preserve the spatio-temporal coherency for better compression. Boundary padding is then augmented for each image plane that will be encoded using the successful 2D video coding standard - High-Efficiency Video Coding (HEVC). Our extensive simulations have shown the noticeable objective compression gains of View-PCC when compared with the typical octree-based approach and global-projection-based method. The subjective quality assessment suggests the better reconstruction quality of View-PCC, in comparison to the MPEG video-based point cloud compression (V-PCC).
Wenjie Zhu 0004, Zhan Ma 0001, Yiling Xu, Li Li 0040, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Lossy Point Cloud Geometry Compression via Region-Wise Processing
abstract
Point cloud geometry (PCG) is used to precisely represent arbitrary-shaped 3D objects and scenes, is of great interest to vast applications which puts forward the pressing desire of high-efficiency PCG compression for transmission and storage. Existing PCG coding mostly relies on the octree model by which point-wise processing is applied without exploring nonlocal regional geometry similarity across the entire 3D surface. This work, instead, suggests the region-wise processing to leverage the region similarity to exploit inter-region redundancy for efficient lossy point cloud geometry compression. Towards this goal, a given PCG is first segmented into numerous local regions each of which comprises a portion of point cloud surface, and can be represented by a surface vector that describes the geometry shape numerically in a projected principal space. Subsequently, these regions are grouped into several discriminative clusters, assuring that inter-cluster similarity is minimized and intra-cluster similarity is maximized simultaneously, where the similarity is calculated using the regional surface vectors. In each cluster, we set a reference region having the largest similarity score to the others, which enables the non-reference region prediction from the reference one using alignment transform. In the end, we encode the reference regions directly using the lossless mode of the Geometry-based Point Cloud Compression (G-PCC), while corresponding non-reference regions are signaled using associated transform parameters. Compared with the state-of-the-art G-PCC using octree model, our region-wise approach can offer remarkable coding efficiency improvement, e.g., 32.4% and 22.0% Bjontegaard-delta rate (BD-Rate) gains for respective point-to-point ($D1$) and point-to-plane ($D2$) distortion evaluations, across a variety of common test sequences used in standard committee.
Wenjie Zhu 0004, Yiling Xu, Dandan Ding, Zhan Ma 0001, Mike Nilsson
IEEE Trans. Circuits Syst. Video Technol.2
2021 Learned Resolution Scaling Powered Gaming-as-a-Service at Scale
abstract
Built on the explosive advancement of cloud and telecommunication technologies, Gaming-as-a-Service (GaaS) or cloud gaming system is expected to revolutionize the traditional multi-billion video game market in the near future. This wave is analogous to the rise of live-video-streaming-based-Netflix to replace conventional DVD rental business for movies and TVs. In practice, a successful GaaS platform need to operate in a transparent mode without requiring substantial efforts from both content providers and end users, and offer the pristine quality of experience (QoE) at an affordable cost. Our analysis suggests that GaaS provisioning cost can be reduced significantly by enforcing the game video rendering and streaming at a lower resolution (so as to increase the user concurrency in the cloud and reduce the streaming bandwidth over the network). However, streaming video at a lower resolution may deteriorate the QoE. To maintain the client QoE at the level using the default-native resolution for streaming or even enhance it, we introduce the learned resolution scaling (LRS), which leverages the computational capabilities at clients/edges to restore/improve the reconstructed image/video quality via stacked deep neural networks (DNN). We integrate this LRS into a commercialized GaaS platform - AnyGame, to study its efficiency and complexity quantitatively. Extensive real-life experiments have shown that LRS-powered AnyGame offers the state-of-the-art performance, and the lower operational cost, paving the road for a potential success of GaaS over the Internet. Additionally, we dive into proposed LRS via ablation studies to further demonstrate its consistent performance, including the discussions on trade-off between efficiency and complexity, alternative training sets, etc.
Hao Chen 0036, Ming Lu 0003, Zhan Ma 0001, Xu Zhang 0006, Yiling Xu, Qiu Shen, Wenjun Zhang 0001
IEEE Trans. Multim.5
2021 Predicting the Perceptual Quality of Point Cloud: A 3D-to-2D Projection-Based Exploration
abstract
Point cloud is emerged as a promising media format to represent realistic 3D objects or scenes in applications, such as virtual reality, teleportation, etc. How to accurately quantify the subjective point cloud quality for application-driven optimization, however, is still a challenging and open problem. In this paper, we attempt to tackle this problem in a systematic means. First, we produce a fairly large point cloud dataset where ten popular point clouds are augmented with seven types of impairments (e.g., compression, photometry/color noise, geometry noise, scaling) at six different distortion levels, and organize a formal subjective assessment with tens of subjects to collect mean opinion scores (MOS) for all 420 processed point cloud samples (PPCS). We then try to develop an objective metric that can accurately estimate the subjective quality. Towards this goal, we choose to project the 3D point cloud onto six perpendicular image planes of a cube for the color texture image and corresponding depth image, and aggregate image-based global (e.g., Jensen-Shannon (JS) divergence) and local features (e.g., edge, depth, pixel-wise similarity, complexity) among all projected planes for a final objective index. Model parameters are fixed constants after performing the regression using a small and independent dataset previously published. The proposed metric has demonstrated the state-of-the-art performance for predicting the subjective point cloud quality compared with multiple full-reference and no-reference models, e.g., the weighted peak signal-to-noise ratio (PSNR), structural similarity (SSIM), feature similarity (FSIM) and natural image quality evaluator (NIQE). The dataset is made publicly accessible athttp://smt.sjtu.edu.cnorhttp://vision.nju.edu.cnfor all interested audiences.
Qi Yang 0003, Hao Chen 0036, Zhan Ma 0001, Yiling Xu, Rongjun Tang, Jun Sun 0005
IEEE Trans. Multim.4
2020 Fast Video Saliency Detection based on Feature Competition
abstract
In this paper, we propose a light video saliency prediction model, named SalFCM, which achieves fixation detection rate of 110fps. It is known that the human attention is captured by objects that have always been present or newly appeared. To model this dynamic change, we propose an Inter-frame Feature Competition Module (IFCM) to make an adaptive choice between correlated and differential features of consecutive frames. Besides, it is noted that saliency is better explained by low-level rather than high-level features in some visual scenes. Hence, we design a Hierarchical Feature Competition Module (HFCM) to balance the influence of low-level and high-level features. Our model achieves a good trade-off between precision and processing speed. The developed SalFCM is evaluated on three video saliency datasets: DHF1K, Hollywood-2 and UCF-sports. We conduct ablation studies to verify the effectiveness of the proposed model.
Hang Yan 0006, Yiling Xu, Jun Sun 0005, Le Yang 0001, Wei Huang 0012
VCIP2
2020 Modeling the Perceptual Quality of Viewport Adaptive Omnidirectional Video Streaming
abstract
Instead of streaming the entire OmniDirectional Videos (ODVs) that are often sampled at ultra high definition and high frame rate, a viewport adaptive streaming is preferred in practice. We usually stream the High-Quality (HQ) content within current viewport, while Low-Quality (LQ) elsewhere to save the network bandwidth consumption. Such scheme would lead to a quality refinement after user adapts his/her focus to a new viewport. In this paper, we thus model the perceptual impact of the quality variations (through adjusting the Quantization Stepsize (QS or q) and Spatial Resolution (SR or s)) with respect to the Refinement Duration (RD or τ) when performing the refinement from an arbitrary LQ scale to an arbitrary HQ one. A number of quality variations are studied to cover sufficient use cases in practice, resulting in a unified analytical model, as a product of separable exponential functions that measure the QS and SR induced perceptual impacts in terms of the RD, and a perceptual index measuring the subjective quality of corresponding viewport video after refinement. This model is first validated in a managed lab environment via independent subjective assessments by constraining user's navigation to avoid unexpected noise, where both Pearson Correlation Coefficient (PCC) and Spearman's Rank Correlation Coefficient (SRCC) are around 0.97. We then extend the validations in a real-life viewport-dependent streaming system, still yielding PCC and SRCC about 0.96 when comparing collected subjective scores with model predictions.
Shaowei Xie, Yiling Xu, Qiu Shen, Zhan Ma 0001, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Efficient Mobile Video Streaming via Context-Aware RaptorQ-Based Unequal Error Protection
abstract
Mobile video streaming systems typically apply the forward error correction (FEC) at the application layer to cope with packet-level transmission errors, which complements the bit-level correction mechanisms at the physical layer. However, most existing works fail to exploit the block-level dependencies in both intra and interframe coding modes of a single-layer compressed video, and thus are less efficient for the prevailing H.264/AVC and/or H.265/HEVC compatible single-layer video application. To this end, we propose a low-complexity FEC, i.e., context-aware RaptorQ (CA-RQ) with unequal error protection (UEP), to improve the error recovery performance of the singlelayer mobile video streaming, through incorporating the blocklevel dependencies in the compressed video data. We use a packet-level video transmission distortion model that considers the dependencies in both spatial and temporal domains, to quantify the importance of video packets within a group of pictures (GoP). The compressed video packets are categorized and grouped into several classes according to their importance to construct the CA-RQ code with the UEP property. We provide a theoretical analysis on redundancy allocation bounds to demonstrate the superior performance of proposed CA-RQ over the standard RaptorQ code. In the meantime, extensive simulations have shown that our scheme not only offers much better subjective visual quality with less than 50% additional redundant symbols as compared to the Macroblock-Based UEP (MB-UEP) scheme, but also outperforms the MB-UEP and classical equal error protection (EEP)-based schemes, by a 0.45%'5.71% and 0.94%'6.78% margin, respectively, in reconstructed quality evaluated using the structural similarity (SSIM) index, across a reasonable range of redundancy proportions.
Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Zhan Ma 0001, Wenjun Zhang 0001
IEEE Trans. Multim.3
2019 Dynamic Point Cloud Geometry Compression via Patch-wise Polynomial Fitting
abstract
With the boosting requirements of realistic 3D modeling for immersive applications, advent of the newly-developed 3D point cloud has attracted great attention. Frankly, immersive experience using high data volume affirms the importance of efficient compression. Inspired by the video-based point cloud compression (V-PCC), we propose a novel point cloud compression algorithm based on polynomial fitting of proper patches. Moreover, the original point cloud is segmented into various patches. We generated corresponding depth maps via projection of all the patches by focusing on geometry information. Instead of directly compressing the absolute values, we utilized proper polynomial functions to fit in each patch to obtain the differences. Finally, it is satisfying to note that the fitting function effectively represents the patch-wise geometry information. Moreover, new depth maps are obtained with extremely small and stable values, which are more suitable for video-based compression. Different patch-wise fitting parameters are preserved and coded using lossless compression through the open source PAQ project. The proposed approach achieves a noticeable improvement in the compression efficiency while maintaining point cloud quality.
Yingzhan Xu, Wenjie Zhu 0004, Yiling Xu, Zhu Li 0001
ICASSP3
2019 Learned Quality Enhancement via Multi-Frame Priors for HEVC Compliant Low-Delay Applications
abstract
Networked video applications, e.g., video conferencing, often suffer from poor visual quality due to unexpected network fluctuation and limited bandwidth. In this paper, we have developed a Quality Enhancement Network (QENet) to reduce the video compression artifacts, leveraging the spatial and temporal priors generated by respective multi-scale convolutions spatially and warped temporal predictions in a recurrent fashion temporally. We have integrated this QENet as a stand-alone post-processing subsystem to the High-Efficiency Video Coding (HEVC) compliant decoder. Experimental results show that our QENet demonstrates the state-of-the-art performance against default in-loop filters in HEVC and other deep learning based methods with noticeable objective gains in Peak Signal-to-Noise Ratio (PSNR) and subjective gains visually.
Ming Lu 0003, Yiling Xu, Shiliang Pu, Qiu Shen, Zhan Ma 0001
ICIP3
2019 Vehicle Positioning and Ranging with Static Traffic Camera based on 2D-3D Tracking and Re-Projection
abstract
Vehicle positioning and ranging are the current research hotspots. To attain the competition goal of the MMSP Witcomm Challenge 2019 and promote the development of the autonomous driving technology, a novel framework is proposed in this paper, which combines the 2D object tracking and 3D reprojection methodologies. Firstly, a correlation filter using the deep convolutional features is designed to detect the bounding box of the moving objects which is achieved via finding the maximum response of the initial object in the convolutional feature maps. Next the homography matrix is calculated based on the image coordinate points and the corresponding world coordinate points to project the vehicle 2D boundary into the real world. The validity of the proposed framework is verified using the dataset provided by the MMSP Witcomm Challenge 2019 competition, where our method was awarded the second-place prize.
Zhan Song, Yipeng Liu 0003, Yiling Xu, Le Yang 0001
MMSP3
2019 T-Gaming: A Cost-Efficient Cloud Gaming System at Scale
abstract
Cloud gaming (CG) system could pursue both high-quality gaming experience via intensive computing, and ultimate convenience anywhere at anytime through any energy-constrained mobile devices. Despite the abundance of efforts devoted, state-of-the-art CG systems still suffer from multiple key limitations: expensive deployment cost, high bandwidth consumption and unsatisfied quality of experience (QoE). As a result, existing works are not widely adopted in reality. This paper proposes a Transparent Gaming framework called T-Gaming that allows users to play any popular high-end desktop/console games on-the-fly over the Internet. T-Gaming utilizes the off-the-shelf consumer GPUs without resorting to the expensive proprietary GPU virtualization (vGPU) technology to reduce the deployment cost. Moreover, it enables prioritized video encoding based on the human visual feature to reduce the bandwidth consumption without noticeable visual quality degradation. Last but not least, T-Gaming adopts adaptive real-time streaming based on deep reinforcement learning (RL) to improve user's QoE. To evaluate the performance of T-Gaming, we implement and test a prototype system in the real world. Compared with the existing cloud gaming systems, T-Gaming not only reduces the expense per user by 75 percent hardware cost reduction and 14.3 percent network cost reduction, but also improves the normalized average QoE by 3.6-27.9 percent.
Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Ju Ren 0001, Jingtao Fan, Zhan Ma 0001, Wenjun Zhang 0001
IEEE Trans. Parallel Distributed Syst.3
2018 Hierarchical Segmentation Based Point Cloud Attribute Compression
abstract
With the rapid development of 3D capture techniques, point cloud has attracted significant attentions in recent years. Due to the large data volume of point cloud, efficient compression algorithms are essential for reducing bandwidth and storage consumption. In this paper, we present a novel scheme for point cloud attribute compression based on hierarchical segmentation. In this case, both global segmentation in photometric space and local segmentation in geometric space are analyzed to split point cloud into clusters. An octree based traversal algorithm is introduced to obtain the attribute stream of each cluster. Then, an intra-cluster prediction method is applied to achieve lossless compression. Meanwhile, we map the attribute streams to uniform 2D grids and leverage image coding method to achieve satisfying lossy compression performance. Experimental results demonstrate that our scheme outperforms the previous MPEG scheme in terms of coding efficiency.
Wenjie Zhu 0004, Yiling Xu
ICASSP3
2018 Modeling the Perceptual Impact of Viewport Adaptation for Immersive Video
abstract
Immersive video offers the freedom to navigate inside the virtualized environment. Instead of streaming the entire bulky content, a viewport or field of view (FoV) adaptive streaming is preferred. We often stream the high-quality content within current viewport, but degraded-quality representation elsewhere, so as to reduce the network bandwidth consumption. We then could refine the quality when focusing to a new FoV. Therefore, in this work, we have attempted to model the perceptual response of the quality variations (through adapting the quantization and spatial resolution) with respect to the refinement duration, and reach at a product of two closed-form exponential functions that well explain the joint quantization and resolution induced quality impact. Analytical model is also cross-validated using another set of data with both Pearson and Spearman's rank rank correlations over 0.98. Our work would be devised to guide the bandwidth-quality optimized immersive video streaming.
Shaowei Xie, Yiling Xu, Qiaojian Qian, Qiu Shen, Zhan Ma 0001, Wenjun Zhang 0001
ISCAS2
2018 Selective Convolutional Features based Generalized-mean Pooling for Fine-grained Image Retrieval
abstract
Image retrieval with convolutional neural network (CNN) has obtained a lot of attention. In this paper, we focus on a more challenging task: fine-grained image retrieval. We propose a simple and effective feature aggregation method using generalized-mean pooling (GeM pooling), which can make better use of information from the output tensor of the convolutional layer. In addition, we propose a simple feature selection scheme to remove noise and background. Experimental results demonstrate that our aggregation method not only outperformed state-of-the-art aggregation methods for general image retrieval, but also reach up to the same level of existing aggregation method for fine-grained image retrieval, with more compact representation and less memory cost.
Zhuoqun Wang, Zhu Li 0001, Jun Sun 0005, Yiling Xu
VCIP4
2018 QoE based SDN heterogeneous LTE and WLAN multi-radio networks for multi-user access
abstract
The scarcity of the long term evolution (LTE) bandwidth and the increasing data demand call for more wireless local area networks (WLANs) to help the cellular offloading. A software-defined networking (SDN) based control approach is proposed to effectively utilize the heterogeneous LTE and WLAN radio bandwidth by operating the multi-radio interfaces simultaneously. The use of the deployed Wi-Fi access points (APs) for Internet access has been a common practice for most people, however, the total utility or the quality of experience (QoE) drops when users compete for a single Wi-Fi AP or there is no controller to aid them to allocate rates to each radio access technology (RAT). This problem will even aggravate when more APs are involved in wireless mobile networks. This paper investigates how heterogeneous resources should be coordinated and allocated to multi-user access of the LTE and Wi-Fi aggregation (LWA) network. The proposed application layer scheme can adjust the target users and resources adaptively based on estimated users states so that the total QoE attained by all users is maximized. The developed scheme can help the operator decide the associations of users to appropriate Wi-Fi APs and adaptively adjust rate allocations of each user on the associated Wi-Fi and LTE. The good performance and low complexity of the proposed scheme is validated in a network simulator (NS-3).
Wei Huang 0012, De Meng, Jenq-Neng Hwang, Jounsup Park, Yiling Xu, Wenjun Zhang 0001
WCNC5
2017 Best-effort projection based attribute compression for 3D point cloud
abstract
According to the characteristics of exquisite presentation and rapid capture, 3D point cloud has been widely applied in immersive media industry. Aiming at vivid rendering of objects or dynamic scenarios, each point is associated with corresponding attributes like color, normal or intensity which induce massive data capacity. Consequently, efficient attribute compression methods are essential for effective immersive media transmission and consumption in the current media system. In this paper, we propose a best-effort projection based compression method for point cloud attributes. To take advantage of the well-developed 2D compression algorithm, regularized 3D point cloud is projected onto specified planes as different views while position information and related attributes are preserved. Joint depth- and color-dependent block-wise prediction has been utilized to further reduce the inter-view redundancy between projected 2D images. The experimental results have shown that the point cloud is successfully reconstructed via corresponding decoding process. Our method has presented strong competitiveness in both lossless and lossy compression of attributes of 3D point cloud.
Lanyi He, Wenjie Zhu 0004, Yiling Xu
APCC3
2017 An End-to-End View of IoT Security and Privacy
abstract
In this paper, we present an end-to-end view of IoT security and privacy and a case study. Our contribution is twofold. First, we present our end-to-end view of an IoT system and this view can guide risk assessment and design of an IoT system. We identify 10 basic IoT functionalities that are related to security and privacy. Based on this view, we systematically present security and privacy requirements in terms of IoT system, software, networking and big data analytics in the cloud. Second, using the end-to-end view of IoT security and privacy, we present a vulnerability analysis of the Edimax IP camera system. We are the first to exploit this system and have identified various attacks that can fully control all the cameras from the manufacturer. Our real- world experiments demonstrate the effectiveness of the discovered attacks and raise the alarms again for the IoT manufacturers.
Zhen Ling 0001, Kaizheng Liu, Yiling Xu, Yier Jin, Xinwen Fu
GLOBECOM3
2017 Optimal DASH-multicasting over LTE
abstract
Dynamic Adaptive Streaming over HTTP (DASH) is a fast growing video streaming platform which enables adaptive rate selection based on channel conditions. File Delivery over Unidirectional Transport (FLUTE) further enables multicasting of the DASH segments over LTE eMBMS systems. In this paper, an optimal DASH-multicasting solution is proposed to allow more DASH clients in an LTE network to receive better videos by optimizing the resource allocation, Forward Error Correction (FEC) code rate and modulation and coding scheme (MCS) of each multicasting group, which corresponds to a FLUTE session. Multiple FLUTE sessions are considered to deliver multiple videos and multiple video rates for enhancing the overall utility. We have applied the convex optimization method to find the optimal resource allocation in terms of utility for multiple FLUTE sessions. We also find the optimal FEC code rates to add redundancies to protect the video segments for each FLUTE session. Moreover, an efficient MCS selection is introduced to reduce the complexity of the algorithm. Simulation results, with realistic LTE parameters, are shown to prove the proposed scheme is optimal, with more DASH clients receiving better video representations within limited resources when compared to other existing algorithms.
Jounsup Park, Aliasghar Tarkhan, Jenq-Neng Hwang, Qiyue Li 0001, Yiling Xu, Wei Huang 0012
ICC5
2017 Compact scalable hash from deep learning features aggregation for content de-duplication
abstract
Unprecedented growth in media content generation, communication and consumption has taken over the vast majority of storage spaces in devices, network caches, and clouds. How to identify duplications from network caches is an important issue for fast and efficient content delivery network (CDN) communication and storage. In this work, we developed a novel hash scheme which is scalable and robust to typical CDN induced transcoding and manipulations. Scalable hash design is constructed in essentially two stages: images are first represented as 512 channels of thumbnail images from the deep learning VGG-16 networks, and then a Fisher Vector aggregation is performed on the features which offer scalability in both underlying Gaussian Mixture Model (GMM) PCA embedding and component posterior likelihood. Hash is generated by direct binarizing the Fisher Vector with component/dimensionality priority optimization. Simulation results have demonstrated that this is a very compact and accurate scheme for CDN content de-duplication.
Shan Feng, Zhu Li 0001, Yiling Xu, Jun Sun 0005
MMSP3
2017 Lossless point cloud geometry compression via binary tree partition and intra prediction
abstract
Characterized by geometry and photometry attributes, point cloud has become widely applied in the real-time presentation of various 3D objects and scenes. The development of even more precise capture devices and the increasing requirements for vivid rendering inevitably induce huge point capacity, thus making the point cloud compression a demanding issue. Considering the non-uniform sampling and time-variant geometry, appropriate structural representation for point cloud is important. In this paper, we propose a lossless geometry compression algorithm for 3D point cloud which serves as the basis of future adaptive improvement. We utilize the binary tree structure for effectively partitioning unorganized points into block structure. This hierarchical representation obtains roughly the same quantity level for each leaf node. Further analysis is conducted on an intra-geometry prediction via extended Travelling Salesman Problem (TSP), achieving an impressive performance in eliminating point-wise redundancy while preserving one single reference position for each block. The residual encoding is accomplished via a shallow neural network-based lossless compression algorithm, PAQ. Simulation results confirm the lossless compression of geometry from high quality capture, achieving approximately 3.5 times efficiency gain over the state of art algorithm implemented as MPEG Point Cloud Compression (PCC) reference software.
Wenjie Zhu 0004, Yiling Xu, Li Li 0040, Zhu Li 0001
MMSP2
2017 NDMP - An emerging MPEG standard for network distributed media processing
abstract
In this paper, we introduce a novel video codec and distribution standard developed by the Moving Picture Experts Group (MPEG), called network-distributed media processing (NDMP). Compared with existing video codec and distribution standards, which assume one encoder and one decoder, NDMP inserts a media processing unit, based on network entities, which implements media processing in a distributed manner, between the encoder and the decoder. NDMP has advantages in the following aspects:(1) it reduces the need of storage resources in the network. (2) it reduces the occupation of the bandwidth in backhaul network and (3) it reduces Round-Trip Time of remote media processing.
Yiling Xu, Jun Sun 0005, Jaeyeon Song, Kyungmo Park
VCIP2
2017 A Practical System Towards the Secure, Robust and Pervasive Mobile Workstyle
abstract
We develop an innovative PC2PC (personal computer to pervasive computing) system to enable the secure, robust and pervasive mobile workstyle. PC2PC server compresses the desktop screens of any virtualized system, and delivers the stream through any popular networks to PC2PC client remotely for stream decoding, rendering and end-user interaction (such as keyboard/mouse commands). We have implemented the overall system from the scratch, where the emerging screen content coding (SCC) extension of the High-Efficiency Video Coding (HEVC) is implemented to compress and stream the desktop screens in real-time, and three core asset channels (i.e., system, display, inputs, etc) are defined to enable systematic end-to-end communication. Compared with the commercial Red Hat SPICE virtual desktop infrastructure (VDI) scheme, our PC2PC could save the network bandwidth by a factor of 2, 7 and 4 respectively for typical video streaming, web browsing and stationary office applications at same visual quality. Meanwhile, we have also measured the delays in the system and presented the preliminary study on the user experience impact. A simple network estimation is applied to optimize the quality-bandwidth adaptation for both single user and multiuser scenarios to combat the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
VTC Spring4
2017 Security Vulnerabilities of Internet of Things: A Case Study of the Smart Plug System
abstract
With the rapid development of the Internet of Things, more and more small devices are connected into the Internet for monitoring and control purposes. One such type of devices, smart plugs, have been extensively deployed worldwide in millions of homes for home automation. These smart plugs, however, would pose serious security problems if their vulnerabilities were not carefully investigated. Indeed, we discovered that some popular smart home plugs have severe security vulnerabilities which could be fixed but unfortunately are left open. In this paper, we case study a smart plug system of a known brand by exploiting its communication protocols and successfully launching four attacks: 1) device scanning attack; 2) brute force attack; 3) spoofing attack; and 4) firmware attack. Our real-world experimental results show that we can obtain the authentication credentials from the users by performing these attacks. We also present guidelines for securing smart plugs.
Zhen Ling 0001, Junzhou Luo, Yiling Xu, Kui Wu 0001, Xinwen Fu
IEEE Internet Things J.3
2017 Interactive Screen Video Streaming-Based Pervasive Mobile Workstyle
abstract
In this paper, we develop an interactive screen video streaming-based system to enable the ubiquitous mobile workstyle, which is referred to as personal computer to pervasive computing (PC2PC). The desktop screens of virtualized systems are compressed in the PC2PC servers and delivered to remote end users for stream decoding, rendering, and interactions. We have implemented a system from the scratch, where the emerging screen content coding extension of high-efficiency video coding is implemented to compress and stream the desktop screens of the virtualized system in real time. Three core asset channels, system, display, and inputs, are defined to enable systematic end-to-end communication. Compared with Red Hat SPICE virtual desktop infrastructure scheme, the proposed PC2PC could save network bandwidth consumption by a factor of 2, 7, and 4, respectively, in terms of typical video streaming, web browsing, and stationary office applications at the same visual quality. Meanwhile, we have also measured the delays of the system and presented preliminary results on the user experience aspect. A simple network estimation is applied to optimize the quality bandwidth adaptation for both single user and multiuser scenarios to consider the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
IEEE Trans. Multim.4
2016 Light Reference Video Quality Estimation with Embedding of Compact Features
abstract
One key issue in the content delivery network service is efficient and effective video quality assessment. Yet the legacy method obtaining the results via detailed calculation based on the original sequence or subjective experiments which lacks efficiency. In this work, we developed a compact feature based solution that allows for on-demand quality assessment with limited light references. Compact feature subspace embedding and SVM classifiers are trained to derive Quality of Experience (QoE) labels from the combination of the original and corresponding reconstructed compact features. Overhead in computation complexity and communication cost is minimum. Simulation results demonstrated the robustness and universality of the proposed solution.
Wenjie Zhu 0004, Zhu Li 0001, Yiling Xu
ISM3
2016 A new AL-FEC coding scheme with limited feedback
abstract
For the next generation mobile video broadcasting, especially in-band solutions that serves the mobile devices, a limited feedback scheme via cellular channel polling is feasible to give accurate real-time information on the broadcast receivers' channel erasure rate, and decoding buffer status. In this work, we propose an AL-FEC coding degree scheme based on this feedback, to achieve a better decode efficiency and save the code redundancy. Simulation results demonstrate the effectiveness of this solution, and open up new opportunities in the next generation broadcasting system design.
Wei Huang 0012, Hao Chen 0036, Yiling Xu, Zhu Li 0001, Wenjun Zhang 0001
MMSP3
2016 Single-input-multiple-ouput transcoding for video streaming
abstract
In this work, a single input multiple output (SIMO) transcoding architecture is proposed. SIMO will benefit the mobile edge computing (such as HTTP Live Streaming requiring multiple copies of the video streams at different quality levels) without resorting to the legacy transcoding that video stream is completed decoded and encoded multiple times without exploring the compressed information. Leveraging the information encoded in the existing video streams, we could reduce the search candidates when transcoding the high quality bitstream to other versions with reduced quality level. As the first step, we have demonstrated the SIMO idea with bit rate shaping (i.e., bit rate transcoding) only scenario. It has shown more than 2x complexity reduction without quality loss using the common test conditions.
Hao Zhang 0032, Hao Chen 0036, Yiling Xu, Zhan Ma 0001
MMSP6
2015 2-D Index Map Coding for HEVC Screen Content Compression
abstract
This paper introduces a 2-D index map coding of the palette mode in screen content coding extension of the High-Efficiency Video Coding (HEVC SCC) standard to further improve the compression performance. In contrast to the current 1-D search using RUN to represent the length of matched string, we bring the block width and height to describe the arbitrary rectangle shape. We also use the block vector displacement to signal the matched block distance efficiently. By enlarging the search range from current coding tree unit (CTU) to a small neighbor CTU window (i.e., 3×5 CTUs), it provides the coding efficiency comparable to the case that full-frame intra block copy is used. It is more practical to use the local search window in real life considering the trade-off between the coding efficiency and implementation cost.
Yiling Xu, Wei Huang 0012, Fanyi Duanmu, Zhan Ma 0001
DCC1