VLDB 2026 Research / reviewers in the wild / expert
Yingzhan Xu
dblp:243/6667
· DBLP profile ↗
17ranked-venue papers
3as first author
16since 2021 · last 2026
0009-0002-8205-2316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 14 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pre-Trained Prior-Assisted Dynamic Point Cloud Compression via Feature PredictionabstractPoint cloud sequences play a pivotal role in immersive applications as they provide 6 degree of freedom experience into dynamic 3D environments. However, the dynamic point cloud sequences have an overwhelming amount of data, which poses challenges for storage and transmission. Therefore, effective point cloud sequence compression is crucial for immersive applications. In this paper, we propose a dynamic point cloud compression framework, which enhances the compression efficiency of point cloud sequences from the perspectives of representation and prediction. To improve the representation and reconstruction ability of a single frame, we propose a specialized pre-training mechanism for compression, which enables a well-designed autoencoder to acquire substantial prior knowledge. The pre-training mechanism is based on the proposed mask operation and is trained on a large amount of data of multiple categories. This enables the pre-trained encoder to identify and extract features with strong representational capabilities, and allows the pre-trained decoder to accurately reconstruct the complete point cloud. For precise prediction, we develop tailored methods for different frame types. For intra-frames (I frames), we introduce a dual-branch spatial intra-prediction network by leveraging coordinate and occupancy information from downsampled sparse point clouds to predict representative features. For inter-frames (P frames), we propose a multi-scale context-based spatio-temporal inter-prediction network. Specifically, for each scale, we utilize the proposed Transformer-style spatio-temporal modeling module to analyze and model the temporal characteristics among multiple reference frames, providing rich context information. Integrating three-scale contexts, we achieve comprehensive feature prediction, significantly improving accuracy and eliminating spatio-temporal redundancy. Experimental results show that our method outperforms existing compression methods in both I-frame compression and P-frame compression. Xinfeng Zhang 0001, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Cross-Component Residual Prediction for Geometry-Based Point Cloud CompressionabstractPoint cloud compression is pivotal for the success of immersive multimedia applications. For attribute compression in geometry-based point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Inspired by the significant impact of cross-component prediction in traditional image and video coding, we investigate and present our pioneering work on cross-component residual prediction for RAHT in G-PCC. The method builds on the core observation that cross-component correlations are observed locally in some regions in some sequences. Accordingly, the prediction is employed for last few layers of RAHT which capture local characteristics. We employ a simple linear model, that predicts chroma residues from reconstructed luma residue. The prediction coefficients are learnt on the fly from reconstructed residues of the neighbors. The method gives 1% luma coding gain and around 2-3% chroma coding gain with negligible increase in complexity. The method is adopted to the Geometric Solid Test Model (GeS-TM v7.0), a dedicated codec being developed for solid point clouds. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0136 |
ICASSP | 2 |
| 2025 | Rate-Distortion Optimized Chroma Quantization for Point Cloud CompressionabstractPoint cloud compression is pivotal for the success of immersive multimedia applications. For attribute compression in the geometry-based point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. G-PCC performs rate-distortion optimized quantization of luma and chroma residue where they are jointly quantized to zero if deemed to be R-D optimal. However, it is often beneficial to zero out the chroma residue while retaining the luma residue, since chroma exhibits less variations. To address this, we propose rate-distortion (R-D) optimized chroma quantization. Each chroma residue sample is decided to be quantized to zero based on the R-D cost. For accurate R-D cost evaluation, we propose a method to estimate the Lagrange multiplier λ on the fly and further scale it according to the RAHT layers, achieving sequence and layer-wise adaptivity. The method gives 1% effective luma coding gain with negligible increase in complexity. The method has been adopted to the next version of Geometric Solid Test Model (GeS-TM v8.0), a dedicated codec being developed for solid point clouds. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007 |
ICIP | 2 |
| 2025 | Weighted Average Prediction for Region Adaptive Hierarchical Transform in Solid Geometry Point Cloud CompressionabstractPoint cloud attribute compression is critical in real-time application scenarios of immersive multimedia data. Within the geometry based point cloud compression (G-PCC) standard, the region-adaptive hierarchical transform (RAHT) based methodology is employed to perform attribute compression with intra or inter prediction. However, current inter prediction or intra prediction cannot provide optimal performance for all RAHT nodes. To further improve prediction efficiency, this paper proposes a weighted average prediction for RAHT based attribute compression. The proposed strategy includes using the weighted average of inter prediction and intra prediction as the additional candidate prediction value, selecting different candidate prediction lists according to the octree depth to determine the final prediction value. We also propose using the prediction information of nodes in the parent layer or the same layer to adaptively derive the weights in average prediction for each node. Experiments on top of Ges-TM v4.0 demonstrate the superior performance of our method. The average coding gain on attribute BD-Rate is -3.2% (Luma), -4.5% (Chroma Cb) and -4.8% (Chroma Cr). The proposed scheme has been adopted to the Solid G-PCC standard. Yingzhan Xu, Kai Zhang 0007 |
ICIP | 1 |
| 2025 | Structure-Aware Generative Point Cloud Compression for Visual PerceptionabstractIn recent years, there has been a rapid growth in applications that rely on point clouds to represent the 3D world, driven by the increasing demand for immersive and other related scenarios. However, compressing the large and high-precision point cloud data efficiently while maintaining high perceptual quality for human vision remains a challenge. To solve the problem, we propose a new structure-aware generative point cloud compression framework for human vision. In the encoder, we focus on information that is more sensitive to the human vision and obtain this type of information from different scale. This allows us to capture structural importance information from global scale and local scale, which are more difficult to reconstruct. For the decoder, we introduce a progressive generative reconstruction approach that utilizes acquired information from the encoder to guide the generation of point cloud surfaces. Moreover, we propose a novel probability cloud-based discriminator. Instead of directly assessing the authenticity of the generated point clouds, our discriminator evaluates the probability distribution of the existence of points within the generated point cloud. This approach reduces the difficulty of discrimination while effectively improving the accuracy of the generator in generating probability distributions. According to the correct probability, we can obtain a high accuracy point cloud by pruning the points with low probability. Through comprehensive experiments, we demonstrate the effectiveness and superiority of our proposed framework in terms of encoding efficiency, high perceptual quality, and generation quality. Xinfeng Zhang 0001, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006, Qingming Huang |
IEEE Trans. Image Process. | 3 |
| 2024 | Dynamic point cloud compression with spatio-temporal transformer-style modelingabstractThe essence of dynamic point cloud compression lies in the effective modeling of temporal context information, which poses significant challenges owing to the unstructured and sparse characteristics of point clouds. Existing dynamic compression methods exhibit a limited capacity to capture and leverage inter-frame information. Consequently, in this paper, we propose a Dynamic Point Cloud Compression framework with Spatio-Temporal Transformer-style Modeling (DPCC-STTM) to compress point cloud sequences within a latent space. To effectively extract and fully utilize temporal context, we introduce a spatio-temporal transformer-style modeling module, which performs effective modeling of the rich temporal information based on the correlation of temporal content. Furthermore, we introduce a multi-scale temporal processing module that captures temporal correlations across short and long ranges of multi-frame point clouds. This module also fuses modeled temporal information to enhance the prediction accuracy of potential features for the current frame. Extensive experiments demonstrate the superiority of our proposed framework, validated through both objective evaluation and subjective perception. Xinfeng Zhang 0001, Xiaoqi Ma, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006 |
DCC | 4 |
| 2024 | Sample Domain Prediction and Transform Skip for Region Adaptive Hierarchical Transform in Geometric Point Cloud CompressionabstractPoint cloud compression is critical for the success of immersive multimedia applications. For attribute compression in geometric point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Although RAHT was initially introduced as a pure transform coding tool, recent advancements have introduced intra and inter prediction for RAHT. However, these methods perform prediction in transform domain which is sub-optimal since: ${i}$) fixed-point RAHT introduces distortion to the prediction signal and $i {i}$) transforming prediction signal leads to additional decoding complexity. To address this, we propose to perform prediction in sample domain, thereby retaining crisp prediction signal and alleviating decoder of unnecessary computations. Performing prediction in sample domain opens door to completely skip the transform stage at the decoder when all the residue of a block are quantized to zero, leading to further complexity reduction. The proposed methods achieve an average chroma coding gain of around $1 \%$ and reduces the overall decoding complexity by $3-5 \%$. The method is adopted to the next version of Geometric Solid Test Model (GeS-TM v5.0) and is being evaluated on G-PCC test model TMC13v25. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0136 |
ICIP | 3 |
| 2024 | Adaptive Downsampling and Spatial Upconversion for Point Cloud CompressionabstractAlthough ultra-high resolution point clouds have been acquired more easily, the huge amount of the data makes it more challenging to store and transmit, and also difficult to be applied in lightweight terminal. To address this challenge, we propose an Adaptive Downsampling and Spatial Upconversion framework for point cloud compression (ADSU). In the proposed adaptive downsampling, we introduce two key components, feature-aware augmented graph convolution (FAGC) and adaptive-sampling-based global graph aggregation module (ASGGA), to capture correlations between local and global features. For upsampling, we propose a high-frequency feature generation module (HFFG) to generate detailed information, which plays a crucial role in achieving precise reconstruction. Experimental results demonstrate that the combination of our proposed ADSU with popular point cloud compression methods can significantly improve compression performance. Xinfeng Zhang 0001, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 3 |
| 2024 | Improved Geometry Coding for Spinning LiDAR Point Cloud CompressionabstractPoint cloud compression has emerged as a hot research topic in recent years. Due to applications such as autonomous driving, LiDAR point cloud compression is an important research aspect of this field. Moving Picture Experts Group (MPEG) is developing a standard called Geometry-based Point Cloud Compression (G-PCC) to meet the compression requirements of point clouds from different collection devices including LiDAR. In current G-PCC, the prior information of spinning LiDAR is not fully utilized in octree geometry coding. In this paper, we address this issue and effectively account for the prior information of spinning LiDAR to improve the compression efficiency of octree geometry coding. Specifically, the angle information provided by capture laser scanner is utilized for Inferred Direct Coding Mode (IDCM) eligibility criterion and z coordinate compensation of the reconstructed point cloud. Experimental results demonstrate that the proposed method achieves 6.7% and 16.4% average coding gain under D1 and D2 quality metrics, respectively, with a negligible increase in complexity. The major part of the proposed method has been adopted in G-PCC. Yingzhan Xu, Bharath Vishwanath, Kai Zhang 0007, Li Zhang 0136 |
ISCAS | 2 |
| 2024 | Standardization Status of MPEG Geometry-Based Point Cloud Compression (G-PCC) Edition 2abstractPoint clouds, crucial for representing 3D objects and scenes, offer immersive and precise depictions of the real world. Despite their superiority, the substantial data volume challenges current multimedia ecosystems. To address this, the Moving Picture Expert Group (MPEG) initiated the point cloud compression project in 2017, leading to two branches: Video-based Point Cloud Compression (V-PCC) and Geometry-based Point Cloud Compression (G-PCC). The first edition of G-PCC was published in March 2023, and ongoing efforts over the past three years have advanced towards G-PCC Edition 2. This paper aims to present recent technical achievements and the current status of G-PCC standardization activities. Wei Zhang 0072, Fuzheng Yang 0001, Yingzhan Xu, Marius Preda |
PCS | 3 |
| 2023 | HM-PCGC: A Human-Machine Balanced Point Cloud Geometry Compression SchemeabstractPoint cloud compression has various purposes in different application scenarios, such as requiring visual fidelity in human vision tasks and pursuing semantic fidelity in machine vision tasks. To accommodate these diverse requirements, we propose a Human-Machine balanced point cloud geometry compression scheme (HM-PCGC) which considers the characteristics of various tasks. Our proposed scheme starts from a pre-trained, lightweight point cloud compression backbone and employs a Learned Semantic Mining module to aggregate multi-tasks features. By leveraging the aggregated features, HM-PCGC is able to retain the geometry and semantic properties of the point clouds during compression. To better balance between the signal distortion and semantic distortion, we integrated a multi-task learning mechanism during the training phase. Our approach is extensively evaluated and analyzed, and the results demonstrate its superiority over state-of-the-art traditional and deep learning based point cloud codecs for both signal reconstruction and machine vision tasks. Xiaoqi Ma, Yingzhan Xu, Xinfeng Zhang 0001, Lv Tang, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 2 |
| 2023 | Peer Upsampled Transform Domain Prediction for G-PCCabstractTo meet the growing demand for point cloud compression, MPEG is developing a point cloud compression standard called as G-PCC. In G-PCC, upsampled transform domain prediction (UTDP) is used to improve attribute coding performance. However, only the attributes in the previous level can be used to predict the attributes of transform sub-blocks in UTDP, which limits the efficiency of UTDP. To address this limitation, we propose a method called peer-UTDP to improve UTDP by using peer neighbors in this paper. With peer-UTDP, attributes of co-plane or co-line peer neighbors in the level same as that of the transform sub-block can be used as prediction in the upsampling process. Experimental results show that our method outperforms G-PCC with an average coding gain of -5.1%, -5.4%, -5.1% and -1.4% under C1 condition, and -5.1%, -5.6%, -5.6% and -1.7% under C2 condition for Y, Cb, Cr and reflectance, respectively. The proposed peer-UTDP has been adopted by G-PCC. Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006 |
ICME | 2 |
| 2023 | Temporal Filtering for Region Adaptive Hierarchical Transform in Geometric Point Cloud CompressionabstractDynamic point cloud compression is critical for the success of immersive multimedia applications and autonomous driving. For attribute compression in geometric point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Recently, inter-prediction was introduced for RAHT in G-PCC. In inter-RAHT, the transform coefficients of the top layers (low frequency coefficients) are predicted by a simple copy of the transform coefficients from the reference frame. Such a prediction assumes reference and current RAHT layers are perfectly correlated, which is barely true. To address this, the paper introduces a temporal filtering mechanism for RAHT. Specifically, we scale the reference layer by a filtering coefficient. Different filtering coefficients are designed for different RAHT layers, thereby emulating a low-pass filter. Experiments show an average 2% bit-rate savings over G-PCC test model tmc13-v22 with negligible increase in complexity. The method has been adopted to the second version of G-PCC. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0136 |
VCIP | 2 |
| 2022 | Bi-Directional Inter-Prediction For Geometry-Based Point Cloud CompressionabstractEfficient 3D point cloud compression plays critical role in immersive multimedia presentation and autonomous driving. The temporal redundancy between consecutive point cloud frames is obvious, but the exploration of inter-prediction is insufficient in the geometry-based point cloud compression (G-PCC) framework. In inter exploration model (Inter-EM) of G-PCC, the reference information can only come from one previous frame and a consistent quantization parameter (QP) value is used for all frames. To perform a more efficient inter-prediction for geometry and attribute, a bidirectional inter-prediction (bi-prediction) scheme is proposed for G-PCC based on Inter-EM. With the bi-prediction scheme, the reference information can come from two reference frames. For attribute compression, neighboring search is started from one search center derived by a Morton code distance and a two-threshold method is applied to constrain inter-prediction. Meanwhile, the dependencies between frames are designed according to a hierarchical group of frames (GOF) structure with corresponding hierarchical QP values. Experimental results demonstrate that our method outperforms G-PCC and Inter-EM for both lossless and lossy compression. More specifically, the average coding gain over G-PCC is 14.1% on attribute and 11.8% on geometry under lossy condition. The designs on attribute compression have been adopted to the latest Inter-EM. Yingzhan Xu, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 1 |
| 2021 | Visual Quality Optimization for View-Dependent Point Cloud CompressionabstractThe video-based point cloud compression (V-PCC) is the state-of-the-art dynamic point cloud compression technique. V-PCC projects the 3D point cloud data patch by patch to its bounding box and organizes projected patches into a video frame, making full use of the well-developed video coding tools. Despite its high efficiency, cracks easily exist in the reconstructed point cloud in various viewing angles, which seriously degrades the visual quality. In this paper, we propose an efficient method to improve the visual quality of dynamic point cloud, especially for the main view from the content provider. The relationship between patches and views is exploited, and an algorithm intelligently reserving points that may be discarded in V-PCC is proposed. According to our subjective and perceptual objective evaluation experiments, compared with V-PCC, the overall visual quality of the reconstructed point could is evidently improved. In particular, cracks are mended with our proposed method. The Bjontegaard delta bit-rate reduction of up to 3.1% is achieved with respect to Point Cloud Quality Metric (PCQM), which partially verifies the improvement of subjective quality when adopting the proposed method. Danying Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001 |
ISCAS | 3 |
| 2021 | Point-Voting based Point Cloud Geometry CompressionabstractThe Geometry-based Point Cloud Compression (G-PCC) proposed by the Moving Picture Experts Group (MPEG) is the state-of-art point cloud compression algorithm. It provides an efficient lossy geometry compression technique called triangle soup (Trisoup) for static point clouds. Based on the pruned octree structure, Trisoup provides a local surface model consisting of multiple triangles and compresses vertices of the triangles instead of directly compressing the positions of the original points. Accordingly, we propose a point-voting based method to improve the triangle-construction within each leaf node. This new method leverages the node-based points distribution for more precise vertices determination, which better fits the local surface. Experimental results demonstrate the effectiveness of our point-voting based method for both objective and subjective quality evaluation. Chaofei Wang, Wenjie Zhu 0004, Yingzhan Xu, Yiling Xu, Le Yang 0001 |
MMSP | 3 |
| 2019 | Dynamic Point Cloud Geometry Compression via Patch-wise Polynomial FittingabstractWith the boosting requirements of realistic 3D modeling for immersive applications, advent of the newly-developed 3D point cloud has attracted great attention. Frankly, immersive experience using high data volume affirms the importance of efficient compression. Inspired by the video-based point cloud compression (V-PCC), we propose a novel point cloud compression algorithm based on polynomial fitting of proper patches. Moreover, the original point cloud is segmented into various patches. We generated corresponding depth maps via projection of all the patches by focusing on geometry information. Instead of directly compressing the absolute values, we utilized proper polynomial functions to fit in each patch to obtain the differences. Finally, it is satisfying to note that the fitting function effectively represents the patch-wise geometry information. Moreover, new depth maps are obtained with extremely small and stable values, which are more suitable for video-based compression. Different patch-wise fitting parameters are preserved and coded using lossless compression through the open source PAQ project. The proposed approach achieves a noticeable improvement in the compression efficiency while maintaining point cloud quality. Yingzhan Xu, Wenjie Zhu 0004, Yiling Xu, Zhu Li 0001 |
ICASSP | 1 |