Yiting Shao

dblp:207/7680 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-9625-0124ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Lossless Dynamic Point Cloud Geometry Compression via Rate-Distortion Optimized Motion Estimation
abstract
Dynamic point clouds are valid representations of three-dimensional moving entities in diverse application scenarios. The comprehensive redundancy in uncompressed point clouds necessitates efficient compression methods. Motion estimation (ME) plays a crucial role in eliminating the temporal redundancy of point cloud sequences. However, existing ME methods suffer from inaccurate geometry compensation distortion measures and imbalanced rate-distortion modeling, significantly impacting the coding performance. To address these challenges, we propose a rate-distortion (R-D) optimized ME scheme for dynamic point cloud geometry compression. First, a point cloud is hierarchically decomposed into a set of macroblocks and blocks via the octree, and variable block-size ME is introduced to capture complex local motions in point clouds. Second, we propose a block-matching motion search method that integrates a new geometry compensation distortion measure, combining voxel-wise Hamming distance and point-wise Euclidean distance as a joint criterion to improve the accuracy of the ME. To mitigate complexity, local motion features are extracted to efficiently accelerate the geometry distortion measure. Third, a rate-distortion model with an adaptive Lagrange multiplier is designed to facilitate better selection of inter predictors in the motion decision stage, thereby achieving efficient coding performance. Experimental results demonstrate the improvements of our proposed scheme over rival platforms in coding performance and computational efficiency. Additional results also validate the effectiveness of key modules within the proposed R-D optimized ME framework.
Qi Zhang 0029, Yiting Shao, Lixuan Meng, Shan Liu 0001, Ge Li 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 Rate-Distortion Optimized Motion Estimation for Dynamic Point Cloud Geometry Compression
abstract
Dynamic point clouds serve as crucial representations of three-dimensional moving entities across diverse applications. The substantial amount of data in point clouds necessitates the development of efficient compression techniques. Motion estimation (ME) plays a crucial role in eliminating the temporal redundancy of point cloud sequences. However, prevailing ME methods suffer from the inaccurate geometry distortion measure and the imbalanced rate-distortion modeling, significantly impacting the coding performance. To address these challenges, we propose a rate-distortion (R-D) optimized ME scheme for dynamic point cloud geometry compression.
Qi Zhang 0029, Yiting Shao, Lixuan Meng, Hailong Jiao, Shan Liu 0001, Ge Li 0002
DCC2
2024 3D Point Cloud Attribute Compression Using Diffusion-Based Texture-Aware Intra Prediction
abstract
There is an urgent need from various multimedia applications to efficiently compress point clouds. The Moving Picture Experts Group has released a standard platform called geometry-based point cloud compression (G-PCC). However, itsk-nearest neighbor (k-NN) based attribute prediction has limited efficiency for point clouds with rich texture and directional information. To overcome this problem, we propose a texture-aware attribute predictive coding framework in a point cloud diffusion model. In our work, attribute intra prediction is solved as a diffusion-based interpolation problem, and a general attribute predictor is developed. It is theoretically proven that G-PCCk-NN based predictor is a degraded case of the proposed diffusion-based solution. First, a point cloud is represented as two levels of details with seeds as the inpainting mask and non-seed points to be predicted. Second, we design point cloud partial difference operators to perform energy-minimizing attribute inpainting from seeds to unknowns. Smooth attribute interpolation can be achieved via an iterative diffusion process, and an adaptive early termination is proposed to reduce complexity. Third, we propose a structure-adaptive attribute predictive coding scheme, where edge-enhancing anisotropic diffusion is employed to perform texture-aware attribute prediction. Finally, attributes of seeds are beforehand encoded and prediction residuals of left points are progressively encoded into bitstream. Experiments show the proposed scheme surpasses the state-of-the-art by an average of 14.14%, 17.52%, and 17.87% BD-BR gains on the coding of Y, U, and V components, respectively. Subjective results on attribute reconstruction quality also verify the advantage of our scheme.
Yiting Shao, Wei Gao 0003, Shan Liu 0001, Ge Li 0002
IEEE Trans. Circuits Syst. Video Technol.1
2023 PDE-based Progressive Prediction Framework for Attribute Compression of 3D Point Clouds
abstract
In recent years, the diffusion-based image compression scheme has achieved significant success, which inspires us to use diffusion theory to employ the diffusion model for point cloud attribute compression. However, the relevant existing methods cannot be used to deal with our task due to the irregular structure of point clouds. To handle this, we propose the partial differential equation (PDE) based progressive prediction framework for attribute compression of 3D point clouds. Firstly, we propose a PDE-based prediction module, which performs prediction by optimizing attribute gradients, allowing the geometric distribution of adjacent areas to be fully utilized and explaining the weighting method for prediction. Besides, we propose a low-complexity method for calculating partial derivative operations on point clouds to address the uncertainty of neighbor occupancy in three-dimensional space. In the proposed prediction framework, we design a two-layer level of detail (LOD) structure, where the attribute information in the high level is used for interpolating the low level by edge-enhancing anisotropic diffusion (EED) to infer local features from the high-level information. After the diffusion-based interpolation, we design a texture-wise prediction method making use of interpolated values and texture information. Experiment results show that our proposed framework achieves an average of 12.00% BD-rate reduction and 1.59% bitrate saving compared with Predlift (PLT) under attribute near-lossless and attribute lossless conditions, respectively. Furthermore, additional experiments demonstrate our proposed scheme has better texture preservation and subjective quality.
Yiting Shao, Shan Liu 0001, Thomas H. Li, Ge Li 0002
ACM Multimedia2
2023 Nonrigid Registration-Based Progressive Motion Compensation for Point Cloud Geometry Compression
abstract
There is a critical requirement for efficiently compressing point cloud geometries representing three-dimensional (3D) moving objects in various applications. The Moving Picture Experts Group 3D Graphics coding group (MPEG 3DG) set up an inter-exploration model for geometry-based point cloud compression (G-PCC interEM). However, the block-matching motion compensation scheme with a translational motion model has limited ability to handle dense point clouds with complex local motions. To overcome this problem, we propose a progressive non-rigid motion compensation framework for point cloud geometry compression, where the point cloud registration technique is introduced and tailored with our designed rate-distortion cost. In the coarse-grained stage, a point cloud is represented as deformable point patches, and the patch-wise non-rigid motion estimation task is formulated as a registration-based optimization problem that can be efficiently solved by the majorization-minimization method. In the fine-grained stage, we propose a block-based motion refinement to enhance the estimated motion field in the local region, followed by a multi-hypothesis motion compensation scheme enabling smooth reference reconstruction with patch-wise deformation and block-wise refined motions. Experiments demonstrate our proposed scheme outperforms several competitive platforms in terms of both coding performance and compensation quality. Compared with G-PCC interEM, our proposed framework achieves significant bitrate savings, i.e., 4.71% (32 frames) and 4.22% (200 frames), for point cloud lossless geometry compression.
Yiting Shao, Ge Li 0002, Qi Zhang 0029, Wei Gao 0003, Shan Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 A Fast Motion Estimation Method With Hamming Distance for LiDAR Point Cloud Compression
abstract
With more three-dimensional space information, Light detection and ranging (LiDAR) point clouds, which are promising to play more roles in the future, have an urgent need to be efficiently compressed. There are lots of compression methods based on spatial correlations, whereas few studies consider exploiting temporal correlations. In this paper, we propose a different perspective for the motion estimation. In most previous works, geometric distance between matching points was used as the criterion, which has an expensive computational cost and is not accurate. We first propose the Hamming distance between the octree's nodes, instead of the geometric distance between per point which is a more direct criterion. We have implemented our method in the MPEG (Moving Picture Expert Group) Geometry-based PCC (Point Cloud Compression) inter-exploration (G-PCC Inter-EM). Experimental results show our method can provide the average 3.5 % bitrate savings and 92.5 % encoding speed increase in lossless geometric coding, compared to the G-PCC Inter-EM.
Yuhao An, Yiting Shao, Ge Li 0002, Wei Gao 0003, Shan Liu 0001
VCIP2
2021 OralViewer: 3D Demonstration of Dental Surgeries for Patient Education with Oral Cavity Reconstruction from a 2D Panoramic X-ray
abstract
Patient’s understanding on forthcoming dental surgeries is required by patient-centered care and helps reduce anxiety. Due to the complexity of dental surgeries and the patient-dentist expertise gap, conventional techniques of patient education are usually not effective for explaining surgical steps. In this paper, we present OralViewer—the first interactive application that enables dentist’s demonstration of dental surgeries in 3D to promote patients’ understanding. OralViewer takes a single 2D panoramic dental X-ray to reconstruct patient-specific 3D teeth structures, which are then assembled with registered gum and jaw bone models for complete oral cavity modeling. During the demonstration, OralViewer enables dentists to show surgery steps with virtual dental instruments that can animate effects on a 3D model in real-time. A technical evaluation shows that our deep learning model achieves a mean Intersection over Union (IoU) of 0.771 for 3D teeth reconstruction. A patient study with 12 participants shows OralViewer can improve patients’ understanding of surgeries. A preliminary expert study with 3 board-certified dentists further verifies the clinical validity of our system.
Yuan Liang 0001, Liang Qiu 0001, Tiancheng Lu, Zhujun Fang, Dezhan Tu, Jiawei Yang 0002, Yiting Shao, Kun Wang 0005, Xiang 'Anthony' Chen, Lei He 0001
IUI7
2021 Exploring Forensic Dental Identification with Deep Learning
abstract
Dental forensic identification targets to identify persons with dental traces.The task is vital for the investigation of criminal scenes and mass disasters because of the resistance of dental structures and the wide-existence of dental imaging. However, no widely accepted automated solution is available for this labour-costly task. In this work, we pioneer to study deep learning for dental forensic identification based on panoramic radiographs. We construct a comprehensive benchmark with various dental variations that can adequately reflect the difficulties of the task. By considering the task's unique challenges, we propose FoID, a deep learning method featured by: (\textit{i}) clinical-inspired attention localization, (\textit{ii}) domain-specific augmentations that enable instance discriminative learning, and (\textit{iii}) transformer-based self-attention mechanism that dynamically reasons the relative importance of attentions. We show that FoID can outperform traditional approaches by at least \textbf{22.98\%} in terms of Rank-1 accuracy, and outperform strong CNN baselines by at least \textbf{10.50\%} in terms of mean Average Precision (mAP). Moreover, extensive ablation studies verify the effectiveness of each building blocks of FoID. Our work can be a first step towards the automated system for forensic identification among large-scale multi-site databases. Also, the proposed techniques, \textit{e.g.}, self-attention mechanism, can also be meaningful for other identification tasks, \textit{e.g.}, pedestrian re-identification.Related data and codes can be found at \href{https://github.com/liangyuandg/FoID}{https://github.com/liangyuandg/FoID}.
Yuan Liang 0001, Weikun Han, Liang Qiu 0001, Yiting Shao, Kun Wang 0005, Lei He 0001
NeurIPS5
2021 Layer-Wise Geometry Aggregation Framework for Lossless LiDAR Point Cloud Compression
abstract
Point cloud compression is critical to deploy 3D applications like autonomous driving. However, LiDAR point clouds contain many disconnected regions, where redundant bits for unoccupied 3D space and weak correlations between points make it a troublesome problem to achieve efficient compression. This paper aims to aggregate LiDAR point clouds to get compact representations with full consideration of the point distribution characteristics. Specifically, we propose a novel Layer-wise Geometry Aggregation (LGA) framework for LiDAR point cloud lossless geometry compression, which adaptively partitions point clouds into three layers based on the content properties, including a ground layer, an object layer, and a noise layer. The aggregation algorithms are delicately designed for each layer. Firstly, the ground layer is fitted to a Gaussian Mixture Model, which can uniformly represent ground points using much fewer model parameters than adopting the original 3D coordinates. Then, the object layer is tightly packed to reduce the space between objects effectively, and a dense layout for points can benefit compression efficiency. Finally, in the noise layer, the difference between neighbor points is reduced by reordering using Morton Code, and the reduced residuals can help saving bit consumption. Experimental results demonstrate that the proposed LGA significantly outperforms competitive methods without prior knowledge by 12.05~23.37% compression ratio gains. Furthermore, the enhanced LGA with prior knowledge shows consistent performance gains than the latest reference software. Additional results also validate the robustness and stability of our proposed scheme with acceptable time complexity.
Yiting Shao, Wei Gao 0003, Haiqiang Wang, Thomas Li
IEEE Trans. Circuits Syst. Video Technol.2
2020 Point Cloud Attribute Compression via Successive Subspace Graph Transform
abstract
Inspired by the recently proposed successive subspace learning (SSL) principles, we develop a successive subspace graph transform (SSGT) to address point cloud attribute compression in this work. The octree geometry structure is utilized to partition the point cloud, where every node of the octree represents a point cloud subspace with a certain spatial size. We design a weighted graph with self-loop to describe the subspace and define a graph Fourier transform based on the normalized graph Laplacian. The transforms are applied to large point clouds from the leaf nodes to the root node of the octree recursively, while the represented subspace is expanded from the smallest one to the whole point cloud successively. It is shown by experimental results that the proposed SSGT method offers better R-D performances than the previous Region Adaptive Haar Transform (RAHT) method.
Yueru Chen, Yiting Shao, Jing Wang 0115, Ge Li 0002, C.-C. Jay Kuo
VCIP2
2020 A point cloud compression framework via spherical projection
abstract
In this paper, we propose a sphere-projection-based framework for point cloud geometry and attribute lossless and lossy coding. The original point cloud is adaptively divided into blocks, and then we create a fitting sphere in each block for modeling the local geometry structure of the point cloud. Sphere coordination transform and spherical projection scheme are introduced to transfer a 3D point cloud to a set of the range images. A novel compact representation of generated range images based on Morton codes is proposed to separate the range images into occupancy images and attributes vectors for further compression. Experimental results demonstrate that for the LiDAR point clouds datasets in lossless compression, the proposed method offers better performance than geometry-based point cloud compression (G-PCC). For the object point clouds datasets in lossy compression, the proposed method has better rate-distortion (R-D) performance than Draco.
Yingshen He, Ge Li 0002, Yiting Shao, Jing Wang 0115, Yueru Chen, Shan Liu 0001
VCIP3
2020 Fast Recolor Prediction Scheme in Point Cloud Attribute Compression
abstract
Due to the emerging requirement of point cloud applications, efficient point cloud compression methods are in high demand for compact point cloud representation in limited bandwidth transmission. The compression standard GPCC (Geometry-based Point Cloud Compression) is led by the MPEG (Moving Picture Expert Group) in respond to industrial requirements. KNN (K-Nearest Neighbors) search based prediction method is adopted for point cloud attribute compression in current G-PCC, which only exploits Euclidean distance-based geometric relationship without fully consideration of underlying geometric distribution. In this paper, we propose a novel prediction scheme based on fast recolor technique for attribute lossless and near-lossless compression. Our method has been implemented upon G-PCC reference software of the latest version. Experimental results show that our method can take advantage of the correlation between the attributes of neighbors, which leads to better rate-distortion (R-D) performance than G-PCC anchor on point cloud dataset with negligible encode and decode time increase under the common test conditions.
Ge Li 0002, Qi Zhang 0029, Yiting Shao, Jing Wang 0115, Shan Liu 0001
VCIP4
2019 Enhanced Intra Prediction Scheme in Point Cloud Attribute Compression
abstract
3D point cloud compression (PCC) has been an attractive field with increasing applications in recent years. Moving Picture Experts Group (MPEG) is building an open standard for point cloud compression, consisting of two solutions, video-based point cloud compression (V-PCC) and geometry-based point cloud compression (G-PCC). As an essential process in G-PCC, K-nearest neighbors (KNN) algorithm is adopted to perform intra prediction, which only considering distance-based local similarity but neglecting the overall geometric distribution of the neighbor set. In this paper, we propose an enhanced intra prediction scheme based on point-cloud geometric distribution. The centroid-based criterion is introduced to measure the uniformity of spatial distribution of predictive reference points. Our scheme is implemented in G-PCC reference software. Experimental results demonstrate that our proposed method can optimize the selection of predictors, which leads to better rate-distortion (R-D) performance than the G-PCC anchor on point cloud datasets under the common test conditions (CTC).
Honglian Wei, Yiting Shao, Jing Wang 0115, Shan Liu 0001, Ge Li 0002
VCIP2
2018 Hybrid Point Cloud Attribute Compression Using Slice-based Layered Structure and Block-based Intra Prediction
abstract
Point cloud compression is a key enabler for the emerging applications of immersive visual communication, autonomous driving and smart cities, etc. In this paper, we propose a hybrid point cloud attribute compression scheme built on an original layered data structure. First, a slice partition scheme and a geometry-adaptive k-dimensional tree (k-d tree) method are devised to generate layer structures. Second, we introduce an efficient block-based intra prediction scheme containing to exploit spatial correlations among adjacent points. Third, an adaptive transform scheme based on Graph Fourier Transform (GFT) is Lagrangian optimized to achieve better transform efficiency. The Lagrange multiplier is off-line derived based on the statistics of attribute coding. Last but not least, multiple scan modes are dedicated to improve coding efficiency for entropy coding. Experimental results demonstrate that our method performs better than the state-of-the-art region-adaptive hierarchical transform (RAHT) system, and on average a 37.21% BD-rate gain is achieved. Comparing with the test model for category 1 (TMC1) anchors, which were recently published by MPEG-3DG group on 121st MPEG meeting, a 8.81% BD-rate gain is obtained.
Yiting Shao, Qi Zhang 0029, Ge Li 0002, Zhu Li 0001, Li Li 0040
ACM Multimedia1
2018 Point Clouds Attribute Compression Using Data-Adaptive Intra prediction
abstract
In recent years, 3D sensing and capture technologies have made constant progress, leading to point clouds with higher resolution and fidelity. Since most applications demand compact storage and fast transmission, the issue of how to compress point clouds efficiently becomes an intractable problem. While previous GFT-based solutions use the transform tool to decorrelate attributes directly, ignoring the overall attribute's data spatial redundancy, Graph Fourier Transform (GFT) has shown good performance on point cloud attribute compression. So, motivated by coding tools in traditional image and video coding, we propose a block-based data-adaptive intra prediction tool before graph transform processing to further reduce the redundancy. We adopt uniform quantizing and context-based arithmetic coding to get the final bitstream. Experimental results on different datasets demonstrate that our method improves the compression efficiency of other GFT-based schemes and has much better BD-BR performance than the state-of-the-art Region-Adaptive Hierarchical Transform (RAHT) approach on most specified point cloud contents.
Qi Zhang 0029, Yiting Shao, Ge Li 0002
VCIP2
2017 Attribute compression of 3D point clouds using Laplacian sparsity optimized graph transform
abstract
3D sensing and content capturing have made significant progress in recent years and the MPEG standardization organization is launching a new project on immersive media with point cloud compression (PCC) as one key corner stone. In this work, we introduce a new binary tree based point cloud partition and explore the graph signal processing tools, especially the graph transform with optimized Laplacian sparsity, to achieve better energy compaction and compression efficiency. The resulting rate-distortion operating points are convex-hull optimized over the existing Lagrangian solutions. Simulation results on the latest high quality point cloud content from the MPEG PCC demonstrate the transform efficiency and rate-distortion (R-D) optimal potential of the proposed solutions.
Yiting Shao, Zhaobin Zhang, Zhu Li 0001, Kui Fan, Ge Li 0002
VCIP1