Ruoke Yan

dblp:272/6466 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-4570-5402ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Feedforward Human-Centric Video Compression via 3D Gaussian Generation
abstract
In this paper, we propose a feed-forward framework (Fig. 1) for human video compression based on 3D generative reconstruction. Recent surveys [1], [2] highlight the need for more efficient and semantically aligned solutions. Our approach disentangles video content into complementary structural and motion layers: the structural layer encodes regularized texture from a single frame, while the motion layer leverages the SMPL-X prior to represent complex dynamics with a compact set of pose and shape parameters. A hierarchical coding scheme exploits the heterogeneity of these representations for improved efficiency. After decoding, a feed-forward 3D reconstruction pipeline with facial feature extraction is employed, in which a multimodal transformer and a Gaussian head synthesize parametric cues that are fused with motion signals for accurate animation and high-fidelity rendering. Experiments show over$1000 \times$compression while preserving structural and semantic fidelity. The method consistently outperforms strong baselines (especially at$0.04-0.1 \text{kbpp})$, with significant gains in rate-distortion, FVD, and perceptual quality, as well as robust generalization across identities and scenes.
Haocheng Tang, Ruoke Yan, Jiaqi Zhang 0007, Siwei Ma 0001
DCC2
2025 HGC-Avatar: Hierarchical Gaussian Compression for Streamable Dynamic 3D Avatars
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have enabled fast, photorealistic rendering of dynamic 3D scenes, showing strong potential in immersive communication. However, in digital human encoding and transmission, the compression methods based on general 3DGS representations are limited by the lack of human priors, resulting in suboptimal bitrate efficiency and reconstruction quality at the decoder side, which hinders their application in streamable 3D avatar systems. We propose HGC-Avatar, a novel Hierarchical Gaussian Compression framework designed for efficient transmission and high-quality rendering of dynamic avatars. Our method disentangles the Gaussian representation into a structural layer, which maps poses to Gaussians via a StyleUNet-based generator, and a motion layer, which leverages the SMPL-X model to represent temporal pose variations compactly and semantically. This hierarchical design supports layer-wise compression, progressive decoding, and controllable rendering from diverse pose inputs such as video sequences or text. Since people are most concerned with facial realism, we incorporate a facial attention mechanism during StyleUNet training to preserve identity and expression details under low-bitrate constraints. Experimental results demonstrate that HGC-Avatar provides a streamable solution for rapid 3D avatar rendering, while significantly outperforming prior methods in both visual quality and compression efficiency.
Haocheng Tang, Ruoke Yan, Xinhui Yin, Qi Zhang 0042, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia
ACM Multimedia2
2025 Prediction Enhancement for Point Cloud Attribute Compression Using Smoothing Filter
abstract
In recent years, 3D point cloud compression (PCC) has emerged as a prominent research area, attracting widespread attention from both academia and industry. As one of the PCC standards released by the moving picture expert group (MPEG), the geometry-based PCC (G-PCC) adopts two attribute lossy coding schemes, namely the prediction-based Lifting Transform and the region adaptive hierarchical transform (RAHT). Based on statistical analysis, it can be observed that the increase in predictive distance gradually weakens the attribute correlation between points, resulting in larger prediction errors. To address this issue, we propose a prediction enhancement method by using the smoothing filter to improve the attribute coding efficiency, which is both integrated into the Lifting Transform and RAHT. For the former, the neighbor point smoothing method based on the prediction order is proposed via a weighted average strategy. The proposed smoothing is only applied to points in the lower level of details (LoDs) by adjusting the distance-based predicted attribute values. For the latter, we design a neighbor node smoothing method after the inter depth up-samping (IDUS) prediction, where the sub-nodes in the same unit node are filtered for lower levels. Experimental results have demonstrated that compared with two latest MPEG G-PCC reference software TMC13-v23.0 and GeSTM-v3.0, our proposed enhanced prediction method exhibits superior Bjøntegaard delta bit rate (BDBR) gains with small increase in time complexity.
Qian Yin 0002, Ruoke Yan, Xinfeng Zhang 0001, Siwei Ma 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Joint Structure-Texture Scan-Order for Point Cloud Attribute Compression Using Affine Transformation
abstract
Existing geometry-based Point Cloud Compression (PCC) frameworks are typically designed to code the geometric coordinates first, followed by compressing the attributes (e.g., colors, reflectances) according to the order derived from geometric structures, such as the Morton codes. Although geometry-based reordering methods can eliminate the redundancy of attributes, the errors caused by dramatic variations of attributes in the non-smooth areas potentially limit the efficiency of the point cloud attribute coding. To tackle this challenge, a novel joint structure-texture scan-order and coding scheme is proposed, which aims to explore a better attribute coding order from the viewpoint of improving the geometry-attribute consistency. Specifically, we formulate the attribute reordering problem as a geometry-attribute alignment task, and utilize the affine-transform model to find the optimal correspondence between geometry and attribute information by minimizing attributed prediction residuals. Then, the Morton codes-based point cloud reordering is conducted on the transformed point cloud. Note that our residual-based mode decision scheme implicitly embodies that the proposed reordering method further incorporates attribute textures based on the geometric structure. However, the exhaustive search for the optimal transformation space introduces the extremely high complexity to the encoder. Therefore, we also propose a fast pruning algorithm to narrow the search space for the approximate solution as an alternative. Experimental results conducted on the various benchmark datasets have illustrated that our proposed method outperforms the MPEG standard G-PCC with gains of up to 2%, 9%, and 7% in luma, chroma, and reflectance, respectively.
Qian Yin 0002, Xinfeng Zhang 0001, Ruoke Yan, Yuhuai Zhang, Shanshe Wang, Siwei Ma 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Pose-Driven Compression for Dynamic 3D Human via Human Prior Models
abstract
To cost-effectively transmit high-quality dynamic 3D human images in immersive multimedia applications, efficient data compression is crucial. Unlike existing methods that focus on reducing signal-level reconstruction errors, we propose the first dynamic 3D human compression framework based on human priors. The layered coding architecture significantly enhances the perceptual quality while also supporting a variety of downstream tasks, including visual analysis and content editing. Specifically, a high-fidelity pose-driven Avatar is generated from the original frames as the basic structure layer to implicitly represent the human shape. Then, human movements between frames are parameterized via a commonly-used human prior model, i.e., the Skinned Multi-Person Linear Model (SMPL), to form the motion layer and drive the Avatar. Furthermore, the normals are also introduced as an enhancement layer to preserve fine-grained geometric details. Finally, the Avatar, SMPL parameters, and normal maps are efficiently compressed into layered semantic bitstreams. Extensive qualitative and quantitative experiments show that the proposed framework remarkably outperforms other state-of-the-art 3D codecs in terms of subjective quality with only a few bits. More notably, as the size or frame number of the 3D human sequence increases, the superiority of our framework in perceptual quality becomes more significant while saving more bitrates.
Ruoke Yan, Qian Yin 0002, Xinfeng Zhang 0001, Qi Zhang 0042, Gai Zhang, Siwei Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Model-Driven Compression for Digital Human Using Multi-Granularity Representations
abstract
With the popularity of the "metaverse", the applications of virtual digital humans are emerging in entertainment and communication fields, etc, where effective digital human compression schemes are urgently needed to process huge amounts of data. Most of the existing coding methods focus on reducing spatial redundancy in terms of the signal level but ignore the consideration of visual perception. In this paper, we propose a novel digital human model-driven compression framework by using multi-granularity representations. A coarse-grained layer based on the Skinned Multi-Person Linear model (SMPL) is introduced to extract general structures, while the normal images are used to represent fine-grained details via pose consistency-based predictions. Then, the SMPL parameters and normal images are encoded to achieve high-quality compression at extremely low bitrates. Experimental results demonstrate that the proposed method provides better coding performance with superior perceptual quality compared to the state-of-the-art 3D model compression methods.
Ruoke Yan, Qian Yin 0002, Xinfeng Zhang 0001, Siwei Ma 0001
ICME1
2020 Cascaded Detail-Aware Network for Unsupervised Monocular Depth Estimation
abstract
Existing unsupervised learning methods usually reformulate the depth estimation into the image reconstruction problem by training on stereo image pairs to circumvent the need of dense labeled ground truth depth information. Most of them are designed based on a simple encoder-decoder backbone architecture, which has limited expression for context information and suffers from the loss of depth details. In this paper, we propose a cascaded detail-aware network which contains a contextual network (CN) followed by consecutive spatial networks (SNs) to make an unsupervised coarse-to-fine prediction. CN aims to provide good initialized depth estimation results by introducing a multi-scale attention fusion module to enhance the ability of feature representation. Then, SN is progressively applied on the coarse depth map to produce refined depth outputs by exploiting abundant spatial details from input color image. Moreover, we design a robust loss function that further considers the penalty of photometric errors and the occlusion, and strengthens the recovery of spatial details for better depth estimation. Experimental results show that the proposed method achieves promising performance.
Xinchen Ye, Mingliang Zhang 0002, Xin Fan 0001, Rui Xu 0002, Juncheng Pu, Ruoke Yan
ICME6