Sicheng Li 0003

dblp:129/7659-3 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0001-2468-5863ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 GSCodec Studio: A Modular Framework for Gaussian Splat Compression
abstract
3D Gaussian Splatting and its extension to 4D dynamic scenes enable photorealistic, real-time rendering from real-world captures, positioning Gaussian Splats (GS) as a promising format for next-generation immersive media. However, their high storage requirements pose significant challenges for practical use in sharing, transmission, and storage. Despite various studies exploring GS compression from different perspectives, these efforts remain scattered across separate repositories, complicating benchmarking and the integration of best practices. To address this gap, we present GSCodec Studio, a unified and modular framework for GS reconstruction, compression, and rendering. The framework incorporates a diverse set of 3D/4D GS reconstruction methods and GS compression techniques as modular components, facilitating flexible combinations and comprehensive comparisons. By integrating best practices from community research and our own explorations, GSCodec Studio supports the development of compact representation and compression solutions for static and dynamic Gaussian Splats. Specifically, we present Static and Dynamic GSCodec: Static GSCodec achieves competitive 3D Gaussian Splat rate-distortion performance with low decoding complexity, while Dynamic GSCodec delivers advanced 4D Gaussian Splat compression performance. The code for our framework is publicly available at https://github.com/JasonLSC/GSCodec_Studio, to advance the research on Gaussian Splats compression.
Sicheng Li 0003, Chengzhen Wu, Hao Li 0069, Yiyi Liao, Lu Yu 0003
IEEE Trans. Circuits Syst. Video Technol.1
2026 Low-Rank Approximation for Efficient Compression of Gaussian Splatting Spherical Harmonics
abstract
3D Gaussian Splatting (3DGS) enables real-time, high-fidelity rendering but suffers from large model sizes, mainly due to the Spherical Harmonics (SH) coefficients used for view-dependent appearance modeling, whose parameter count scales quadratically(O(L2))with the degreeL. This paper presents the first systematic study that reveals and leverages the intrinsic low-rank structure of SH coefficients in 3DGS. Unlike prior approaches that truncate spectral energy, the proposed low-rank paradigm compactly preserves spectral information, achieving high visual quality with substantially reduced storage. Two complementary approaches are introduced. (1) SHAC-PCA (Principal Component Analysis) is a plug-and-play post-hoc compressor that retains principal spectral variance for high-fidelity compression. (2) SHAC-LST (Learned Subset Transformation) is a training-integrated approach that decomposes Alternating Current (AC) of SH coefficients into low-dimensional subset coefficients and a shared transformation matrix, offering superior compression and even quality improvements through regularization. Both methods effectively reduce SH coefficients storage complexity toO(L). Extensive experiments demonstrate that our approaches significantly reduce memory usage while maintaining or even improving rendering quality. The proposed techniques are highly versatile: SHAC-PCA can be applied to any pre-trained 3DGS model, while SHAC-LST supports end-to-end training or fine-tuning. Both methods are compatible with existing 3DGS compression pipelines, providing a practical and general solution for efficient compression of SH coefficients in 3DGS.
Sicheng Li 0003, Yiyi Liao, Lu Yu 0003
IEEE Trans. Circuits Syst. Video Technol.2
2025 GIFStream: 4D Gaussian-based Immersive Video with Feature Stream
abstract
Immersive video offers a 6-Dof-Free viewing experience, potentially playing a key role in future video technology. Recently, 4D Gaussian Splatting has gained attention as an effective approach for immersive video due to its high rendering efficiency and quality, though maintaining quality with manageable storage remains challenging. To address this, we introduce GIFStream, a novel 4D Gaussian representation using a canonical space and a deformation field enhanced with time-dependent feature streams. These feature streams enable complex motion modeling and allow efficient compression by leveraging their motion-awareness and temporal correspondence. Additionally, we incorporate both temporal and spatial compression networks for end-to-end compression. Experimental results show that GIFStream delivers high-quality immersive video at 30 Mbps, with real-time rendering and fast decoding on an RTX 4090.
Hao Li 0069, Sicheng Li 0003, Abudouaihati Batuer, Lu Yu 0003, Yiyi Liao
CVPR2
2025 Eliminating Geometric Representation Redundancy for 3D Gaussian Splat Coding
abstract
3D Gaussian Splatting (3DGS) enables photorealistic, real-time rendering, yet its native representation imposes significant storage and transmission overhead, hindering widespread deployment. Most existing approaches focus on reducing data volume or improving the efficiency of lossy coding. However, they overlook the high data entropy caused by inherent representation ambiguity, where multiple geometric attribute values can define the same geometry. To address this, we introduce a lightweight, plug-and-play preprocessing method that lowers raw data entropy by canonicalizing scale and quaternion attributes. Specifically, our method first resolves geometric representation ambiguity via a deterministic regularization rule that enforces a unique representation; reduces dimensionality by converting 4D quaternions to minimal 3D Rodrigues parameters; and addresses numerical redundancy by clamping perceptually insignificant scale values. Our method is a generic, plug-and-play preprocessing module, fully orthogonal to existing 3DGS compression schemes, and effectively boosts their coding efficiency. When combined with a baseline compression pipeline, it yields an average BD-Rate reduction of 15.81% compared to the same pipeline without our preprocessing.
Shanchuan Liu, Sicheng Li 0003, Yiyi Liao, Lu Yu 0003
VCIP3
2024 NeRFCodec: Neural Feature Compression Meets Neural Radiance Fields for Memory-Efficient Scene Representation
Sicheng Li 0003, Hao Li 0069, Yiyi Liao, Lu Yu 0003
CVPR1
2023 SteerNeRF: Accelerating NeRF Rendering via Smooth Viewpoint Trajectory
abstract
Neural Radiance Fields (NeRF) have demonstrated superior novel view synthesis performance but are slow at rendering. To speed up the volume rendering process, many acceleration methods have been proposed at the cost of large memory consumption. To push the frontier of the efficiency-memory trade-off, we explore a new perspective to accelerate NeRF rendering, leveraging a key fact that the view-point change is usually smooth and continuous in interactive viewpoint control. This allows us to leverage the information of preceding viewpoints to reduce the number of rendered pixels as well as the number of sampled points along the ray of the remaining pixels. In our pipeline, a low-resolution feature map is rendered first by volume rendering, then a lightweight 2D neural renderer is applied to generate the output image at target resolution leveraging the features of preceding and current frames. We show that the proposed method can achieve competitive rendering quality while reducing the rendering time with little memory overhead, enabling 30FPS at 1080P image resolution with a low memory footprint.
Sicheng Li 0003, Hao Li 0069, Yue Wang 0020, Yiyi Liao, Lu Yu 0003
CVPR1
2021 Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks have been widely used for skeleton-based action recognition due to their excellent modeling ability of non-Euclidean data. As the graph convolution is a local operation, it can only utilize the short-range joint dependencies and short-term trajectory but fails to directly model the distant joints relations and long-range temporal information that are vital to distinguishing various actions. To solve this problem, we present a multi-scale spatial graph convolution (MS-GC) module and a multi-scale temporal graph convolution (MT-GC) module to enrich the receptive field of the model in spatial and temporal dimensions. Concretely, the MS-GC and MT-GC modules decompose the corresponding local graph convolution into a set of sub-graph convolution, forming a hierarchical residual architecture. Without introducing additional parameters, the features will be processed with a series of sub-graph convolutions, and each node could complete multiple spatial and temporal aggregations with its neighborhoods. The final equivalent receptive field is accordingly enlarged, which is capable of capturing both short- and long-range dependencies in spatial and temporal domains. By coupling these two modules as a basic block, we further propose a multi-scale spatial temporal graph convolutional network (MST-GCN), which stacks multiple blocks to learn effective motion representations for action recognition. The proposed MST-GCN achieves remarkable performance on three challenging benchmark datasets, NTU RGB+D, NTU-120 RGB+D and Kinetics-Skeleton, for skeleton-based action recognition.
Sicheng Li 0003, Bing Yang 0004, Qinghan Li, Hong Liu 0008
AAAI2