Zhehao Shen

dblp:344/4748 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance
abstract
Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In this paper, we present RePerformer, a novel Gaussian-based representation that unifies playback and re-performance for high-fidelity human-centric volumetric videos. Specifically, we hierarchically disentangle the dynamic scenes into motion Gaussians and appearance Gaussians which are associated in the canonical space. We further employ a Morton-based parameterization to efficiently encode the appearance Gaussians into 2D position and attribute maps. For enhanced generalization, we adopt 2D CNNs to map position maps to attribute maps, which can be assembled into appearance Gaussians for high-fidelity rendering of the dynamic scenes. For re-performance, we develop a semantic-aware alignment module and apply deformation transfer on motion Gaussians, enabling photo-real rendering under novel motions. Extensive experiments validate the robustness and effectiveness of RePerformer, setting a new benchmark for playback-then-reperformance paradigm in human-centric volumetric videos. Project page: https://moqiyinlun.github.io/Reperformer/.
Yuheng Jiang, Zhehao Shen, Zhuo Su 0006, Yingliang Zhang, Marc Habermann, Lan Xu 0003
CVPR2
2025 BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video
Yize Wu, Zhehao Shen, Yuheng Jiang, Yingliang Zhang, Qiang Hu 0003, Jingyi Yu 0001, Lan Xu 0003
ACM Multimedia3
2025 Dynamic Gaussian Streams for Volumetric Video via Codebook-Based Quantization
abstract
Volumetric video is rapidly emerging as a next-generation media format for immersive VR/AR applications, offering free-viewpoint rendering and unprecedented realism. While 3D Gaussian Splatting (3DGS) has recently demonstrated impressive rendering quality and real-time performance, existing dynamic extensions often struggle with long sequences due to the lack of efficient and codec-friendly compression schemes. In particular, current methods are not yet VR/AR-ready, as they fail to balance high-fidelity rendering, compact storage, and real-time decoding across heterogeneous platforms. To address these challenges, we propose Dynamic Gaussian Streams, a compact, video-compatible representation for real-time immersive playback. Given multi-view video inputs, we leverage the DualGS framework to reconstruct a temporally coherent 4D Gaussian sequence, introducing key modifications that directly optimize the 3D positions of dense skin Gaussians to improve compressibility and rendering quality. Each frame is converted into structured 2D maps, where key appearance attributes are compressed using per-channel codebooks with uint8 index maps. Hierarchical index reordering and Morton layout optimize spatial and temporal locality, ensuring compatibility with standard H.264 codecs. For spatial attributes like position, a lossless uint16 quantization preserves sub-pixel accuracy. Our system strikes a strong balance between compression and visual fidelity, enabling real-time decoding and immersive rendering on platforms including mobile devices, and XR headsets such as Apple Vision Pro.
Zhehao Shen, Yiwen Cai, Yuanji Lu, Yize Wu, Meihan Zheng, Yingliang Zhang, Lan Xu 0003
MMSP1
2025 Topology-Aware Optimization of Gaussian Primitives for Human-Centric Volumetric Videos
abstract
Volumetric video is emerging as a key medium for digitizing the dynamic physical world, creating the virtual environments with six degrees of freedom to deliver immersive user experiences. However, robustly modeling general dynamic scenes, especially those involving topological changes while maintaining long-term tracking remains a fundamental challenge. In this paper, we present TaoGS, a novel topology-aware dynamic Gaussian representation that disentangles motion and appearance to support, both, long-range tracking and topological adaptation. We represent scene motion with a sparse set of motion Gaussians, which are continuously updated by a spatio-temporal tracker and photometric cues that detect structural variations across frames. To capture fine-grained texture, each motion Gaussian anchors and dynamically activates a set of local appearance Gaussians, which are non-rigidly warped to the current frame to provide strong initialization and significantly reduce training time. This activation mechanism enables efficient modeling of detailed textures and maintains temporal coherence, allowing high-fidelity rendering even under challenging scenarios such as changing clothes. To enable seamless integration into codec-based volumetric formats, we introduce a global Gaussian Lookup Table that records the lifespan of each Gaussian and organizes attributes into a lifespan-aware 2D layout. This structure aligns naturally with standard video codecs and supports up to 40× compression. TaoGS provides a unified, adaptive solution for scalable volumetric video under topological variation, capturing moments where “elegance in motion” and “Power in Stillness”— delivering immersive experiences that harmonize with the physical world. Project page: https://guochch.github.io/TaoGS/.
Yuheng Jiang, Yize Wu, Shengkun Zhu, Zhehao Shen, Yingliang Zhang, Shaohui Jiao, Zhuo Su 0006, Lan Xu 0003, Marc Habermann, Christian Theobalt
SIGGRAPH Asia6
2025 EfficientPEAL: Efficient prior-embedded attention learning for partially overlapping point cloud registration
abstract
Learning discriminative point-wise features is critical for partially overlapping point cloud registration. In recent years, the integration of a Transformer into point cloud feature representation has demonstrated remarkable success, which typically involves a self-attention module to learn intra-point-cloud features, followed by a cross-attention module for feature exchange between input point clouds. Transformer models mainly benefit from the use of self-attention to capture the global correlations in feature space. However, the global correlations involved in self-attention may not only result in a significant amount of redundant computational overhead but also introduce feature ambiguities, especially in low-overlap scenarios. This is because overlapping regions of point clouds typically do not span a wide range but are rather concentrated around a localized area. Therefore, the correlations with an extensive range of non-overlapping points are ineffective and may degrade the discriminability of features. To address this issue, we present a E fficient P rior- E mbedded A ttention L earning model ( E fficientPEAL). By incorporating overlap prior to the learning process, the point clouds are divided into two parts. One part includes points lying in the putative overlapping region and the other includes points located in the putative non-overlapping region. Then, EfficientPEAL performs localized attention with the putative overlapping points. The proposed attention module significantly reduces the computational complexity of the model while achieving competitive performance. Extensive experiments on 3DMatch/3DLoMatch, ScanNet, and KITTI datasets demonstrate its effectiveness.
Junle Yu, Zhehao Shen, Yongwei Miao
Expert Syst. Appl.3
2024 HiFi4G: High-Fidelity Human Performance Rendering via Compact Gaussian Splatting
abstract
We have recently seen tremendous progress in photo-real human modeling and rendering. Yet, efficiently ren-dering realistic human performance and integrating it into the rasterization pipeline remains challenging. In this pa-per, we present HiFi4G, an explicit and compact Gaussian-based approach for high-fidelity human performance ren-dering from dense footage. Our core intuition is to marry the 3D Gaussian representation with non-rigid tracking, achieving a compact and compression-friendly representation. We first propose a dual-graph mechanism to obtain motion priors, with a coarse deformation graph for effective initialization and a fine-grained Gaussian graph to en-force subsequent constraints. Then, we utilize a 4D Gaus-sian optimization scheme with adaptive spatial-temporal regularizers to effectively balance the non-rigid prior and Gaussian updating. We also present a companion compression scheme with residual compensation for immersive experiences on various platforms. It achieves a substantial compression rate of approximately 25 times, with less than 2MB of storage per frame. Extensive experiments demon-strate the effectiveness of our approach, which significantly outperforms existing approaches in terms of optimization speed, rendering quality, and storage overhead. Project page: https://nowheretrix.github.io/HiFi4G/.
Yuheng Jiang, Zhehao Shen, Penghao Wang 0003, Zhuo Su 0006, Yingliang Zhang, Jingyi Yu 0001, Lan Xu 0003
CVPR2
2024 Robust Dual Gaussian Splatting for Immersive Human-centric Volumetric Videos
abstract
Volumetric video represents a transformative advancement in visual media, enabling users to freely navigate immersive virtual experiences and narrowing the gap between digital and real worlds. However, the need for extensive manual intervention to stabilize mesh sequences and the generation of excessively large assets in existing workflows impedes broader adoption. In this paper, we present a novel Gaussian-based approach, dubbed DualGS , for real-time and high-fidelity playback of complex human performance with excellent compression ratios. Our key idea in DualGS is to separately represent motion and appearance using the corresponding skin and joint Gaussians. Such an explicit disentanglement can significantly reduce motion redundancy and enhance temporal coherence. We begin by initializing the DualGS and anchoring skin Gaussians to joint Gaussians at the first frame. Subsequently, we employ a coarse-to-fine training strategy for frame-by-frame human performance modeling. It includes a coarse alignment phase for overall motion prediction as well as a fine-grained optimization for robust tracking and high-fidelity rendering. To integrate volumetric video seamlessly into VR environments, we efficiently compress motion using entropy encoding and appearance using codec compression coupled with a persistent codebook. Our approach achieves a compression ratio of up to 120 times, only requiring approximately 350KB of storage per frame. We demonstrate the efficacy of our representation through photo-realistic, free-view experiences on VR headsets, enabling users to immersively watch musicians in performance and feel the rhythm of the notes at the performers' fingertips. Project page: https://nowheretrix.github.io/DualGS/.
Yuheng Jiang, Zhehao Shen, Yize Wu, Yingliang Zhang, Jingyi Yu 0001, Lan Xu 0003
ACM Trans. Graph.2
2023 Instant-NVR: Instant Neural Volumetric Rendering for Human-object Interactions from Monocular RGBD Stream
abstract
Convenient 4D modeling of human-object interactions is essential for numerous applications. However, monocular tracking and rendering of complex interaction scenarios remain challenging. In this paper, we propose Instant-NVR, a neural approach for instant volumetric human-object tracking and rendering using a single RGBD camera. It bridges traditional non-rigid tracking with recent instant radiance field techniques via a multi-thread tracking-rendering mechanism. In the tracking front-end, we adopt a robust human-object capture scheme to provide sufficient motion priors. We further introduce a separated instant neural representation with a novel hybrid deformation module for the interacting scene. We also provide an on-the-fly reconstruction scheme of the dynamic/static radiance fields via efficient motion-prior searching. Moreover, we introduce an online key frame selection scheme and a rendering-aware refinement strategy to significantly improve the appearance details for online novel-view synthesis. Extensive experiments demonstrate the effectiveness and efficiency of our approach for the instant generation of human-object radiance fields on the fly, notably achieving real-time photo-realistic novel view synthesis under complex human-object interactions. Project page: https://nowheretrix.github.io/Instant-NVR/.
Yuheng Jiang, Kaixin Yao, Zhuo Su 0006, Zhehao Shen, Haimin Luo, Lan Xu 0003
CVPR4