Shihe Shen

dblp:369/7707 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Compressing Streamable Free-Viewpoint Videos to 0.1 MB per Frame
abstract
The success of 3D Gaussian Splatting (3DGS) in static scenes has inspired numerous attempts to construct Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos. Despite advancements in current techniques, simultaneously achieving photo-realistic view synthesis results, fast on-the-fly training, real-time rendering, and low storage costs remains a formidable problem. To address these challenges, we propose the first Gaussian-based streamable FVV intelligent compression framework named iFVC. Specifically, we utilize an anchor-based Gaussian representation to model the scene. To achieve on-the-fly training, we propose a Binary Transformation Cache (BTC) to model the dynamic changes between adjacent timesteps, which not only ensures compactness but also supports precise bit rate estimation. Furthermore, we carefully design a high-resolution transformation tri-plane assisted by a saliency grid as our BTC, allowing for accurate dynamic capture. The entire pipeline is regarded as a joint optimization of rate and distortion to achieve optimal compression performance. Experiments on widely used datasets demonstrate the state-of-the-art performance of our framework in both synthesis quality and efficiency, i.e., achieving per-frame training in 13 seconds with a storage cost of 0.1 MB and real-time rendering at 120 FPS.
Luyang Tang, Yongqi Zhai, Shihe Shen, Ronggang Wang
AAAI5
2025 ADC-GS: Pose-Free 3D Gaussian Splatting with Adaptive Depth Consistency
abstract
Recently proposed 3D Gaussian Splatting (3DGS) has achieved state-of-the-art results in the fields of novel view synthesis, but it heavily relies on pre-computed camera poses. Although recent methods mitigate by leveraging explicit representations achieve novel view synthesis without requiring camera poses, Gaussians lack an appropriate growth strategy and accurate geometric information, which may lead to noticeable artifacts and distortions, particularly in complex scenes with large camera movements. To address these issues, we propose ADC-GS, where we introduce an Adaptive Depth Alignment (ADA) strategy to utilize monocular depth information from adjacent frames as guidance for camera pose optimization and provides geometric supervision through depth consistency for training global Gaussians. Furthermore, we demonstrate the drawbacks of the progressive growth strategy in Gaussians and further enhance reconstruction quality by refining the growth strategy. Extensive experiments on the challenging Tanks and Temples dataset show that our method achieves state-of-the-art results in both novel view synthesis quality and pose estimation accuracy.
Runling Liu, Shihe Shen, Guanhua Wu, Zhanke Wang, Ronggang Wang
ICASSP2
2024 Surface-Centric Modeling for High-Fidelity Generalizable Neural Surface Reconstruction
Shihe Shen, Kaiqiang Xiong, Huachen Gao, Jianbo Jiao, Ronggang Wang
ECCV (32)2
2024 Disentangled Generation and Aggregation for Robust Radiance Fields
Shihe Shen, Huachen Gao, Wangze Xu, Luyang Tang, Kaiqiang Xiong, Jianbo Jiao, Ronggang Wang
ECCV (49)1
2024 MVPGS: Excavating Multi-view Priors for Gaussian Splatting from Sparse Input Views
Wangze Xu, Huachen Gao, Shihe Shen, Jianbo Jiao, Ronggang Wang
ECCV (47)3
2024 FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth Consistency
abstract
Learning neural radiance fields (NeRF) without camera poses has been widely studied. However, recent methods lack explicit and effective supervision for pose estimation, resulting in ambiguous optimization of camera pose and NeRF geometry during joint training, particularly in scenarios involving large camera movements. In this paper, we propose FDCNeRF that leverages the direction information contained in the RGB-based optical flow and depth-based virtual flow as a direct guidance for camera pose optimization to reduce pose-geometry ambiguity. Additionally, we introduce Adaptive Pose-Aware Sampling (APAS) to replace the previous random ray sampling strategy, which reduces the difficulty of pose learning in early stages and preserves the diversity of rays in later stages. Experiments on the challenging Tanks and Temples dataset demonstrate that our method achieves state-of-the-art results in both novel view synthesis quality and pose estimation accuracy.
Huachen Gao, Shihe Shen, Kaiqiang Xiong, Zhirui Gao, Yugui Xie, Ronggang Wang
ICASSP2
2023 GenS: Generalizable Neural Surface Reconstruction from Multi-View Images
abstract
Combining the signed distance function (SDF) and differentiable volume rendering has emerged as a powerful paradigm for surface reconstruction from multi-view images without 3D supervision. However, current methods are impeded by requiring long-time per-scene optimizations and cannot generalize to new scenes. In this paper, we present GenS, an end-to-end generalizable neural surface reconstruction model. Unlike coordinate-based methods that train a separate network for each scene, we construct a generalized multi-scale volume to directly encode all scenes. Compared with existing solutions, our representation is more powerful, which can recover high-frequency details while maintaining global smoothness. Meanwhile, we introduce a multi-scale feature-metric consistency to impose the multi-view consistency in a more discriminative multi-scale feature space, which is robust to the failures of the photometric consistency. And the learnable feature can be self-enhanced to continuously improve the matching accuracy and mitigate aggregation ambiguity. Furthermore, we design a view contrast loss to force the model to be robust to those regions covered by few viewpoints through distilling the geometric prior from dense input to sparse input. Extensive experiments on popular benchmarks show that our model can generalize well to new scenes and outperform existing state-of-the-art methods even those employing ground-truth depth supervision. Code will be available at https://github.com/prstrive/GenS.
Luyang Tang, Shihe Shen, Fanqi Yu, Ronggang Wang
NeurIPS4