Yuze Wang 0006

dblp:219/1012-6 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0000-7676-3408ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
abstract
Reconstructing 3D objects from a single image is a long-standing challenge, particularly under real-world occlusions. While recent diffusion-based view synthesis models can generate consistent novel views from a single RGB image, they generally assume fully visible inputs and struggle when parts of the object are occluded, leading to inconsistent views and degraded 3D reconstruction quality. To address this limitation, we propose DeOcc-1-to-3, an end-to-end framework for occlusion-aware multi-view generation. Our method directly synthesizes six structurally consistent novel views from a single partially occluded image, enabling downstream 3D reconstruction without requiring prior inpainting or manual annotations. We design a self-supervised training pipeline that leverages occluded–unoccluded image pairs and pseudo-ground-truth views to guide structure-aware completion and view consistency. Without modifying the original architecture, we fully fine-tune the diffusion model to jointly learn completion and multi-view generation. Additionally, we introduce the first benchmark for occlusion-aware reconstruction, covering diverse occlusion levels, object categories, and mask patterns, providing a standardized evaluation protocol.
Yansong Qu, Shaohui Dai, Yuze Wang 0006, You Shen, Shengchuan Zhang, Liujuan Cao
AAAI4
2026 GSBrief: A Globally Consistent Descriptor with 3D Gaussian Splatting for Visual Localization
abstract
Visual localization is a critical component in a wide range of applications. Recent advancements in scene representation, particularly the use of 3D Gaussian Splatting (3D GS), have introduced promising opportunities for enhancing localization pipelines. However, effectively leveraging the features of 3D GS while maintaining full integration of texture, geometric context, and global consistency remains a significant challenge. In this paper, we introduce GSBrief, a novel, globally consistent descriptor designed specifically for 3D GS based visual localization. GSBrief captures scene features through a structured extraction process that seamlessly incorporates both texture and geometry information. The resulting descriptors are engineered to be invariant to scale, position, and rotation, ensuring robust performance across a wide range of conditions. Building on GSBrief, we propose GSBriefNet, a regression network designed to predict GSBrief descriptors, which is based on the Swin Transformer architecture. GSBriefNet employs a Siamese network design to enforce global consistency and simultaneously regresses point maps to recover the ground truth scale. The predicted GSBrief descriptors can be directly applied to tasks such as relative camera pose estimation, relocalization, and Simultaneous Localization and Mapping (SLAM). We demonstrate the effectiveness of our approach through experiments on benchmark datasets, including ScanNet, 7 Scenes, Cambridge Landmarks, TUM RGB-D and Bonn. The results show that our method achieves state-of-the-art performance across all three tasks, providing a robust solution for visual localization.
Junyi Wang 0001, Yuze Wang 0006, Wantong Duan
VR2
2026 Visual Camera Localization by Globally Consistent Descriptor Learning and Combined Bundle Adjustment
abstract
In this paper, we introduce a brand-new localization pipeline designed to comprehensively leverage both hand-crafted and learned features, operating at two distinct levels (point-level and object-level) while simultaneously recovering scene scale from the RGB input. The pipeline integrates a learned globally consistent descriptor matching process for initial camera pose estimation, followed by a pose optimization phase that synergistically combines various features. To generate learned descriptors, we propose a siamese Globally Consistent Feature Descriptor Network (GCFDNet), which accepts a pair of images, Inertial Measurement Unit (IMU) data, and pose sequences as inputs, producing both the image descriptors and the relative camera pose as outputs. The strengths of GCFDNet manifest in two key aspects. First, by incorporating a spatial-to-temporal feature fusion module, GCFDNet enhances relative pose regression, meanwhile enabling accurate scene scale estimation. Second, we devise a loss function that balances descriptor similarity and distance, thereby improving the quality of descriptor learning. Using the initial camera poses derived from GCFDNet, we establish data associations across multiple frames and subsequently propose a combined Bundle Adjustment (BA) optimization framework that integrates hand-crafted features, learned descriptors, and semantic objects. To evaluate the localization performance, we conduct experiments across diverse datasets, including EuRoC, ScanNet, 7 Scenes, TUM RGB-D, and Bonn. The results demonstrate state-of-the-art performance in both static and dynamic scenes, outperforming existing methods. Additionally, we present ablation studies on GCFDNet and the combined BA process to further substantiate the efficacy of our approach.
Junyi Wang 0001, Yuze Wang 0006, Chen Wang 0043
IEEE Trans. Image Process.2
2026 Taking Language Embedded 3D Gaussian Splatting into the Wild
abstract
Recent advances in leveraging large-scale Internet photo collections for 3D reconstruction have enabled immersive virtual exploration of landmarks and historic sites worldwide. However, existing methods primarily focus on visual appearance reconstruction, often overlooking the interactive semantic understanding of these 3D scenes (e.g., identifying specific building parts or scene details), which remains largely confined to browsing static text-image pairs. Therefore, can we draw inspiration from 3D in-the-wild reconstruction techniques and use unconstrained photo collections to create an immersive approach for comprehensive 3D scene understanding beyond mere visual appearance? To this end, we extend language embedded 3D Gaussian splatting (3DGS) and propose a novel framework for open-vocabulary scene understanding from unconstrained photo collections. Specifically, we first render multiple appearance images from the same viewpoint as the unconstrained image with the reconstructed radiance field, then extract multi-appearance CLIP features and two types of language feature uncertainty maps-transient and appearance uncertainty-derived from the multi-appearance features to guide the subsequent optimization process. Next, we propose a transient uncertainty-aware autoencoder, a multi-appearance language field 3DGS representation, and a post-ensemble strategy to effectively compress, learn, and fuse language features from multiple appearances. Finally, to quantitatively evaluate our method, we introduce PT-OVS, a new benchmark dataset for assessing open-vocabulary segmentation performance on unconstrained photo collections. Experimental results show that our method outperforms existing methods, delivering accurate open-vocabulary segmentation and enabling applications such as interactive roaming with open-vocabulary queries, architectural style pattern recognition, and 3D scene editing. Visit our project page at Project Page.
Yuze Wang 0006, Junyi Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2025 Seg-Wild: Interactive Segmentation based on 3D Gaussian Splatting for Unconstrained Image Collections
Yongtang Bao, Chengjie Tang, Yuze Wang 0006
ACM Multimedia3
2025 3D Gaussian Splatting based Scene-independent Relocalization with Unidirectional and Bidirectional Feature Fusion
abstract
Visual localization is a critical component across various domains. The recent emergence of novel scene representations, such as 3D Gaussian Splatting (3D GS), introduces new opportunities for advancing localization pipelines. In this paper, we propose a novel 3D GS-based framework for RGB based, scene-independent camera relocalization, with three main contributions. First, we design a two-stage pipeline with fully exploiting 3D GS. The pipeline consists of an initial stage, which utilizes 2D-3D correspondences between image pixels and 3D Gaussians, followed by pose refinement using the rendered image by 3D GS. Second, we introduce a 3D GS based Relocalization Network, termed GS-RelocNet, to establish correspondences for initial camera pose estimation. Additionally, we present a refinement network that further optimizes the camera pose. Third, we propose a unidirectional 2D-3D feature fusion module and a bidirectional image feature fusion module, integrated into GS-RelocNet and the refinement network, respectively, to enhance feature sharing across the two stages. Experimental results on public 7 Scenes, Cambridge Landmarks, TUM RGB-D and Bonn demonstrate state-of-the-art performance. Furthermore, the beneficial effects of the two feature fusion modules and pose refinement are also highlighted. In summary, we believe that the proposed framework can be a novel universal localization pipeline for further research.
Junyi Wang 0001, Yuze Wang 0006, Wantong Duan
NeurIPS2
2025 Efficient interactive segmentation of three-dimensional Gaussians with optimal view selection
Yongtang Bao, Chengjie Tang, Yuze Wang 0006, Yutong Qi
Eng. Appl. Artif. Intell.3
2025 RISE-Editing: Rotation-invariant neural point fields with interactive segmentation for fine-grained and efficient editing
Yuze Wang 0006, Junyi Wang 0001, Chen Wang 0043
Neural Networks1
2025 Look at the Sky: Sky-Aware Efficient 3D Gaussian Splatting in the Wild
abstract
Photos taken in unconstrained tourist environments often present challenges for accurate 3D scene reconstruction due to variable appearances and transient occlusions, which can introduce artifacts in novel view synthesis. Recently, in-the-wild 3D scene reconstruction has been achieved realistic rendering with Neural Radiance Fields (NeRFs). With the advancement of 3D Gaussian Splatting (3DGS), some methods also attempt to reconstruct 3D scenes from unconstrained photo collections and achieve real-time rendering. However, the rapid convergence of 3DGS is misaligned with the slower convergence of neural network-based appearance encoder and transient mask predictor, hindering the reconstruction efficiency. To address this, we propose a novel sky-aware framework for scene reconstruction from unconstrained photo collection using 3DGS. Firstly, we observe that the learnable per-image transient mask predictor in previous work is unnecessary. By introducing a simple yet efficient greedy supervision strategy, we directly utilize the pseudo mask generated by a pretrained semantic segmentation network as the transient mask, thereby achieving more efficient and higher quality in-the-wild 3D scene reconstruction. Secondly, we find that separately estimating appearance embeddings for the sky and building significantly improves reconstruction efficiency and accuracy. We analyze the underlying reasons and introduce a neural sky module to generate diverse skies from latent sky embeddings extract from unconstrained images. Finally, we propose a mutual distillation learning strategy to constrain sky and building appearance embeddings within the same latent space, further enhancing reconstruction efficiency and quality. Extensive experiments on multiple datasets demonstrate that the proposed framework outperforms existing methods in novel view and appearance synthesis, offering superior rendering quality with faster convergence and rendering speed.
Yuze Wang 0006, Junyi Wang 0001, Ruicheng Gao, Yansong Qu, Wantong Duan
IEEE Trans. Vis. Comput. Graph.1
2024 SCARF: Scalable Continual Learning Framework for Memory-efficient Multiple Neural Radiance Fields
abstract
Abstract This paper introduces a novel continual learning framework for synthesising novel views of multiple scenes, learning multiple 3D scenes incrementally, and updating the network parameters only with the training data of the upcoming new scene. We build on Neural Radiance Fields (NeRF), which uses multi‐layer perceptron to model the density and radiance field of a scene as the implicit function. While NeRF and its extensions have shown a powerful capability of rendering photo‐realistic novel views in a single 3D scene, managing these growing 3D NeRF assets efficiently is a new scientific problem. Very few works focus on the efficient representation or continuous learning capability of multiple scenes, which is crucial for the practical applications of NeRF. To achieve these goals, our key idea is to represent multiple scenes as the linear combination of a cross‐scene weight matrix and a set of scene‐specific weight matrices generated from a global parameter generator. Furthermore, we propose an uncertain surface knowledge distillation strategy to transfer the radiance field knowledge of previous scenes to the new model. Representing multiple 3D scenes with such weight matrices significantly reduces memory requirements. At the same time, the uncertain surface distillation strategy greatly overcomes the catastrophic forgetting problem and maintains the photo‐realistic rendering quality of previous scenes. Experiments show that the proposed approach achieves state‐of‐the‐art rendering quality of continual learning NeRF on NeRF‐Synthetic, LLFF, and TanksAndTemples datasets while preserving extra low storage cost.
Yuze Wang 0006, Junyi Wang 0001, Chen Wang 0043, Wantong Duan, Yongtang Bao
Comput. Graph. Forum1
2023 SG-NeRF: Semantic-guided Point-based Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) can successfully reconstruct room-scale scenes and achieve photo-realistic novel view synthesis results with densely captured input images. However, capturing hundreds of high-quality images in a single room is extremely laborious. We tackle this problem by greatly reducing the number of images input to NeRF while maintaining high-quality rendering results in a room-scale scene. In this paper, we propose semantic-guided point-based NeRF (SG-NeRF), which is capable of reconstructing the radiance field of a room-scale scene with tens of images. To this end, we leverage sparse 3D point clouds with neural features to be the geometry constraints of NeRF optimization and semantic prediction of both 2D images and 3D point clouds to guide the neighboring neural points searching at the ray marching procedure. With the semantic guidance, the sampled query points are capable of searching for neighboring neural points, which are structurally related to the query points accurately in a large area since of the unevenly distributed sparse point clouds. Extensive experimental results demonstrate that our approach outperforms previous state-of-the-art methods.
Yansong Qu, Yuze Wang 0006
ICME2
2023 RIP-NeRF: Learning Rotation-Invariant Point-based Neural Radiance Field for Fine-grained Editing and Compositing
abstract
Neural Radiance Field (NeRF) shows dramatic results in synthesising novel views. However, existing controllable and editable NeRF methods are still incapable of both fine-grained editing and cross-scene compositing, greatly limiting their creative editing as well as potential applications. When the radiance field is fine-grained edited and composited, a severe drawback is that varying the orientation of the corresponding explicit scaffold, such as point, mesh, volume, etc., may lead to the degradation of rendering quality. In this work, by taking the respective strengths of the implicit NeRF-based representation and the explicit point-based representation, we present a novel Rotation-Invariant Point-based NeRF (RIP-NeRF) for both fine-grained editing and cross-scene compositing of the radiance field. Specifically, we introduce a novel point-based radiance field representation to replace the Cartesian coordinate as the network input. This rotation-invariant representation is met by carefully designing a Neural Inverse Distance Weighting Interpolation (NIDWI) module to aggregate neural points, significantly improving the rendering quality for fine-grained editing. To achieve cross-scene compositing, we disentangle the rendering module and the neural point-based representation in NeRF. After simply manipulating the corresponding neural points, a cross-scene neural rendering module is applied to achieve controllable cross-scene compositing without retraining. The advantages of our RIP-NeRF on editing quality and capability are demonstrated by extensive editing and compositing experiments on room-scale real scenes and synthetic objects with complex geometry.
Yuze Wang 0006, Junyi Wang 0001, Yansong Qu
ICMR1