Junyi Wang 0001

dblp:14/948-1 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-3191-1662ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GSBrief: A Globally Consistent Descriptor with 3D Gaussian Splatting for Visual Localization
abstract
Visual localization is a critical component in a wide range of applications. Recent advancements in scene representation, particularly the use of 3D Gaussian Splatting (3D GS), have introduced promising opportunities for enhancing localization pipelines. However, effectively leveraging the features of 3D GS while maintaining full integration of texture, geometric context, and global consistency remains a significant challenge. In this paper, we introduce GSBrief, a novel, globally consistent descriptor designed specifically for 3D GS based visual localization. GSBrief captures scene features through a structured extraction process that seamlessly incorporates both texture and geometry information. The resulting descriptors are engineered to be invariant to scale, position, and rotation, ensuring robust performance across a wide range of conditions. Building on GSBrief, we propose GSBriefNet, a regression network designed to predict GSBrief descriptors, which is based on the Swin Transformer architecture. GSBriefNet employs a Siamese network design to enforce global consistency and simultaneously regresses point maps to recover the ground truth scale. The predicted GSBrief descriptors can be directly applied to tasks such as relative camera pose estimation, relocalization, and Simultaneous Localization and Mapping (SLAM). We demonstrate the effectiveness of our approach through experiments on benchmark datasets, including ScanNet, 7 Scenes, Cambridge Landmarks, TUM RGB-D and Bonn. The results show that our method achieves state-of-the-art performance across all three tasks, providing a robust solution for visual localization.
Junyi Wang 0001, Yuze Wang 0006, Wantong Duan
VR1
2026 Visual Camera Localization by Globally Consistent Descriptor Learning and Combined Bundle Adjustment
abstract
In this paper, we introduce a brand-new localization pipeline designed to comprehensively leverage both hand-crafted and learned features, operating at two distinct levels (point-level and object-level) while simultaneously recovering scene scale from the RGB input. The pipeline integrates a learned globally consistent descriptor matching process for initial camera pose estimation, followed by a pose optimization phase that synergistically combines various features. To generate learned descriptors, we propose a siamese Globally Consistent Feature Descriptor Network (GCFDNet), which accepts a pair of images, Inertial Measurement Unit (IMU) data, and pose sequences as inputs, producing both the image descriptors and the relative camera pose as outputs. The strengths of GCFDNet manifest in two key aspects. First, by incorporating a spatial-to-temporal feature fusion module, GCFDNet enhances relative pose regression, meanwhile enabling accurate scene scale estimation. Second, we devise a loss function that balances descriptor similarity and distance, thereby improving the quality of descriptor learning. Using the initial camera poses derived from GCFDNet, we establish data associations across multiple frames and subsequently propose a combined Bundle Adjustment (BA) optimization framework that integrates hand-crafted features, learned descriptors, and semantic objects. To evaluate the localization performance, we conduct experiments across diverse datasets, including EuRoC, ScanNet, 7 Scenes, TUM RGB-D, and Bonn. The results demonstrate state-of-the-art performance in both static and dynamic scenes, outperforming existing methods. Additionally, we present ablation studies on GCFDNet and the combined BA process to further substantiate the efficacy of our approach.
Junyi Wang 0001, Yuze Wang 0006, Chen Wang 0043
IEEE Trans. Image Process.1
2026 Taking Language Embedded 3D Gaussian Splatting into the Wild
abstract
Recent advances in leveraging large-scale Internet photo collections for 3D reconstruction have enabled immersive virtual exploration of landmarks and historic sites worldwide. However, existing methods primarily focus on visual appearance reconstruction, often overlooking the interactive semantic understanding of these 3D scenes (e.g., identifying specific building parts or scene details), which remains largely confined to browsing static text-image pairs. Therefore, can we draw inspiration from 3D in-the-wild reconstruction techniques and use unconstrained photo collections to create an immersive approach for comprehensive 3D scene understanding beyond mere visual appearance? To this end, we extend language embedded 3D Gaussian splatting (3DGS) and propose a novel framework for open-vocabulary scene understanding from unconstrained photo collections. Specifically, we first render multiple appearance images from the same viewpoint as the unconstrained image with the reconstructed radiance field, then extract multi-appearance CLIP features and two types of language feature uncertainty maps-transient and appearance uncertainty-derived from the multi-appearance features to guide the subsequent optimization process. Next, we propose a transient uncertainty-aware autoencoder, a multi-appearance language field 3DGS representation, and a post-ensemble strategy to effectively compress, learn, and fuse language features from multiple appearances. Finally, to quantitatively evaluate our method, we introduce PT-OVS, a new benchmark dataset for assessing open-vocabulary segmentation performance on unconstrained photo collections. Experimental results show that our method outperforms existing methods, delivering accurate open-vocabulary segmentation and enabling applications such as interactive roaming with open-vocabulary queries, architectural style pattern recognition, and 3D scene editing. Visit our project page at Project Page.
Yuze Wang 0006, Junyi Wang 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Visual Localization using Hybrid Feature Grid and Learned Weighted Global Point Cloud
abstract
To fully leverage diverse scene representations for visual relocalization, we propose a novel localization framework that systematically establishes inter-frame relationships and integrates multiple feature modalities. Our localization pipeline comprises three key stages, containing initial pose estimation using local point cloud structure, pose refinement by hand-crafted features and 3D Gaussians, and pose confidence estimation through a leaned global representation. Specifically, the initial stage begins with aligning a known source point cloud to a predicted local Target Point Cloud (TPC) using a registration algorithm. For pose refinement, we introduce the Hybrid Feature Grid (HFG), which fuses hand-crafted points and 3D Gaussians to enrich texture cues. To assess pose reliability, we propose the learned Weighted Global Point Cloud (WGPC), aggregating multi-frame information to enhance confidence estimation. To jointly learn TPC, HFG, and WGPC, we design a Siamese Localization Network (SiaLocNet) featuring three core innovations, including learning trajectory-based features for the limitation of single-view inputs, a feature fusion module to facilitate the construction of the three core structures. and an inverse self Chamfer Distance along with a shape-aware term to improve the robustness of WGPC. Extensive experiments on the 7 Scenes and Cambridge Landmarks datasets demonstrate that our method achieves state-ofthe-art performance across both indoor and outdoor environments.
Junyi Wang 0001
ACM Multimedia1
2025 3D Gaussian Splatting based Scene-independent Relocalization with Unidirectional and Bidirectional Feature Fusion
abstract
Visual localization is a critical component across various domains. The recent emergence of novel scene representations, such as 3D Gaussian Splatting (3D GS), introduces new opportunities for advancing localization pipelines. In this paper, we propose a novel 3D GS-based framework for RGB based, scene-independent camera relocalization, with three main contributions. First, we design a two-stage pipeline with fully exploiting 3D GS. The pipeline consists of an initial stage, which utilizes 2D-3D correspondences between image pixels and 3D Gaussians, followed by pose refinement using the rendered image by 3D GS. Second, we introduce a 3D GS based Relocalization Network, termed GS-RelocNet, to establish correspondences for initial camera pose estimation. Additionally, we present a refinement network that further optimizes the camera pose. Third, we propose a unidirectional 2D-3D feature fusion module and a bidirectional image feature fusion module, integrated into GS-RelocNet and the refinement network, respectively, to enhance feature sharing across the two stages. Experimental results on public 7 Scenes, Cambridge Landmarks, TUM RGB-D and Bonn demonstrate state-of-the-art performance. Furthermore, the beneficial effects of the two feature fusion modules and pose refinement are also highlighted. In summary, we believe that the proposed framework can be a novel universal localization pipeline for further research.
Junyi Wang 0001, Yuze Wang 0006, Wantong Duan
NeurIPS1
2025 RISE-Editing: Rotation-invariant neural point fields with interactive segmentation for fine-grained and efficient editing
Yuze Wang 0006, Junyi Wang 0001, Chen Wang 0043
Neural Networks2
2025 Look at the Sky: Sky-Aware Efficient 3D Gaussian Splatting in the Wild
abstract
Photos taken in unconstrained tourist environments often present challenges for accurate 3D scene reconstruction due to variable appearances and transient occlusions, which can introduce artifacts in novel view synthesis. Recently, in-the-wild 3D scene reconstruction has been achieved realistic rendering with Neural Radiance Fields (NeRFs). With the advancement of 3D Gaussian Splatting (3DGS), some methods also attempt to reconstruct 3D scenes from unconstrained photo collections and achieve real-time rendering. However, the rapid convergence of 3DGS is misaligned with the slower convergence of neural network-based appearance encoder and transient mask predictor, hindering the reconstruction efficiency. To address this, we propose a novel sky-aware framework for scene reconstruction from unconstrained photo collection using 3DGS. Firstly, we observe that the learnable per-image transient mask predictor in previous work is unnecessary. By introducing a simple yet efficient greedy supervision strategy, we directly utilize the pseudo mask generated by a pretrained semantic segmentation network as the transient mask, thereby achieving more efficient and higher quality in-the-wild 3D scene reconstruction. Secondly, we find that separately estimating appearance embeddings for the sky and building significantly improves reconstruction efficiency and accuracy. We analyze the underlying reasons and introduce a neural sky module to generate diverse skies from latent sky embeddings extract from unconstrained images. Finally, we propose a mutual distillation learning strategy to constrain sky and building appearance embeddings within the same latent space, further enhancing reconstruction efficiency and quality. Extensive experiments on multiple datasets demonstrate that the proposed framework outperforms existing methods in novel view and appearance synthesis, offering superior rendering quality with faster convergence and rendering speed.
Yuze Wang 0006, Junyi Wang 0001, Ruicheng Gao, Yansong Qu, Wantong Duan
IEEE Trans. Vis. Comput. Graph.2
2024 SCARF: Scalable Continual Learning Framework for Memory-efficient Multiple Neural Radiance Fields
abstract
Abstract This paper introduces a novel continual learning framework for synthesising novel views of multiple scenes, learning multiple 3D scenes incrementally, and updating the network parameters only with the training data of the upcoming new scene. We build on Neural Radiance Fields (NeRF), which uses multi‐layer perceptron to model the density and radiance field of a scene as the implicit function. While NeRF and its extensions have shown a powerful capability of rendering photo‐realistic novel views in a single 3D scene, managing these growing 3D NeRF assets efficiently is a new scientific problem. Very few works focus on the efficient representation or continuous learning capability of multiple scenes, which is crucial for the practical applications of NeRF. To achieve these goals, our key idea is to represent multiple scenes as the linear combination of a cross‐scene weight matrix and a set of scene‐specific weight matrices generated from a global parameter generator. Furthermore, we propose an uncertain surface knowledge distillation strategy to transfer the radiance field knowledge of previous scenes to the new model. Representing multiple 3D scenes with such weight matrices significantly reduces memory requirements. At the same time, the uncertain surface distillation strategy greatly overcomes the catastrophic forgetting problem and maintains the photo‐realistic rendering quality of previous scenes. Experiments show that the proposed approach achieves state‐of‐the‐art rendering quality of continual learning NeRF on NeRF‐Synthetic, LLFF, and TanksAndTemples datasets while preserving extra low storage cost.
Yuze Wang 0006, Junyi Wang 0001, Chen Wang 0043, Wantong Duan, Yongtang Bao
Comput. Graph. Forum2
2024 Multi-level feature fusion and joint refinement for simultaneous object pose estimation and camera localization
Junyi Wang 0001
Neural Networks1
2024 Visual camera relocalization using both hand-crafted and learned features
Junyi Wang 0001
Pattern Recognit.1
2023 Scene-independent Localization by Learning Residual Coordinate Map with Cascaded Localizers
abstract
Visual localization plays an essential role in a variety of different fields. The indirect learning based method obtains an excellent performance, but it requests a training process in the target scene before the localization. To achieve deep scene-independent localization, we start by proposing the representation called residual coordinate map between a pair of images. Based on the structure, we put forward a network called SILocNet with the proposed residual coordinate map as the output. The network consists of feature extraction, multi-level feature fusion and transformer based coordinate decoder. Moreover, considering the dynamic scene, we introduce an additional segmentation branch that distinguishes fixed and dynamic parts to promote network perception. With SILocNet in place, a cascaded localizer design is presented for reducing the accumulative error. Meanwhile, the simple mathematical analysis behind the cascaded localizers is also provided. To verify how well our algorithm could perform, we conduct experiments on static 7 Scenes, ScanNet and dynamic TUM RGB-D. In particular, we train the network on ScanNet and test it on 7 Scenes and TUM RGB-D to demonstrate the generality performance. All experiments demonstrate superior performance to other existing methods. Additionally, the effects of the cascaded localizer design, feature fusion, transformer based coordinate decoder and segmentation loss are also discussed.
Junyi Wang 0001
ISMAR1
2023 RIP-NeRF: Learning Rotation-Invariant Point-based Neural Radiance Field for Fine-grained Editing and Compositing
abstract
Neural Radiance Field (NeRF) shows dramatic results in synthesising novel views. However, existing controllable and editable NeRF methods are still incapable of both fine-grained editing and cross-scene compositing, greatly limiting their creative editing as well as potential applications. When the radiance field is fine-grained edited and composited, a severe drawback is that varying the orientation of the corresponding explicit scaffold, such as point, mesh, volume, etc., may lead to the degradation of rendering quality. In this work, by taking the respective strengths of the implicit NeRF-based representation and the explicit point-based representation, we present a novel Rotation-Invariant Point-based NeRF (RIP-NeRF) for both fine-grained editing and cross-scene compositing of the radiance field. Specifically, we introduce a novel point-based radiance field representation to replace the Cartesian coordinate as the network input. This rotation-invariant representation is met by carefully designing a Neural Inverse Distance Weighting Interpolation (NIDWI) module to aggregate neural points, significantly improving the rendering quality for fine-grained editing. To achieve cross-scene compositing, we disentangle the rendering module and the neural point-based representation in NeRF. After simply manipulating the corresponding neural points, a cross-scene neural rendering module is applied to achieve controllable cross-scene compositing without retraining. The advantages of our RIP-NeRF on editing quality and capability are demonstrated by extensive editing and compositing experiments on room-scale real scenes and synthetic objects with complex geometry.
Yuze Wang 0006, Junyi Wang 0001, Yansong Qu
ICMR2
2021 Camera Relocalization using Deep Point Cloud Generation and Hand-crafted Feature Refinement
abstract
Visual localization plays an indispensable role in robotics. Both learning and hand-crafted feature based methods for relocalization process keep their effectiveness and weakness. However, current algorithms seldom consider these two kinds of features under one framework. In this paper, focusing on this task, we propose a novel relocalization framework for RGB or RGB-D data source, which is composed of coarse localization process by learning features and pose refinement by hand-crafted features. In particular, coarse stage contains deep point cloud generation and registration. In this stage, instead of regressing camera pose directly, the paper novelly designs a neural network called PGNet to construct sparse point cloud with RGB or RGB-D as inputs. Further more, by means of training set, hand-crafted feature space is established. Based on the obtained camera pose in coarse stage, accurate point-to-point correspondences are set up through searching the space. Then accurate camera pose is obtained by applying RANSAC to correspondences or solving PnP. Finally, experiments on both outdoor and indoor benchmark datasets demonstrate state-of-the-art performance over other existing methods.
Junyi Wang 0001
ICRA1