Yijia He

dblp:195/8944 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-4358-7393ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Rotation-Translation Decoupled Solution for Visual-Inertial Initialization and Online Spatial-Temporal Calibration
abstract
We propose a novel initialization and online spatial-temporal calibration method for visual-inertial odometry (VIO), which decouples rotation and translation estimation to achieve higher accuracy and better robustness. Existing initialization methods suffer from limited accuracy or robustness (e.g., in scenarios with small translational motion) and rarely integrate simultaneous spatial-temporal calibration during initialization, despite its considerable practical value. Our proposed method leverages rotation-translation decoupling constraints to enable simultaneous estimation of gyroscope bias, extrinsic rotation, and camera-IMU time offset-even under pure rotational motion. Moreover, we are the first to conduct observability analysis on rotational constraints in rotation-translation decoupling methods, experimentally identifying the unobservable state-space directions under three degenerate motions within our approach. We also perform extensive experiments to delineate practical parameter solution boundaries for our method, with both efforts substantially enhancing the overall practical applicability of decoupling-based methods. Extensive experiments on simulated and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in accuracy and robustness while maintaining computational efficiency. Furthermore, experiments verify that it significantly improves convergence in VIO systems.
Bo Xu 0022, Zewen Xu, Yijia He, Zhanpeng Ouyang, Hao Wei 0008, Yihong Wu 0002, Jiancheng Li, Hongdong Li
IEEE Trans. Robotics3
2025 PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding
abstract
Recently, 3D Gaussian Splatting (3DGS) has shown encouraging performance for open vocabulary scene understanding tasks. However, previous methods cannot distinguish 3D instance-level information, which usually predicts a heatmap between the scene feature and text query. In this paper, we propose PanoGS, a novel and effective 3D panoptic open vocabulary scene understanding approach. Technically, to learn accurate 3D language features that can scale to large indoor scenarios, we adopt the pyramid tri-plane to model the latent continuous parametric feature space and use a 3D feature decoder to regress the multi-view fused 2D feature cloud. Besides, we propose language-guided graph cuts that synergistically leverage reconstructed geometry and learned language cues to group 3D Gaussian primitives into a set of super-primitives. To obtain 3D consistent instance, we perform graph clustering based segmentation with SAM-guided edge affinity computation between different super-primitives. Extensive experiments on widely used datasets show better or more competitive performance on 3D panoptic open vocabulary scene understanding. Project page: https://zju3dv.github.io/panogs.
Hongjia Zhai, Zhenzhe Li, Xiaokun Pan, Yijia He, Guofeng Zhang 0001
CVPR5
2025 DOGE: An Extrinsic Orientation and Gyroscope Bias Estimation for Visual-Inertial Odometry Initialization
abstract
Most existing visual-inertial odometry (VIO) initialization methods rely on accurate pre-calibrated extrinsic parameters. However, during long-term use, irreversible structural deformation caused by temperature changes, mechanical squeezing, etc. will cause changes in extrinsic parameters, especially in the rotational part. Existing initialization methods that simultaneously estimate extrinsic parameters suffer from poor robustness, low precision, and long initialization latency due to the need for sufficient translational motion. To address these problems, we propose a novel VIO initialization method, which jointly considers extrinsic orientation and gyroscope bias within the normal epipolar constraints, achieving higher precision and better robustness without delayed rotational calibration. First, a rotation-only constraint is designed for extrinsic orientation and gyroscope bias estimation, which tightly couples gyroscope measurements and visual observations and can be solved in pure-rotation cases. Second, we propose a weighting strategy together with a failure detection strategy to enhance the precision and robustness of the estimator. Finally, we leverage Maximum A Posteriori to refine the results before enough translation parallax comes. Extensive experiments have demonstrated that our method outperforms the state-of-the-art methods in both accuracy and robustness while maintaining competitive efficiency.
Zewen Xu, Yijia He, Hao Wei 0008, Yihong Wu 0002
ICRA2
2025 Neuraloc: Visual Localization in Neural Implicit Map With Dual Complementary Features
abstract
Recently, neural radiance fields (NeRF) have gained significant attention in the field of visual localization. However, existing NeRF-based approaches either lack geometric constraints or require extensive storage for feature matching, limiting their practical applications. To address these challenges, we propose an efficient and novel visual localization approach based on the neural implicit map with complementary features. Specifically, to enforce geometric constraints and reduce storage requirements, we implicitly learn a 3D keypoint descriptor field, avoiding the need to explicitly store point-wise features. To further address the semantic ambiguity of descriptors, we introduce additional semantic contextual feature fields, which enhance the quality and reliability of 2D-3D correspondences. Besides, we propose descriptor similarity distribution alignment to minimize the domain gap between 2D and 3D feature spaces during matching. Finally, we construct the matching graph using both complementary descriptors and contextual features to establish accurate 2D3D correspondences for 6-DoF pose estimation. Compared with the recent NeRF-based approaches, our method achieves a$3 \times$faster training speed and a$45 \times$reduction in model storage. Extensive experiments on two widely used datasets demonstrate that our approach outperforms or is highly competitive with other state-of-the-art NeRF-based visual localization methods. Project page: https://zju3dv.github.io/neuraloc
Hongjia Zhai, Boming Zhao, Xiaokun Pan, Yijia He, Zhaopeng Cui, Hujun Bao, Guofeng Zhang 0001
ICRA5
2025 MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM
abstract
Recent advancements in 3D Gaussian Splatting (3DGS) have made a significant impact on rendering and reconstruction techniques. Current research predominantly focuses on improving rendering performance and reconstruction quality using high-performance desktop GPUs, largely overlooking applications for embedded platforms like micro air vehicles (MAVs). These devices, with their limited computational resources and memory, often face a trade-off between system performance and reconstruction quality. In this paper, we improve existing methods in terms of GPU memory usage while enhancing rendering quality. Specifically, to address redundant 3D Gaussian primitives in SLAM, we propose merging them in voxel space based on geometric similarity. This reduces GPU memory usage without impacting system runtime performance. Furthermore, rendering quality is improved by initializing 3D Gaussian primitives via Patch-Grid (PG) point sampling, enabling more accurate modeling of the entire scene. Quantitative and qualitative evaluations on publicly available datasets demonstrate the effectiveness of our improvements.
Yinlong Bai, Junkai Niu, Yijia He, Yi Zhou 0010
IROS6
2025 Normal Reorientation for Scene Consistency
abstract
With the remarkable progress of 3D scanning technique, the captured indoor scenes appear increasingly in last decade. Generating orientation-consistent normals for indoor point clouds is a fundamental and important task. The existing orientation rectification methods pay more attention to object-level targets with connected surface. However, it is challenging to compute consistent surface orientation for real scanned indoor point clouds. In this paper, we analyze the causes of this difficulty and propose a new normal reorienting framework for indoor scene consistency, namely NRSC. It first estimates normals for an indoor point cloud and extracts all the connected regions. We then design and construct an abstract orientation bridging tree (OBT) to organize the extracted regions in a hierarchical way. For all node regions, NRSC iteratively implements a set of orientation propagations to generate locally orientation-consistent regions. Moreover, we define an auxiliary viewpoint set for each pairwise parent-child node regions and introduce a voting mechanism to rectify the region orientation of child node according to its parent. After processing all the child node regions along OBT, we finally eliminate the orientation inconsistencies between related regions. Multi-groups of experimental results on both fused indoor scenes and single-view-scenes show that our method generates globally consistent orientation for indoor point clouds.
Long Yang 0001, Yijia He, Shaojun Hu, Chunxia Xiao, Zhiyi Zhang 0002
IEEE Trans. Vis. Comput. Graph.4
2025 SplatLoc: 3D Gaussian Splatting-based Visual Localization for Augmented Reality
abstract
Visual localization plays an important role in the applications of Augmented Reality (AR), which enable AR devices to obtain their 6-DoF pose in the pre-build map in order to render virtual content in real scenes. However, most existing approaches can not perform novel view rendering and require large storage capacities for maps. To overcome these limitations, we propose an efficient visual localization method capable of high-quality rendering with fewer parameters. Specifically, our approach leverages 3D Gaussian primitives as the scene representation. To ensure precise 2D-3D correspondences for pose estimation, we develop an unbiased 3D scene-specific descriptor decoder for Gaussian primitives, distilled from a constructed feature volume. Additionally, we introduce a salient 3D landmark selection algorithm that selects a suitable primitive subset based on the saliency score for localization. We further regularize key Gaussian primitives to prevent anisotropic effects, which also improves localization performance. Extensive experiments on two widely used datasets demonstrate that our method achieves superior or comparable rendering and localization performance to state-of-the-art implicit-based visual localization approaches. Code and data are available at project page: https://zju3dv.github.io/splatloc.
Hongjia Zhai, Xiyu Zhang 0003, Boming Zhao, Yijia He, Zhaopeng Cui, Hujun Bao, Guofeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose Tracking
abstract
Ensuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (LED). This novel controller design allows users to free up their hands for more immersive experiences. To track the controller’s motion accurately and robustly, we resort to various forms of visual measurements, including 6 DoF and 5 DoF pose measurements from hand gesture detection, as well as 3 DoF position measurement and 2 DoF image measurement derived from the LED. We theoretically analyze the performances of these observation models and propose an optimal observation model combination scheme. Moreover, the necessity and rationale of online estimating system gravity are illustrated. The effectiveness of our tracking method is validated through extensive experiments.
Zhuqing Zhang, Dongxuan Li, Yijia He, Pan Ji, Rong Xiong, Hongdong Li, Yue Wang 0020
ICRA4
2023 A Rotation-Translation-Decoupled Solution for Robust and Efficient Visual-Inertial Initialization
abstract
We propose a novel visual-inertial odometry (VIO) initialization method, which decouples rotation and translation estimation, and achieves higher efficiency and better robustness. Existing loosely-coupled VIO-initialization methods suffer from poor stability of visual structure-from-motion (SfM), whereas those tightly-coupled methods often ignore the gyroscope bias in the closed-form solution, resulting in limited accuracy. Moreover, the aforementioned two classes of methods are computationally expensive, because 3D point clouds need to be reconstructed simultaneously. In contrast, our new method fully combines inertial and visual measurements for both rotational and translational initialization. First, a rotation-only solution is designed for gyroscope bias estimation, which tightly couples the gyroscope and camera observations. Second, the initial velocity and gravity vector are solved with linear translation constraints in a globally optimal fashion and without reconstructing 3D point clouds. Extensive experiments have demonstrated that our method is 8 ~ 72 times faster (w.r.t. a 10-frame set) than the state-of-the-art methods, and also presents significantly higher robustness and accuracy. The source code is available at https://github.com/boxuLibrary/drt-vio-init.
Yijia He, Bo Xu 0022, Zhanpeng Ouyang, Hongdong Li
CVPR1
2022 MVSTER: Epipolar Transformer for Efficient Multi-view Stereo
Guan Huang 0003, Fangbo Qin, Yijia He, Xu Chi, Xingang Wang 0003
ECCV (31)6
2021 ELSD: Efficient Line Segment Detector and Descriptor
abstract
We present the novel Efficient Line Segment Detector and Descriptor (ELSD) to simultaneously detect line segments and extract their descriptors in an image. Unlike the traditional pipelines that conduct detection and description separately, ELSD utilizes a shared feature extractor for both detection and description, to provide the essential line features to the higher-level tasks like SLAM and image matching in real time. First, we design a one-stage compact model, and propose to use the mid-point, angle and length as the minimal representation of line segment, which also guarantees the center-symmetry. The non-centerness suppression is proposed to filter out the fragmented line segments caused by lines’ intersections. The fine offset prediction is designed to refine the mid-point localization. Second, the line descriptor branch is integrated with the detector branch, and the two branches are jointly trained in an end-to-end manner. In the experiments, the proposed ELSD achieves the state-of-the-art performance on the Wireframe dataset and YorkUrban dataset, in both accuracy and efficiency. The line description ability of ELSD also outperforms the previous works on the line matching task.
Yicheng Luo, Fangbo Qin, Yijia He, Xiao Liu 0042
ICCV4
2021 M3VSNET: Unsupervised Multi-Metric Multi-View Stereo Network
abstract
The present Multi-view stereo (MVS) methods with supervised learning-based networks have an impressive performance comparing with traditional MVS methods. However, the ground-truth depth maps for training are hard to be obtained and are within limited kinds of scenarios. In this paper, we propose a novel unsupervised multi-metric MVS network, named M3VSNet, for dense point cloud reconstruction without any supervision. To improve the robustness and completeness of point cloud reconstruction, we propose a novel multi-metric loss function that combines pixel-wise and feature-wise loss function to learn the inherent constraints from different perspectives of matching correspondences. Besides, we also incorporate the normal-depth consistency in the 3D point cloud format to improve the accuracy and continuity of the estimated depth maps. Experimental results show that M3VSNet establishes the state-of-the-arts unsupervised method and achieves better performance than previous supervised MVSNet on the DTU dataset and demonstrates the powerful generalization ability on the Tanks & Temples benchmark with effective improvement.
Baichuan Huang, Hongwei Yi, Yijia He, Jingbin Liu, Xiao Liu 0042
ICIP4
2021 Structure Reconstruction Using Ray-Point-Ray Features: Representation and Camera Pose Estimation
abstract
Straight line features have been increasingly utilized in visual SLAM and 3D reconstruction systems. The straight lines’ parameterization, parallel constraint, and coplanar constraint are studied in many recent works. In this paper, we explore the novel intersection constraint of straight lines for structure reconstruction. First, a minimum parameterized representation of ray-point-ray (RPR) structures is proposed to represent the intersection of two straight lines in the 3D space. Second, an efficient solver is designed for the camera pose estimation, which leverages the perpendicularity and intersection of straight lines. Third, we build a stereo visual odometry based on RPR features and evaluate it on the simulation and real datasets. The experimental results verify that the intersection constraints from RPR can effectively improve the accuracy and efficiency of line-based SLAM and reconstruction system.
Yijia He, Xiao Liu 0042, Ji Zhao 0001
ICRA1
2020 TP-LSD: Tri-Points Based Line Segment Detector
Siyu Huang, Fangbo Qin, Pengfei Xiong, Yijia He, Xiao Liu 0042
ECCV (27)5
2020 Leveraging Planar Regularities for Point Line Visual-Inertial Odometry
abstract
With monocular Visual-Inertial Odometry (VIO) system, 3D point cloud and camera motion can be estimated simultaneously. Because pure sparse 3D points provide a structureless representation of the environment, generating 3D mesh from sparse points can further model the environment topology and produce dense mapping. To improve the accuracy of 3D mesh generation and localization, we propose a tightly-coupled monocular VIO system, PLP-VIO, which exploits point features and line features as well as plane regularities. The co-planarity constraints are used to leverage additional structure information for the more accurate estimation of 3D points and spatial lines in state estimator. To detect plane and 3D mesh robustly, we combine both the line features with point features in the detection method. The effectiveness of the proposed method is verified on both synthetic data and public datasets and is compared with other state-of-the-art algorithms.
Xin Li 0084, Yijia He, Jinlong Lin, Xiao Liu 0042
IROS2
2020 Minimal Case Relative Pose Computation Using Ray-Point-Ray Features
abstract
Corners are popular features for relative pose computation with 2D-2D point correspondences. Stable corners may be formed by two 3D rays sharing a common starting point. We call such elements ray-point-ray (RPR) structures. Besides a local invariant keypoint given by the lines' intersection, their reprojection also defines a corner orientation and an inscribed angle in the image plane. The present paper investigates such RPR features, and aims at answering the fundamental question of what additional constraints can be formed from correspondences between RPR features in two views. In particular, we show that knowing the value of the inscribed angle between the two 3D rays poses additional constraints on the relative orientation. Using the latter enables the solution of the relative pose problem with as few as 3 correspondences across the two images. We provide a detailed analysis of all minimal cases distinguishing between 90-degree RPR-structures and structures with an arbitrary, known inscribed angle. We furthermore investigate the special cases of a known directional correspondence and planar motion, the latter being solvable with only a single RPR correspondence. We complete the exposition by outlining an image processing technique for robust RPR-feature extraction. Our results suggest high practicality in man-made environments, where 90-degree RPR-structures naturally occur.
Ji Zhao 0001, Laurent Kneip, Yijia He, Jiayi Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Correlational examples for convolutional neural networks to detect small impurities
Yue Guo 0010, Yijia He, Haitao Song 0004, Kui Yuan
Neurocomputing2