Yanhao Zhang 0003

dblp:84/10486-3 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-2722-8019ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Self-supervised 3D Reconstruction of Tibia and Fibula from Biplanar X-rays
abstract
With the growing number of patients experiencing knee-related conditions, total knee arthroplasty (TKA) has become a common procedure, where a 3D visualisation of the patient’s tibia and fibula is essential for preoperative planning. Traditional imaging techniques, such as computed tomography (CT), often expose patients to high levels of radiation or impose significant financial costs. As an alternative, this paper proposes a novel approach that reconstructs a 3D model of the tibia and fibula using only two X-ray images (taken from the coronal and sagittal planes) and a general template, significantly reducing radiation exposure and financial burden. Our algorithm of 3D reconstruction for patient-specific anatomies combines point-based deformation with deep learning techniques. Initially, the general model undergoes a preliminary deformation to match the patient tibia and fibula dimensions. This pre-deformed model then serves as a template, followed by a fine deformation process via a self-supervised graph convolutional network (GCN), whose parameters are trained iteratively by comparing the template projection and the X-ray measurements. Following tests in simulations, cadaver experiments, and in-vivo experiments, our proposed algorithm demonstrates state-of-the-art accuracy and exceptional robustness across different evaluation metrics. Our code is available at https://github.com/DrKaiPan/tfDeform_GCN.git
Kai Pan, Yanhao Zhang 0003, Liang Zhao 0003, Shoudong Huang
IROS2
2025 Nonrigid Structure-From-Motion via Differential Geometry With Recoverable Conformal Scale
abstract
Non-rigid structure-from-motion (NRSfM), a promising technique for addressing the mapping challenges in monocular visual deformable simultaneous localization and mapping (SLAM), has attracted growing attention. We introduce a novel method, called Con-NRSfM, for NRSfM under conformal deformations, encompassing isometric deformations as a subset. Our approach performs point-wise reconstruction using 2D selected image warps optimized through a graph-based framework. Unlike existing methods that rely on strict assumptions, such as locally planar surfaces or locally linear deformations, and fail to recover the conformal scale, our method eliminates these constraints and accurately computes the local conformal scale. Additionally, our framework decouples constraints on depth and conformal scale, which are inseparable in other approaches, enabling more precise depth estimation. To address the sensitivity of the formulated problem, we employ a parallel separable iterative optimization strategy. Furthermore, a self-supervised learning framework, utilizing an encoder-decoder network, is incorporated to generate dense 3D point clouds with texture. Simulation and experimental results using both synthetic and real datasets demonstrate that our method surpasses existing approaches in terms of reconstruction accuracy and robustness. The code for the proposed method will be made publicly available on the project website:https://sites.google.com/view/con-nrsfm.
Yongbo Chen 0001, Yanhao Zhang 0003, Shaifali Parashar, Liang Zhao 0003, Shoudong Huang
IEEE Trans. Robotics2
2024 View from Above: Orthogonal-View Aware Cross-View Localization
abstract
This paper presents a novel aerial-to-ground feature ag-gregation strategy, tailored for the task of cross- view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database, often through establishing planar homography correspondences via a planar ground assumption. As such, they tend to ignore features that are off-ground and not suited for handling visual occlusions, leading to unreliable localization in challenging scenarios. We propose a Top-to-Ground Aggregation (T2GA) module that capitalizes aerial orthographic views to aggregate features down to the ground level, leveraging reliable off-ground information to improve feature alignment. Furthermore, we introduce a Cycle Domain Adaptation (CycDA) loss that ensures feature extraction robustness across do-main changes. Additionally, an Equidistant Re-projection (ERP) loss is introduced to equalize the impact of all key-points on orientation error, leading to a more extended distribution of keypoints which benefits orientation estimation. On both KITTI and Ford Multi-AV datasets, our method consistently achieves the lowest mean longitudinal and lateral translations across different settings and obtains the smallest orientation error when the initial pose is less ac-curate, a more challenging setting. Further, it can complete an entire route through continual vehicle pose estimation with initial vehicle pose given only at the starting point.11Code is available at https://github.com/ShanWang-Shan/View FromAbove.
Shan Wang 0010, Jiawei Liu 0005, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Kaihao Zhang, Hongdong Li
CVPR4
2024 Increasing SLAM Pose Accuracy by Ground-to-Satellite Image Registration
abstract
Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to error accumulation, visual SLAM (vSLAM) usually suffers from long-term drift. This paper proposes a framework to increase the localization accuracy by fusing the vSLAM with a deep-learning based ground-to-satellite (G2S) image registration method. In this framework, a coarse (spatial correlation bound check) to fine (visual odometry consistency check) method is designed to select the valid G2S prediction. The selected prediction is then fused with the SLAM measurement by solving a scaled pose graph problem. To further increase the localization accuracy, we provide an iterative trajectory fusion pipeline. The proposed framework is evaluated on two well-known autonomous driving datasets, and the results demonstrate the accuracy and robustness in terms of vehicle localization. The code will be available at https://github.com/YanhaoZhang/SLAM-G2S-Fusion.
Yanhao Zhang 0003, Yujiao Shi 0002, Shan Wang 0010, Ankit Vora, Akhil Perincherry, Yongbo Chen 0001, Hongdong Li
ICRA1
2024 R2-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstruction
abstract
3D Gaussian splatting (3DGS) has shown promising results in image rendering and surface reconstruction. However, its potential in volumetric reconstruction tasks, such as X-ray computed tomography, remains under-explored. This paper introduces R$^2$-Gaussian, the first 3DGS-based framework for sparse-view tomographic reconstruction. By carefully deriving X-ray rasterization functions, we discover a previously unknown \emph{integration bias} in the standard 3DGS formulation, which hampers accurate volume retrieval. To address this issue, we propose a novel rectification technique via refactoring the projection from 3D to 2D Gaussians. Our new method presents three key innovations: (1) introducing tailored Gaussian kernels, (2) extending rasterization to X-ray imaging, and (3) developing a CUDA-based differentiable voxelizer. Experiments on synthetic and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in accuracy and efficiency. Crucially, it delivers high-quality results in 4 minutes, which is 12$\times$ faster than NeRF-based methods and on par with traditional algorithms.
Ruyi Zha, Yuanhao Cai, Jiwen Cao, Yanhao Zhang 0003, Hongdong Li
NeurIPS5
2023 Homography Guided Temporal Fusion for Road Line and Marking Segmentation
abstract
Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance and overall high appearance consistency. To solve these issues, we propose a Homography Guided Fusion (HomoFusion) module to exploit temporally-adjacent video frames for complementary cues facilitating the correct classification of the partially occluded road lines or markings. To reduce computational complexity, a novel surface normal estimator is proposed to establish spatial correspondences between the sampled frames, allowing the HomoFusion module to perform a pixel-to-pixel attention mechanism in updating the representation of the occluded road lines or markings. Experiments on ApolloScape, a large-scale lane mark segmentation dataset, and ApolloScape Night with artificial simulated night-time road conditions, demonstrate that our method outperforms other existing SOTA lane mark segmentation models with less than 9% of their parameters and computational complexity. We show that exploiting available camera intrinsic data and ground plane assumption for cross-frame correspondence can lead to a light-weight network with significantly improved performances in speed and accuracy. We also prove the versatility of our HomoFusion approach by applying it to the problem of water puddle segmentation and achieving SOTA performance1.
Shan Wang 0010, Jiawei Liu 0005, Kaihao Zhang, Wenhan Luo, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Hongdong Li
ICCV6
2023 View Consistent Purification for Accurate Cross-View Localization
abstract
This paper proposes a fine-grained self-localization method for outdoor robotics that utilizes a flexible number of onboard cameras and readily accessible satellite images. The proposed method addresses limitations in existing cross-view localization methods that struggle to handle noise sources such as moving objects and seasonal variations. It is the first sparse visual-only method that enhances perception in dynamic environments by detecting view-consistent key points and their corresponding deep features from ground and satellite views, while removing off-the-ground objects and establishing homography transformation between the two views. Moreover, the proposed method incorporates a spatial embedding approach that leverages camera intrinsic and extrinsic information to reduce the ambiguity of purely visual matching, leading to improved feature matching and overall pose estimation accuracy. The method exhibits strong generalization and is robust to environmental changes, requiring only geo-poses as ground truth. Extensive experiments on the KITTI and Ford Multi-AV Seasonal datasets demonstrate that our proposed method outperforms existing state-of-the-art methods, achieving median spatial accuracy errors below 0.5 meters along the lateral and longitudinal directions, and a median orientation accuracy error below 2°1.
Shan Wang 0010, Yanhao Zhang 0003, Akhil Perincherry, Ankit Vora, Hongdong Li
ICCV2
2023 3D Reconstruction of Tibia and Fibula using One General Model and Two X-ray Images
abstract
The 3D reconstruction of patient specific bone models plays a crucial role in orthopaedic surgery for clinical evaluation, surgical planning and precise implant design or selection. This paper considers the problem of reconstructing a patient-specific 3D tibia and fibula model from only two 2D X-ray images and one 3D general model segmented from the lower leg CT scans of one randomly selected patient. Currently, the bone 3D reconstruction mainly relies on computed tomography (CT) and magnetic resonance imaging (MRI) scanning-based mode segmentation which result in high radiation exposure or expensive costs. While, the proposed algorithm can accurately and efficiently deform a 3D general model to achieve a patient-specific 3D model that matches the patient's tibia and fibula projections in two 2D X-rays. The algorithm undergoes a preliminary deformation, 2D contour registration, and opti-misation based on the deformation graph that represents the shape deformation of models. Evaluations using simulations, cadaver and in-vivo experiments demonstrate that the proposed algorithm can effectively reconstruct the patient's 3D tibia and fibula surface model with high accuracy.
Kai Pan, Shuai Zhang 0029, Liang Zhao 0003, Shoudong Huang, Yanhao Zhang 0003
ICRA5
2023 Satellite Image Based Cross-view Localization for Autonomous Vehicle
abstract
Existing spatial localization techniques for au-tonomous vehicles mostly use a pre-built 3D-HD map, often constructed using a survey-grade 3D mapping vehicle, which is not only expensive but also laborious. This paper shows that by using an off-the-shelf high-definition satellite image as a ready-to-use map, we are able to achieve cross-view vehicle localization up to a satisfactory accuracy, providing a cheaper and more practical way for localization. While the utilization of satellite imagery for cross-view localization is an established concept, the conventional methodology focuses primarily on image re-trieval. This paper introduces a novel approach to cross-view localization that departs from the conventional image retrieval method. Specifically, our method develops (1) a Geometric-align Feature Extractor (GaFE) that leverages measured 3D points to bridge the geometric gap between ground and overhead views, (2) a Pose Aware Branch (PAB) adopting a triplet loss to encourage pose-aware feature extraction, and (3) a Recursive Pose Refine Branch (RPRB) using the Levenberg-Marquardt (LM) algorithm to align the initial pose towards the true vehicle pose iteratively. Our method is validated on KITTI and Ford Multi-AV Seasonal datasets as ground view and Google Maps as the satellite view. The results demonstrate the superiority of our method in cross-view localization with median spatial and angular errors within 1 meter and 1°, respectively.
Shan Wang 0010, Yanhao Zhang 0003, Ankit Vora, Akhil Perincherry, Hengdong Li
ICRA2
2023 Structure-to-Shape Aortic 3-D Deformation Reconstruction for Endovascular Interventions
abstract
Fluoroscopy-guided endovascular interventions by using X-ray images are challenging. The catheter needs to be manipulated precisely inside the aorta, while only 2-D views from the X-ray fluoroscopy are currently used to help the surgeons. Because the catheter is operated in a 3-D space, a visualization of the deforming 3-D aorta will be useful as guidance for catheter manipulation. Existing 3-D reconstruction methods fall short in only focusing on the deformation reconstruction of the aortic 3-D centerline, or using additional prior knowledge of 3-D catheter position for estimating the aortic 3-D deformation. In this article, we propose a novel framework that reconstructs the aortic 3-D deformation by fusing a preoperative 3-D model and two intraoperative X-ray images. Different from existing methods, the proposed framework reconstructs aortic deformation using a coarse-to-fine pipeline by first reconstructing the aortic 3-D centerline and then reconstructing the 3-D shape. To obtain the accurate features for the fluoroscopic-based 3-D reconstruction, we extract semantic features from the X-ray images, and compute the distance field to efficiently calculate the 3-D–2-D nonrigid correspondence. Nonlinear least squares optimization is used to solve the deformation of both centerline and shape. The proposed framework is validated using phantom and patient datasets, whose results demonstrate improved efficiency and accuracy compared with the existing methods. This framework provides a valuable clinical tool for endovascular interventions.
Yanhao Zhang 0003, Raphael Falque, Liang Zhao 0003, Yongbo Chen 0001, Shoudong Huang, Hongdong Li
IEEE Trans. Robotics1
2022 NAF: Neural Attenuation Fields for Sparse-View CBCT Reconstruction
Ruyi Zha, Yanhao Zhang 0003, Hongdong Li
MICCAI (6)2
2022 Anchor Selection for SLAM Based on Graph Topology and Submodular Optimization
abstract
This article considers simultaneous localization and mapping (SLAM) problem for robots in situations where accurate estimates for some of the robot poses, termed anchors, are available. These may be acquired through external means, for example, by either stopping the robot at some previously known locations or pausing for a sufficient period of time to measure the robot poses with an external measurement system. The main contribution is an efficient algorithm for selecting a fixed number of anchors from a set of potential poses that minimizes estimated error in the SLAM solution. Based on a graph-topological connection between the D-optimality design metric and the tree-connectivity of the pose-graph, the anchor selection problem can be formulated approximately as a submatrix selection problem for reduced weighted Laplacian matrix, leading to a cardinality-constrained submodular maximization problem. Two greedy methods are presented to solve this submodular optimization problem with a performance guarantee. These methods are complemented by Cholesky decomposition, approximate minimum degree permutation, order reuse, and rank-1 update that exploit the sparseness of the weighted Laplacian matrix. We demonstrate the efficiency and effectiveness of the proposed techniques on public-domain datasets, Gazebo simulations, and real-world experiments.
Yongbo Chen 0001, Liang Zhao 0003, Yanhao Zhang 0003, Shoudong Huang, Gamini Dissanayake
IEEE Trans. Robotics3
2021 Some Research Questions for SLAM in Deformable Environments
abstract
SLAM in deformable environments is a very challenging research topic. Some research works have been presented by different research groups in the past few years. However, there are still some challenging research questions remaining unanswered. This paper discusses some of these research questions focusing on the case when point features are used to describe the deformable environments. The SLAM problems are formulated as extensions of point feature based SLAM in static environments, including both optimisation based offline SLAM and filter based online SLAM. To illustrate the problems and questions more clearly, some concepts and results using simple 2D examples are presented. The MATLAB source codes of the results are made publicly available (https://github.com/cyb1212/DeformableSLAM2D.git) to help the readers understand the problems more clearly.
Shoudong Huang, Yongbo Chen 0001, Liang Zhao 0003, Yanhao Zhang 0003, Mengya Xu
IROS4
2020 Aortic 3D Deformation Reconstruction using 2D X-ray Fluoroscopy and 3D Pre-operative Data for Endovascular Interventions
abstract
Current clinical endovascular interventions rely on 2D guidance for catheter manipulation. Although an aortic 3D surface is available from the pre-operative CT/MRI imaging, it cannot be used directly as a 3D intra-operative guidance since the vessel will deform during the procedure. This paper aims to reconstruct the live 3D aortic deformation by fusing the static 3D model from the pre-operative data and the 2D live imaging from fluoroscopy. In contrast to some existing deformation reconstruction frameworks which require 3D observations such as RGB-D or stereo images, fluoroscopy only presents 2D information. In the proposed framework, a 2D-3D registration is performed and the reconstruction process is formulated as a non-linear optimization problem based on the deformation graph approach. Detailed simulations and phantom experiments are conducted and the result demonstrates the reconstruction accuracy and robustness, as well as the potential clinical value of this framework.
Yanhao Zhang 0003, Liang Zhao 0003, Shoudong Huang
ICRA1
2020 Deep Learning Assisted Automatic Intra-operative 3D Aortic Deformation Reconstruction
Yanhao Zhang 0003, Raphael Falque, Liang Zhao 0003, Shoudong Huang, Boni Hu
MICCAI (4)1