Zhaoshuai Qi

dblp:308/5792 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-3013-5099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AxisPose: Model-Free Matching-Free Single-Shot 6D Object Pose Estimation via Axis Generation
abstract
Object pose estimation is a fundamental task in computer vision and plays an important role in various applications such as robotics, augmented reality, and autonomous manipulation. Existing studies often demand complex inputs or depend on correspondence-based matching between 2D image features and 3D object representations. While effective, these methods rely strongly on explicit appearance matching, often requiring multi-view inputs, depth sensors, or CAD models, which limits their scalability and robustness. Building on top of the pioneering generative studies, we propose AxisPose, a model-free, matching-free, and single-view 6D pose estimation framework that departs from conventional correspondence-based paradigms. Unlike existing methods, AxisPose directly infers a pose representation by learning a latent distribution of object orientation axes through a diffusion model. Specifically, AxisPose introduces an Axis Generation Module (AGM) that progressively denoises tri-axial orientation fields guided by geometric consistency constraints, and a Triaxial Back-projection Module (TBM) to recover the final 6D pose from the generated orientation axes without relying on explicit 2D- 2D/3D correspondences. AxisPose achieves strong cross-instance generalization, enabling a single model to handle multiple object categories without retraining. Extensive experiments on LINEMOD and YCB-Video datasets demonstrate that AxisPose improves the Average Distance Deviation score from 0.733 to 0.814 over the strong baseline NOPE, using only a single RGB input. The code is available at https://github.com/pubyLu/AxisPose/tree/main.
Yang Zou 0004, Zhaoshuai Qi, Weipeng Sun, Xingyuan Li 0005, Jiaqi Yang 0002, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 SPU-IMR: Self-supervised Arbitrary-scale Point Cloud Upsampling via Iterative Mask-recovery Network
abstract
Point cloud upsampling aims to generate dense and uniformly distributed point sets from sparse point clouds. Existing point cloud upsampling methods typically approach the task as an interpolation problem. They achieve upsampling by performing local interpolation between point clouds or in the feature space, then regressing the interpolated points to appropriate positions. By contrast, our proposed method treats point cloud upsampling as a global shape completion problem. Specifically, our method first divides the point cloud into multiple patches. Then a masking operation is applied to remove some patches, leaving visible point cloud patches. Finally, our custom-designed neural network iterative completes the missing sections of the point cloud through the visible parts. During testing, by selecting different mask sequences, we can restore various complete patches. A sufficiently dense upsampled point cloud can be obtained by merging all the completed patches. We demonstrate the superior performance of our method through both quantitative and qualitative experiments, showing overall superiority against both existing self-supervised and supervised methods.
Ziming Nie, Qiao Wu, Chenlei Lv, Siwen Quan, Zhaoshuai Qi, Muze Wang, Jiaqi Yang 0002
AAAI5
2025 HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D Registration
abstract
Geometric constraints between feature matches are critical in 3D point cloud registration problems. Existing approaches typically model unordered matches as a consistency graph and sample consistent matches to generate hypotheses. However, explicit graph construction introduces noise, posing great challenges for handcrafted geometric constraints to render consistency. To overcome this, we propose HyperGCT, a flexible dynamic Hyper-GNN-learned geometric ConstrainT that leverages high-order consistency among 3D correspondences. To our knowledge, HyperGCT is the first method that mines robust geometric constraints from dynamic hypergraphs for 3D registration. By dynamically optimizing the hypergraph through vertex and edge feature aggregation, HyperGCT effectively captures the correlations among correspondences, leading to accurate hypothesis generation. Extensive experiments on 3DMatch, 3DLoMatch, KITTI-LC, and ETH show that HyperGCT achieves state-of-the-art performance. Furthermore, HyperGCT is robust to graph noise, demonstrating a significant advantage in terms of generalization.
Xiyu Zhang 0001, Jiayi Ma 0001, Zhaoshuai Qi, Fei Hui, Jiaqi Yang 0002, Yanning Zhang 0001
ICCV5
2025 Matching quality-guided model-free satellite pose estimation
Zhaoshuai Qi
Eng. Appl. Artif. Intell.1
2025 TAP-Track: Generalizable Spacecraft Pose Tracking by Tracking Any Points
abstract
Recent learning-based spacecraft pose tracking methods have demonstrated impressive improvement in estimation accuracy and potential scalability to complex space environment. However, most of them still rely on the detection of discriminative keypoints on a known 3D model, limiting the generalization to unknown spacecraft. To this end, we propose, to the best of our knowledge, the first generalizable spacecraft pose tracking method. Instead of requiring a known model, we only assume the existence of at least one planar structure, e.g. solar panels, which holds for most satellites in general scenes. Additionally, the proposed method tracks any points on the plane across multi-frame followed by a re-projection error minimization, rather than detecting keypoints between image pairs, allowing robust capture of “long-term” temporal information among frames even for textureless surfaces without sufficient keypoints. Moreover, we also constructed the first large-scale dataset G-SPET for generalizable spacecraft pose estimation and tracking. It covers 174 satellites with diversity structures and rich annotations, increasing the number of targets in previous datasets by almost two orders of magnitude. Extensive evaluations on the proposed dataset have demonstrated the superiority of our method over state-of-the-art methods. The code and dataset will be made publicly available soon.
Zhaoshuai Qi, Pulin Chen, Huilin Fan, Yu Zhu 0004, Jiaqi Yang 0002, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Edge-Guided Detector-Free Network for Robust and Accurate Visible-Thermal Image Matching
abstract
Recent detector-free models strive to leverage both local and global context for image matching, showcasing enhanced robustness, particularly in scenarios with weak-textured scenes. Despite these advancements, automatically establishing feature correspondences between visible and thermal images still introduces additional challenges. Differences in radiation and geometry between these modalities often result in degraded performance for the majority of existing methods. To this end, we propose edge-guided detector-free model termed EDMatcher for visible-thermal image matching. Besides local and global context in the images, EDMatcher also leverages modality-robust structural information in image edges, which demonstrates promising robustness to images with distinct modalities. Moreover, an edge-masked ground-truth matrix generation strategy is introduced during the training, which helps EDMatcher to further focus on more salient regions while leaving out texture-less regions, leading to more efficient learning. Extensive experiments show that EDMatcher has strong generalization and achieves excellent matching performances.
Zhaoshuai Qi, Xiuwei Zhang 0001, Tao Zhuo, Yanning Zhang 0001
ICME2
2024 A Minimal Solution for Sphere-Based Camera-Projector Pair Calibration
abstract
We propose a minimal solution for sphere-based camera-projector pair (CPP) calibration. Previous works often treated the camera and projector calibration as two independent problems, which exploit only intra-view information from geometric properties of sphere dual image formation and hence require at least three spheres for CPP calibration. However, other than intra-view information, we observe that inter-view information between camera and projector provides additional constraints. Combining these two kinds of information yields a minimal solution for CPP calibration, where only a single sphere is required. Extensive experiments have verified the effectiveness of proposed minimal solver, which demonstrates higher flexibility and comparable accuracy to the state-of-the-art methods. Moreover, the achieved flexibility allows high-quality 3D reconstruction with an uncalibrated CPP, given only a single sphere in the scene.
Zhaoshuai Qi, Jingqi Pang, Yifeng Hao, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Toward Meta-Shape-Based Multi-View 3D Point Cloud Registration: An Evaluation
abstract
Reducing cumulative registration error is critical to accurate 3D multi-view registration. Meta-shape based methods optimize rigid transformations of point clouds by iteratively registering each point cloud with a meta-shape, which remain popular solutions to 3D multi-view registration. However, the merits and demerits of existing meta-shape based methods remain unclear. Moreover, we argue that simpler meta-shape based solutions can achieve even better performance. To this end, we evaluate seven representative meta-shape based methods in this work, including four existing ones and three modified ones, in order to investigate the problem of defining a good meta-shape. In particular, we first abstract the main steps of considered methods. Then, experiments on both object and scene datasets with real and synthetic cumulative registration errors are deployed for an in-depth evaluation. Finally, based on the experimental outcomes, we give a discussion on the advantages and limitations of meta-shape based methods. We demonstrate prior works have used unnecessarily complicated techniques for cumulative error elimination and our slightly modified simpler solutions can achieve competitive performance on experimental datasets.
Shikun Zhang, Jiaqi Yang 0002, Zhaoshuai Qi, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 SAC-COT: Sample Consensus by Sampling Compatibility Triangles in Graphs for 3-D Point Cloud Registration
abstract
Six-degree-of-freedom (6-DOF) pose estimation from feature correspondences remains a popular and robust approach for 3-D registration. However, heavy outliers that existed in the initial correspondence set pose a great challenge to this problem. This article presents a simple yet effective estimator called SAmple Consensus by sampling COmpatibility Triangles in graphs (SAC-COT) for robust 6-DOF pose estimation and 3-D registration. The key novelty is a guided three-point sampling approach. It is based on a novel correspondence sample representation, i.e., COmpatibility Triangle (COT). We first model the correspondence set as a graph with nodes connecting compatible correspondences. Then, by ranking and sampling COTs formed by ternary loops, we show that correct hypotheses can be generated in early iteration stage. Finally, the hypothesis generated by the COT yielding to the maximum consensus is the output of SAC-COT. Extensive experiments on six data sets and extensive comparisons with the state-of-the-art estimators confirm that: 1) SAC-COT can achieve accurate registrations with a few iterations and 2) SAC-COT outperforms all competitors and is ultrarobust when confronted with Gaussian noise, data decimation, holes, clutter, partial overlap, varying scales of input correspondences, and data modality variation.
Jiaqi Yang 0002, Siwen Quan, Zhaoshuai Qi, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4