EDBT 2026 Demo / reviewers in the wild / expert
Jiaqi Yang 0002
dblp:131/7234-2
· DBLP profile ↗
57ranked-venue papers
20as first author
40since 2021 · last 2026
0000-0002-2071-2457ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 8 first-author · 25 since 2021Artificial intelligence and machine learning · 32 · 9 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFF-Matcher: Robust cross-source registration with density-fused feature and bidirectional consensus matching
Zhenxuan Zeng, Xiyu Zhang 0001, Siwen Quan, Zhongwen Hu, Yu Zhu 0004, Jiaqi Yang 0002 |
J. Vis. Commun. Image Represent. | 8 |
| 2026 | Rethinking the refinement stage of 3D object detection: A multi-task learning perspective with Mixture-of-Experts
Bingqian Wu, Pei An, Siwen Quan, Qiao Wu, Chu'ai Zhang, Jiaqi Yang 0002 |
J. Vis. Commun. Image Represent. | 7 |
| 2026 | Single Voter Spreading for Efficient Correspondence Grouping and 3D RegistrationabstractObtaining highly consistent correspondences between point clouds is crucial for computer vision tasks such as 3D registration and recognition. Due to nuisances such as limited overlap and noise, initial correspondences often contain a large number of outliers, imposing a great challenge to downstream tasks. In this paper, we present a novel single voter spreading (SVOS) method for efficient 3D correspondence grouping and 3D registration. Our core insight is to leverage low-order graph constraints only in a single voter spreading voting scheme to achieve comparable constrain-ability as complex constraints without searching them. First, a simple first-order graph is constructed for the initial correspondence set. Second, a two-stage voting method is proposed, including single voter voting and spread voters voting. Each voting stage involves both local and global voting via edge constraints only. This promises good selectivity while making the voting process time- and storage-efficient. Finally, top-scored correspondences are opted for robust transformation estimation. Experiments on U3M, 3DMatch/3DLoMatch, ETH, and KITTI-LC datasets verify that SVOS achieves new state-of-the-art correspondence grouping and registration performance, while being light-weight and robust to graph construction parameters. Siwen Quan, Zhao Zeng, Xiyu Zhang 0001, Jiaqi Yang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | A Hierarchical Prior Mining Approach for Non-Local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of the target from a set of 2D images. However, the reconstructed quality is significantly impacted by the presence of low-textured areas. In this paper, we propose a Hierarchical Prior Mining (HPM) framework for non-local multi-view stereo. Different from most existing works dedicated to focusing on local information and only using a single prior, HPM captures non-local structural cues and leverages multi-source priors for geometry recovery. Based on the framework, we first propose HPM-MVS, which obtains precise initial hypotheses through non-local operations, simultaneously constructing a better planar prior model in an HPM framework to further facilitate hypothesis generation. In addition, we futher propose HPM-MVS++, which excavates the structured region information of images and spatial geometric relationships of hypotheses as prior knowledge. Then, it incorporates them into probabilistic graphical models, ultimately deducing two novel multi-view matching costs. This significantly enhances the robustness to challenging situations and improves the completeness of the reconstruction. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our approach. Jiaqi Yang 0002, Yanan He, Chunlin Ren, Qingshan Xu 0001, Siwen Quan, Xiyu Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | VoMarkSplat: Robust watermarking for 3D Gaussian splatting with patch and multi-convolutional voting
Tianyu Xiong, Rui Li 0013, Jiaqi Yang 0002, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | Robust Context Modeling for Unsupervised Non-Rigid Point Cloud CorrespondenceabstractWe address the “long-range ambiguity” problem for unsupervised non-rigid point cloud correspondence, where corresponding points own inconsistent features while different local regions are spatially or geometrically similar. Previous methods struggle with this problem, since local reference frames (LRF) or coordinate-based methods struggle to exclude locally similar or spatially near mismatches, and widely used independent geometric relations might be inconsistent under non-rigid deformation, introducing extra ambiguity. To this end, we propose a novel robust context modeling module (RCM) to alleviatelong-range ambiguityin two aspects: 1) RCM tackles the ambiguity problem by introducing inter-relation attention (IRA), which mines robust cues from the interplay between relative geometric relations. 2) RCM enhances features with accessiblelong-rangeinformation from IRAs, following a local-to-global manner. Our method shows significant improvements in multiple benchmarks, with accurate correspondence over rotation and large deformation perturbation. Specifically, our method achieves a new state-of-the-art performance with correspondence accuracy of 33.9% and mean error of 4.2 on the SURREAL benchmark. Rui Li 0013, Jiaming Guo, Ya'nan He, Zhengbao Wang, Xian-Feng Han, Kun Sun 0002, Jiaqi Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | AxisPose: Model-Free Matching-Free Single-Shot 6D Object Pose Estimation via Axis GenerationabstractObject pose estimation is a fundamental task in computer vision and plays an important role in various applications such as robotics, augmented reality, and autonomous manipulation. Existing studies often demand complex inputs or depend on correspondence-based matching between 2D image features and 3D object representations. While effective, these methods rely strongly on explicit appearance matching, often requiring multi-view inputs, depth sensors, or CAD models, which limits their scalability and robustness. Building on top of the pioneering generative studies, we propose AxisPose, a model-free, matching-free, and single-view 6D pose estimation framework that departs from conventional correspondence-based paradigms. Unlike existing methods, AxisPose directly infers a pose representation by learning a latent distribution of object orientation axes through a diffusion model. Specifically, AxisPose introduces an Axis Generation Module (AGM) that progressively denoises tri-axial orientation fields guided by geometric consistency constraints, and a Triaxial Back-projection Module (TBM) to recover the final 6D pose from the generated orientation axes without relying on explicit 2D- 2D/3D correspondences. AxisPose achieves strong cross-instance generalization, enabling a single model to handle multiple object categories without retraining. Extensive experiments on LINEMOD and YCB-Video datasets demonstrate that AxisPose improves the Average Distance Deviation score from 0.733 to 0.814 over the strong baseline NOPE, using only a single RGB input. The code is available at https://github.com/pubyLu/AxisPose/tree/main. Yang Zou 0004, Zhaoshuai Qi, Weipeng Sun, Xingyuan Li 0005, Jiaqi Yang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Disentangling Global Orientation With Test-Time Rectification for Unsupervised Non-Rigid Point Cloud Correspondence
Jiaming Guo, Zhengbao Wang, Rui Li 0013, Jiaqi Yang 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | MAC++: Going Further with Maximal Cliques for 3D RegistrationabstractMaximal cliques (MAC) represent a novel state-of-theart approach for 3D registration from correspondences, however, it still suffers from extremely severe outliers. In this paper, we introduce a robust learning-free estimator called MAC++, exploring maximal cliques for$3 D$registration from the following two perspectives: 1)$A$novel hypothesis generation method utilizing putative seeds through voting to guide the construction of maximal clique pools, effectively preserving more potential correct hypotheses. 2) A progressive hypothesis evaluation method that continuously reduces the solution space in a “global-clusters-cluster-individual” manner rather than traditional one-shot techniques, greatly alleviating the issue of missing good hypotheses. Experiments conducted on U3M, 3DMatch/3DLoMatch, and KITTI-LC datasets show the new state-of-the-art performance of MAC++. MAC++ demonstrates the capability to handle extremely low inlier ratio data where MAC fails (e.g., showing 27.1%/30.6% registration recall improvements on 3DMatch/3DLoMatch with$<1 \%$inliers). Xiyu Zhang 0001, Yanning Zhang 0001, Jiaqi Yang 0002 |
3DV | 3 |
| 2025 | SPU-IMR: Self-supervised Arbitrary-scale Point Cloud Upsampling via Iterative Mask-recovery NetworkabstractPoint cloud upsampling aims to generate dense and uniformly distributed point sets from sparse point clouds. Existing point cloud upsampling methods typically approach the task as an interpolation problem. They achieve upsampling by performing local interpolation between point clouds or in the feature space, then regressing the interpolated points to appropriate positions. By contrast, our proposed method treats point cloud upsampling as a global shape completion problem. Specifically, our method first divides the point cloud into multiple patches. Then a masking operation is applied to remove some patches, leaving visible point cloud patches. Finally, our custom-designed neural network iterative completes the missing sections of the point cloud through the visible parts. During testing, by selecting different mask sequences, we can restore various complete patches. A sufficiently dense upsampled point cloud can be obtained by merging all the completed patches. We demonstrate the superior performance of our method through both quantitative and qualitative experiments, showing overall superiority against both existing self-supervised and supervised methods. Ziming Nie, Qiao Wu, Chenlei Lv, Siwen Quan, Zhaoshuai Qi, Muze Wang, Jiaqi Yang 0002 |
AAAI | 7 |
| 2025 | Unlocking Generalization Power in LiDAR Point Cloud RegistrationabstractIn real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety in autonomous driving and other LiDAR-based applications. However, current methods fall short in achieving this level of generalization. To address these limitations, we propose UGP, a pruned framework designed to enhance generalization power for LiDAR point cloud registration. The core insight in UGP is the elimination of cross-attention mechanisms to improve generalization, allowing the network to concentrate on intra-frame feature extraction. Additionally, we introduce a progressive self-attention module to reduce ambiguity in large-scale scenes and integrate Bird’s Eye View (BEV) features to incorporate semantic information about scene elements. Together, these enhancements significantly boost the network’s generalization performance. We validated our approach through various generalization experiments in multiple outdoor scenes. In cross-distance generalization experiments on KITTI and nuScenes, UGP achieved state-of-the-art mean Registration Recall rates of 94.5% and 91.4%, respectively. In cross-dataset generalization from nuScenes to KITTI, UGP achieved a state-of-the-art mean Registration Recall of 90.9%. Code will be available at https://github.com/peakpang/UGP Zhenxuan Zeng, Qiao Wu, Xiyu Zhang 0001, Lin Wu 0001, Pei An, Jiaqi Yang 0002, Peng Wang 0015 |
CVPR | 6 |
| 2025 | MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPabstractImage-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. The differential perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing the projective constraints on 2D-3D correspondences. However, differential PnP is highly sensitive to noise and outliers in the predicted correspondences. This issue hinders the effectiveness of correspondence learning. Inspired by the robustness of blind PnP against noise and outliers in correspondences, we propose an approximated blind PnP based correspondence learning approach. To mitigate the high computational cost of blind PnP, we simplify blind PnP to an amenable task of minimizing Chamfer distance between learned 2D and 3D keypoints, called MinCD-PnP. To effectively solve MinCD-PnP, we design a lightweight multi-task learning module, named as MinCD-Net, which can be easily integrated into the existing I2P registration architectures. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected datasets demonstrate that MinCD-Net outperforms state-of-the-art methods and achieves a higher inlier ratio (IR) and registration recall (RR) in both cross-scene and cross-dataset settings. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Liangliang Nan |
ICCV | 2 |
| 2025 | ArgMatch: Adaptive Refinement Gathering for Efficient Dense Matching
Yuxin Deng 0002, Kaining Zhang, Linfeng Tang, Jiaqi Yang 0002, Jiayi Ma 0001 |
ICCV | 4 |
| 2025 | BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation
Yuanhong Yu 0003, Chen Zhao 0025, Junhao Yu, Jiaqi Yang 0002, Ruizhen Hu, Yujun Shen, Xiaowei Zhou 0001, Sida Peng |
ICCV | 5 |
| 2025 | HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D RegistrationabstractGeometric constraints between feature matches are critical in 3D point cloud registration problems. Existing approaches typically model unordered matches as a consistency graph and sample consistent matches to generate hypotheses. However, explicit graph construction introduces noise, posing great challenges for handcrafted geometric constraints to render consistency. To overcome this, we propose HyperGCT, a flexible dynamic Hyper-GNN-learned geometric ConstrainT that leverages high-order consistency among 3D correspondences. To our knowledge, HyperGCT is the first method that mines robust geometric constraints from dynamic hypergraphs for 3D registration. By dynamically optimizing the hypergraph through vertex and edge feature aggregation, HyperGCT effectively captures the correlations among correspondences, leading to accurate hypothesis generation. Extensive experiments on 3DMatch, 3DLoMatch, KITTI-LC, and ETH show that HyperGCT achieves state-of-the-art performance. Furthermore, HyperGCT is robust to graph noise, demonstrating a significant advantage in terms of generalization. Xiyu Zhang 0001, Jiayi Ma 0001, Zhaoshuai Qi, Fei Hui, Jiaqi Yang 0002, Yanning Zhang 0001 |
ICCV | 7 |
| 2025 | Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology RelationshipabstractImage-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Jie Ma 0003, Liangliang Nan |
IJCAI | 2 |
| 2025 | Enhance Image-to-Point-Cloud Registration with Beltrami Flow
Pei An, You Yang 0002, Jiaqi Yang 0002, Muyao Peng, Qiong Liu 0001, Liangliang Nan |
Int. J. Comput. Vis. | 3 |
| 2025 | Pre-training meets iteration: Learning for robust 3D point cloud denoising
Siwen Quan, Hebin Zhao, Zhao Zeng, Ziming Nie, Jiaqi Yang 0002 |
Pattern Recognit. Lett. | 5 |
| 2025 | TAP-Track: Generalizable Spacecraft Pose Tracking by Tracking Any PointsabstractRecent learning-based spacecraft pose tracking methods have demonstrated impressive improvement in estimation accuracy and potential scalability to complex space environment. However, most of them still rely on the detection of discriminative keypoints on a known 3D model, limiting the generalization to unknown spacecraft. To this end, we propose, to the best of our knowledge, the first generalizable spacecraft pose tracking method. Instead of requiring a known model, we only assume the existence of at least one planar structure, e.g. solar panels, which holds for most satellites in general scenes. Additionally, the proposed method tracks any points on the plane across multi-frame followed by a re-projection error minimization, rather than detecting keypoints between image pairs, allowing robust capture of “long-term” temporal information among frames even for textureless surfaces without sufficient keypoints. Moreover, we also constructed the first large-scale dataset G-SPET for generalizable spacecraft pose estimation and tracking. It covers 174 satellites with diversity structures and rich annotations, increasing the number of targets in previous datasets by almost two orders of magnitude. Extensive evaluations on the proposed dataset have demonstrated the superiority of our method over state-of-the-art methods. The code and dataset will be made publicly available soon. Zhaoshuai Qi, Pulin Chen, Huilin Fan, Yu Zhu 0004, Jiaqi Yang 0002, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | 3D Single-Object Tracking in Point Clouds with High Temporal Variation
Qiao Wu, Kun Sun 0002, Pei An, Mathieu Salzmann, Yanning Zhang 0001, Jiaqi Yang 0002 |
ECCV (7) | 6 |
| 2024 | Mutual Voting for Ranking 3D CorrespondencesabstractConsistent correspondences between point clouds are vital to 3D vision tasks such as registration and recognition. In this paper, we present a mutual voting method for ranking 3D correspondences. The key insight is to achieve reliable scoring results for correspondences by refining both voters and candidates in a mutual voting scheme. First, a graph is constructed for the initial correspondence set with the pairwise compatibility constraint. Second, nodal clustering coefficients are introduced to preliminarily remove a portion of outliers and speed up the following voting process. Third, we model nodes and edges in the graph as candidates and voters, respectively. Mutual voting is then performed in the graph to score correspondences. Finally, the correspondences are ranked based on the voting scores and top-ranked ones are identified as inliers. Feature matching, 3D point cloud registration, and 3D object recognition experiments on various datasets with different nuisances and modalities verify that MV is robust to heavy outliers under different challenging settings, and can significantly boost 3D point cloud registration and 3D object recognition performance. Jiaqi Yang 0002, Xiyu Zhang 0001, Shichao Fan, Chunlin Ren, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | MAC: Maximal Cliques for 3D RegistrationabstractThis paper presents a 3D registration method with maximal cliques (MAC) for 3D point cloud registration (PCR). The key insight is to loosen the previous maximum clique constraint and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each representing a consensus set. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. In addition, we present a variant of MAC if given overlap prior, called MAC-OP. Overlap prior further enhances MAC from many technical aspects, such as graph construction with re-weighted nodes, hypotheses generation from cliques with additional constraints, and hypothesis evaluation with overlap-aware weights. Extensive experiments demonstrate that both MAC and MAC-OP effectively increase registration recall, outperform various state-of-the-art methods, and boost the performance of deep-learned methods. For instance, MAC combined with GeoTransformer achieves a state-of-the-art registration recall of [Formula: see text] on 3DMatch / 3DLoMatch. We perform synthetic experiments on 3DMatch-LIR / 3DLoMatch-LIR, a dataset with extremely low inlier ratios for 3D registration in ultra-challenging cases. Jiaqi Yang 0002, Xiyu Zhang 0001, Peng Wang 0015, Yulan Guo, Kun Sun 0002, Qiao Wu, Shikun Zhang, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Toward Meta-Shape-Based Multi-View 3D Point Cloud Registration: An EvaluationabstractReducing cumulative registration error is critical to accurate 3D multi-view registration. Meta-shape based methods optimize rigid transformations of point clouds by iteratively registering each point cloud with a meta-shape, which remain popular solutions to 3D multi-view registration. However, the merits and demerits of existing meta-shape based methods remain unclear. Moreover, we argue that simpler meta-shape based solutions can achieve even better performance. To this end, we evaluate seven representative meta-shape based methods in this work, including four existing ones and three modified ones, in order to investigate the problem of defining a good meta-shape. In particular, we first abstract the main steps of considered methods. Then, experiments on both object and scene datasets with real and synthetic cumulative registration errors are deployed for an in-depth evaluation. Finally, based on the experimental outcomes, we give a discussion on the advantages and limitations of meta-shape based methods. We demonstrate prior works have used unnecessarily complicated techniques for cumulative error elimination and our slightly modified simpler solutions can achieve competitive performance on experimental datasets. Shikun Zhang, Jiaqi Yang 0002, Zhaoshuai Qi, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Survey of Extrinsic Calibration on LiDAR-Camera System for Intelligent Vehicle: Challenges, Approaches, and TrendsabstractA system with light detection and ranging (LiDAR) and camera (named as LiDAR-camera system) plays the essential role in intelligent vehicle (IV), for it provides 3D spatial and 2D texture features for 3D scene understanding. To leverage LiDAR point cloud and image, extrinsic calibration is a crucial technique, for it can align 2D pixel and 3D point in the pixel-level accuracy. With the rapid development of IV, calibration demand is shifted from offline to online, from the specific scenes to the open scenes. It brings new challenge to the calibration task. Although numbers of approaches have been proposed in the last decade, there lacks an in-depth summary about this topic. Thus, we conduct a survey of extrinsic calibration. Theoretically, the key of calibration is to build correspondence from LiDAR point cloud and optical image. From the viewpoint of correspondence, we attempt to divide the mainstream approaches into explicit and implicit correspondence based methods. After that, we summarize both the strength and weakness of the current works, provide the methods comparison, and list the open-source implementations. Finally, we analyze the tendency of calibration approach, discuss the remained problems in this field. We believe that this survey benefits to the community of autonomous driving. Pei An, Junfeng Ding, Siwen Quan, Jiaqi Yang 0002, You Yang 0002, Qiong Liu 0001, Jie Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Multi-view Inverse Rendering for Large-scale Real-world Indoor ScenesabstractWe present a efficient multi-view inverse rendering method for large-scale real-world indoor scenes that reconstructs global illumination and physically-reasonable SVBRDFs. Unlike previous representations, where the global illumination of large scenes is simplified as multiple environment maps, we propose a compact representation called Texture-based Lighting (TBL). It consists of 3D mesh and HDR textures, and efficiently models direct and infinite-bounce indirect lighting of the entire large scene. Based on TBL, we further propose a hybrid lighting representation with precomputed irradiance, which significantly improves the efficiency and alleviates the rendering noise in the material optimization. To physically disentangle the ambiguity between materials, we propose a three-stage material optimization strategy based on the priors of semantic segmentation and room segmentation. Extensive experiments show that the proposed method outperforms the state-of-the-art quantitatively and qualitatively, and enables physically-reasonable mixed-reality applications such as material editing, editable novel view synthesis and relighting. The project page is at https://lzleejean.github.io/TexIR. Lingli Wang, Mofang Cheng, Cihui Pan, Jiaqi Yang 0002 |
CVPR | 5 |
| 2023 | 3D Registration with Maximal CliquesabstractAs a fundamental problem in computer vision, 3D point cloud registration (PCR) aims to seek the optimal pose to align a point cloud pair. In this paper, we present a 3D registration method with maximal cliques (MAC). The key insight is to loosen the previous maximum clique constraint, and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each of which represents a consensus set. We perform node-guided clique selection then, where each node corresponds to the maximal clique with the greatest graph weight. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. Extensive experiments on U3M, 3DMatch, 3DLoMatch and KITTI demonstrate that MAC effectively increases registration accuracy, outperforms various state-of-the-art methods and boosts the performance of deep-learned methods. MAC combined with deep-learned methods achieves state-of-the-art registration recall of 95.7% /78.9% on 3DMatch /3DLoMatch. Xiyu Zhang 0001, Jiaqi Yang 0002, Shikun Zhang, Yanning Zhang 0001 |
CVPR | 2 |
| 2023 | Hierarchical Prior Mining for Non-local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of a target from a set of 2D images. Recent advances in MVS have shown that it is important to perceive non-local structured information for recovering geometry in low-textured areas. In this work, we propose a Hierarchical Prior Mining for Non-local Multi-View Stereo (HPM-MVS). The key characteristics are the following techniques that exploit non-local information to assist MVS: 1) A Non-local Extensible Sampling Pattern (NESP), which is able to adaptively change the size of sampled areas without becoming snared in locally optimal solutions. 2) A new approach to leverage non-local reliable points and construct a planar prior model based on K-Nearest Neighbor (KNN), to obtain potential hypotheses for the regions where prior construction is challenging. 3) A Hierarchical Prior Mining (HPM) framework, which is used to mine extensive non-local prior information at different scales to assist 3D model recovery, this strategy can achieve a considerable balance between the reconstruction of details and low-textured areas. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our method. Our code will be available at https://github.com/CLinvx/HPM-MVS. Chunlin Ren, Qingshan Xu 0001, Shikun Zhang, Jiaqi Yang 0002 |
ICCV | 4 |
| 2023 | MixCycle: Mixup Assisted Semi-Supervised 3D Single Object Tracking with Cycle Consistencyabstract3D single object tracking (SOT) is an indispensable part of automated driving. Existing approaches rely heavily on large, densely labeled datasets. However, annotating point clouds is both costly and time-consuming. Inspired by the great success of cycle tracking in unsupervised 2D SOT, we introduce the first semi-supervised approach to 3D SOT. Specifically, we introduce two cycle-consistency strategies for supervision: 1) Self tracking cycles, which leverage labels to help the model converge better in the early stages of training; 2) forward-backward cycles, which strengthen the tracker’s robustness to motion variations and the template noise caused by the template update strategy. Furthermore, we propose a data augmentation strategy named SOTMixup to improve the tracker’s robustness to point cloud diversity. SOTMixup generates training samples by sampling points in two point clouds with a mixing rate and assigns a reasonable loss weight for training according to the mixing rate. The resulting MixCycle approach generalizes to appearance matching-based trackers. On the KITTI benchmark, based on the P2B tracker [16], MixCycle trained with 10% labels outperforms P2B trained with 100% labels, and achieves a 28.4% precision improvement when using 1% labels. Our code will be released at https://github.com/Mumuqiao/MixCycle. Qiao Wu, Jiaqi Yang 0002, Kun Sun 0002, Chu'ai Zhang, Yanning Zhang 0001, Mathieu Salzmann |
ICCV | 2 |
| 2023 | VOID: 3D object recognition based on voxelization in invariant distance space
Jiaqi Yang 0002, Shichao Fan, Siwen Quan, Yanning Zhang 0001 |
Vis. Comput. | 1 |
| 2022 | PhyIR: Physics-based Inverse Rendering for Panoramic Indoor ImagesabstractInverse rendering of complex material such as glossy, metal and mirror material is a long-standing ill-posed problem in this area, which has not been well solved. Previous approaches cannot tackle them well due to simplified BRDF and unsuitable illumination representations. In this paper, we present PhyIR, a neural inverse rendering method with a more completed SVBRDF representation and a physics-based in-network rendering layer, which can handle complex material and incorporate physical constraints by re-rendering realistic and detailed specular reflectance. Our framework estimates geometry, material and Spatially-Coherent (SC) illumination from a single indoor panorama. Due to the lack of panoramic datasets with completed SVBRDF and full-spherical light probes, we introduce an artist-designed dataset named FutureHouse with high-quality geometry, SVBRDF and per-pixel Spatially-Varying (SV) lighting. To ensure the coherence of SV lighting, a novel SC loss is proposed. Extensive experiments on both synthetic and real-world data show that the proposed method outperforms the state-of-the-arts quantitatively and qualitatively, and is able to produce photorealistic results for a number of applications such as dynamic virtual object insertion. Lingli Wang, Cihui Pan, Jiaqi Yang 0002 |
CVPR | 5 |
| 2022 | Unsupervised Learning of 3D Semantic Keypoints with Mutual Reconstruction
Haocheng Yuan, Chen Zhao 0025, Shichao Fan, Jiaxi Jiang, Jiaqi Yang 0002 |
ECCV (2) | 5 |
| 2022 | Dual spin-image: A bi-directional spin-image variant using multi-scale radii for 3D local shape description
Daryl L. Bibissi, Jiaqi Yang 0002, Siwen Quan, Yanning Zhang 0001 |
Comput. Graph. | 2 |
| 2022 | Rotation invariant point cloud analysis: Where local geometry meets global topology
Chen Zhao 0025, Jiaqi Yang 0002, Angfan Zhu, Zhiguo Cao 0001, Xin Li 0005 |
Pattern Recognit. | 2 |
| 2022 | Toward Efficient and Robust Metrics for RANSAC Hypotheses and 3D Rigid RegistrationabstractThis paper focuses on developing efficient and robust evaluation metrics for RANSAC hypotheses to achieve accurate 3D rigid registration. Estimating six-degree-of-freedom (6-DoF) pose from feature correspondences remains a popular approach to 3D rigid registration, where random sample consensus (RANSAC) is a well-known solution to this problem. However, existing metrics for RANSAC hypotheses are either time-consuming or sensitive to common nuisances, parameter variations, and different application scenarios, resulting in performance deterioration with respect to overall registration accuracy and speed. We alleviate this problem by first analyzing the contributions of inliers and outliers and then proposing several efficient and robust metrics with different designing motivations for RANSAC hypotheses. Comparative experiments on four standard datasets with different nuisances and application scenarios verify that our considered metrics can significantly improve the registration performance and are more robust than several state-of-the-art competitors, making them good gifts to practical applications. This work also draws an interesting conclusion, i.e., not all inliers are equal while all outliers should be equal, which may shed new light on this research problem. Jiaqi Yang 0002, Siwen Quan, Qian Zhang 0046, Yanning Zhang 0001, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Correspondence Selection With Loose-Tight Geometric Voting for 3-D Point Cloud RegistrationabstractThis article presents a simple yet effective method for 3-D correspondence selection and point cloud registration. It first models the initial correspondence set as a graph with nodes representing correspondences and edges connecting geometrically compatible nodes. Such graphs offer either loose or tight geometric constraints for judging the correctness of correspondence, e.g., edges, loops, and cliques. Then, we render these constraints dynamic voters to judge the correctness of a node. More specifically, we develop a loose–tight geometric voting (LT-GV) method that employs both loose and tight geometric constraints in the graph to score 3-D feature correspondences. The motivation behind this is to strike a balanced performance in terms of precision and recall because loose and tight constraints are complementary to each other. Under the dynamic voting scheme with both loose and tight voters, consistent correspondences can be retrieved based on the voting score. Both feature-matching and 3-D point cloud registration experiments on datasets with different modalities, challenges, application scenarios, and comparisons with state-of-the-art methods (including deep learned methods) verify that our LT-GV is effective for correspondence selection, robust to a number of nuisances, and able to dramatically boost 3-D point cloud registration performance. Jiaqi Yang 0002, Siwen Quan, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | SAC-COT: Sample Consensus by Sampling Compatibility Triangles in Graphs for 3-D Point Cloud RegistrationabstractSix-degree-of-freedom (6-DOF) pose estimation from feature correspondences remains a popular and robust approach for 3-D registration. However, heavy outliers that existed in the initial correspondence set pose a great challenge to this problem. This article presents a simple yet effective estimator called SAmple Consensus by sampling COmpatibility Triangles in graphs (SAC-COT) for robust 6-DOF pose estimation and 3-D registration. The key novelty is a guided three-point sampling approach. It is based on a novel correspondence sample representation, i.e., COmpatibility Triangle (COT). We first model the correspondence set as a graph with nodes connecting compatible correspondences. Then, by ranking and sampling COTs formed by ternary loops, we show that correct hypotheses can be generated in early iteration stage. Finally, the hypothesis generated by the COT yielding to the maximum consensus is the output of SAC-COT. Extensive experiments on six data sets and extensive comparisons with the state-of-the-art estimators confirm that: 1) SAC-COT can achieve accurate registrations with a few iterations and 2) SAC-COT outperforms all competitors and is ultrarobust when confronted with Gaussian noise, data decimation, holes, clutter, partial overlap, varying scales of input correspondences, and data modality variation. Jiaqi Yang 0002, Siwen Quan, Zhaoshuai Qi, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | 3D Correspondence Grouping with Compatibility Features
Jiaqi Yang 0002, Zhiguo Cao 0001, Yanning Zhang 0001 |
PRCV (2) | 1 |
| 2021 | A light-weight, efficient, and general cross-modal image fusion network
Aiqing Fang, Jiaqi Yang 0002, Beibei Qin, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2021 | A Performance Evaluation of Correspondence Grouping Methods for 3D Rigid Data MatchingabstractSeeking consistent point-to-point correspondences between 3D rigid data (point clouds, meshes, or depth maps) is a fundamental problem in 3D computer vision. While a number of correspondence selection methods have been proposed in recent years, their advantages and shortcomings remain unclear regarding different applications and perturbations. To fill this gap, this paper gives a comprehensive evaluation of nine state-of-the-art 3D correspondence grouping methods. A good correspondence grouping algorithm is expected to retrieve as many as inliers from initial feature matches, giving a rise in both precision and recall as well as facilitating accurate transformation estimation. Toward this rule, we deploy experiments on three benchmarks with different application contexts, including shape retrieval, 3D object recognition, and point cloud registration. We also investigate various perturbations such as noise, point density variation, clutter, occlusion, partial overlap, different scales of initial correspondences, and different combinations of keypoint detectors and descriptors. The rich variety of application scenarios and nuisances result in different spatial distributions and inlier ratios of initial feature correspondences, thus enabling a thorough evaluation. Based on the outcomes, we give a summary of the traits, merits, and demerits of evaluated approaches and indicate some potential future research directions. Jiaqi Yang 0002, Ke Xian, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Non-linear and selective fusion of cross-modal images
Aiqing Fang, Jiaqi Yang 0002, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2020 | Compatibility-Guided Sampling Consensus for 3-D Point Cloud RegistrationabstractThis article presents an efficient and robust estimator called compatibility-guided sampling consensus (CG-SAC) to achieve accurate 3-D point cloud registration. For correspondence-based registration methods, the random sample consensus (RANSAC) is served as a de facto solution for rigid transformation estimation from a number of feature correspondences. Unfortunately, RANSAC still suffers from two major limitations. First, it generates a hypothesis with at least three samples and desires a very large number of iterations to attain reasonable results, making it relatively time consuming. Second, the randomness during sampling can result in inaccurate results as it is highly potential to miss the optimal hypothesis. To solve these problems, we propose a compatibility-guided sampling strategy to eliminate randomness during sampling. In particular, only two correspondences are required by our method for hypothesis generation. We then rank correspondence pairs according to their compatibility scores because compatible correspondences are more likely to be correct and can yield more reasonable hypotheses. In addition, we propose a new geometric constraint named the distance between salient points (DSP) to measure the compatibility of two correspondences. Experiments on a set of real-world point cloud data with different application contexts and data modalities confirm the effectiveness of the proposed method. Comparison with several state-of-the-art estimators demonstrates the overall superiority of our CG-SAC estimator with regards to precision and time efficiency. Siwen Quan, Jiaqi Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Evaluating Local Geometric Feature Representations for 3D Rigid Data MatchingabstractLocal geometric descriptors act as an essential component for 3D rigid data matching. A rotational invariant local geometric descriptor usually consists of two components: local reference frame (LRF) and feature representation. However, existing evaluation efforts have mainly been paid on the LRF or the overall descriptor and the quantitative comparison of feature representations remains unexplored. This paper fills the gap by comprehensively evaluating nine state-of-the-art local geometric feature representations. In particular, our evaluation assesses feature representations based on ground-truth LRFs such that the ranking of tested methods is more convincing as compared with existing studies. The experiments are deployed on six standard datasets with various application scenarios (shape retrieval, point cloud registration, and object recognition) and data modalities (LiDAR, Kinect, and Space Time) as well as perturbations including Gaussian noise, shot noise, data decimation, clutter, occlusion, and limited overlap. The evaluated terms cover the major concerns for a feature representation, e.g., distinctiveness, robustness, compactness, and efficiency. The outcomes present interesting findings that may shed new light on this community and provide complementary perspectives to existing evaluations on the topic of local geometric feature description. A summary of evaluated methods regarding their peculiarities is finally presented to guide real-world applications and new descriptor crafting. Jiaqi Yang 0002, Siwen Quan, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Image Feature Correspondence Selection: A Comparative Study and a New ContributionabstractImage feature correspondence selection is pivotal to many computer vision tasks from object recognition to 3D reconstruction. Although many correspondence selection algorithms have been developed in the past decade, there still lacks an in-depth evaluation and comparison in the open literature, which makes it difficult to choose the appropriate algorithm for a specific application. This paper attempts to fill this gap by evaluating eight competing correspondence selection algorithms including both classical methods and current state-of-the-art ones. In addition to preselected correspondences, we have compared different combinations of detector and descriptor on four standard datasets. The diversity of those datasets cover a wide range of uncertainty factors including zoom, rotation, blur, viewpoint change, JPEG compression, light change, different rendering styles and multiple structures. We have measured the quality of competing correspondence selection algorithms in terms of four performance metrics -i.e., precision, recall, F-measure and efficiency. Moreover, we propose to combine the strengths of eight competing methods by combining their correspondence selection results. Extensive experimental results are reported to demonstrate the superiority of several fusion strategies to individual methods, which suggests the possibility of adaptively combining those methods for even better performance. Chen Zhao 0025, Zhiguo Cao 0001, Jiaqi Yang 0002, Ke Xian, Xin Li 0005 |
IEEE Trans. Image Process. | 3 |
| 2019 | NM-Net: Mining Reliable Neighbors for Robust Feature CorrespondencesabstractFeature correspondence selection is pivotal to many feature-matching based tasks in computer vision. Searching spatially k-nearest neighbors is a common strategy for extracting local information in many previous works. However, there is no guarantee that the spatially k-nearest neighbors of correspondences are consistent because the spatial distribution of false correspondences is often irregular. To address this issue, we present a compatibility-specific mining method to search for consistent neighbors. Moreover, in order to extract and aggregate more reliable features from neighbors, we propose a hierarchical network named NM-Net with a series of graph convolutions that is insensitive to the order of correspondences. Our experimental results have shown the proposed method achieves the state-of-the-art performance on four datasets with various inlier ratios and varying numbers of feature consistencies. Chen Zhao 0025, Zhiguo Cao 0001, Xin Li 0005, Jiaqi Yang 0002 |
CVPR | 5 |
| 2019 | Mean-Variance Loss for Monocular Depth EstimationabstractMonocular depth estimation is a widely studied computer vision problem with a vast variety of applications. In this paper, we formulate it as a pixel-wise classification task and use a mean-variance loss for robust depth estimation via distribution learning. More precisely, the mean-variance loss is composed of a mean loss that penalizes the difference between the mean of predicted depth distribution and the ground-truth depth, and a variance loss that penalizes the variance of predicted depth distribution to obtain a more focused distribution. The mean-variance loss is jointly trained with the soft-max loss to supervise a Deep Convolutional Neural Networks (DCNN) for depth estimation. Experimental results on the NYUDv2 dataset show that the proposed method outperforms previous state-of-the-art approaches. Hongwei Zou, Ke Xian, Jiaqi Yang 0002, Zhiguo Cao 0001 |
ICIP | 3 |
| 2019 | Limited Receptive Field Network for Real-Time Driving Scene Semantic Segmentation
Dehui Li, Zhiguo Cao 0001, Ke Xian, Jiaqi Yang 0002, Xinyuan Qi, Wei Li 0132 |
PRICAI (3) | 4 |
| 2019 | Ranking 3D feature correspondences via consistency voting
Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001, Weidong Yang 0006 |
Pattern Recognit. Lett. | 1 |
| 2019 | Aligning 2.5D Scene Fragments With Distinctive Local Geometric Features and Voting-Based CorrespondencesabstractAligning 2.5D views has been extensively explored in the past decades, where most prior works have concentrated on object data with complex structures. This paper presents a method to align real-word scene scans with challenging features such as noise, poor geometric information, and highly repeatable patterns. Our method consists of two modules: pairwise and multiview alignments. Key to the proposed pairwise alignment method is the rotational contour signature geometric feature and voting-based correspondence selection algorithm. The former promises strong discriminative power for 2.5D scene data, while the latter affords high-quality correspondences via a voting process for all raw feature matches using L2distance and point pair affinity constraints. For the multiview alignment method, we first use a connected graph algorithm to establish the connections of all 2.5D views for coarse merging; then, we propose a shape-growing iterative closest point algorithm for further refinement. Experiments are conducted on scene point cloud datasets addressing both the indoor and outdoor scenarios, whereby we demonstrate that the proposed pairwise alignment method clearly outperforms the state of the art. Moreover, the proposed multiview alignment method manages to put multiple unordered 2.5D scene fragments into a unified coordinate system automatically, accurately, and efficiently. Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Scalable Multi-Consistency Feature Matching with Non-Cooperative GamesabstractCorrespondence selection aiming at seeking correct relationships between two images is a fundamental and critical task in computer vision. This paper attempts to select consistent correspondences in the context of dynamic scenarios where multiple matching consistencies are normally incorporated. To this end, we present a grid-based game-theoretic matching (Grid-GTM) method which is divided into three processes, i.e., grid matching, local games and enrichment. Specifically, grid matching translates the multi-consistency problem into several independent single-consistency problems to decrease difficulties of selection and boost the efficiency. Local games extended under the guidance of a novel payoff function guarantee that mismatches are effectively removed. Enrichment is added to recover correct matches neglected by local games. Crucially, our approach achieves the state-of-the-art performance compared with seven algorithms in comprehensive evaluations. In addition, we construct a dataset that involves multiple consistencies under three different scenes in this paper. Chen Zhao 0025, Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
ICIP | 2 |
| 2018 | Toward the Repeatability and Robustness of the Local Reference Frame for 3D Shape Matching: An EvaluationabstractThe local reference frame (LRF), as an independent coordinate system constructed on the local 3D surface, is broadly employed in 3D local feature descriptors. The benefits of the LRF include rotational invariance and full 3D spatial information, thereby greatly boosting the distinctiveness of a 3D feature descriptor. There are numerous LRF methods in the literature; however, no comprehensive study comparing their repeatability and robustness performance under different application scenarios and nuisances has been conducted. This paper evaluates eight state-of-the-art LRF proposals on six benchmarks with different data modalities (e.g., LiDAR, Kinect, and Space Time) and application contexts (e.g., shape retrieval, 3D registration, and 3D object recognition). In addition, the robustness of each LRF to a variety of nuisances, including varying support radii, Gaussian noise, outliers (shot noise), mesh resolution variation, distance to boundary, keypoint localization error, clutter, occlusion, and partial overlap, is assessed. The experimental study also measures the performance under different keypoint detectors, descriptor matching performance when using different LRFs and feature representation combinations, as well as computational efficiency. Considering the evaluation outcomes, we summarize the traits, advantages, and current limitations of the tested LRF methods. Jiaqi Yang 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Performance Evaluation of 3D Correspondence Grouping AlgorithmsabstractThis paper presents a thorough evaluation of several widely-used 3D correspondence grouping algorithms, motived by their significance in vision tasks relying on correct feature correspondences. A good correspondence grouping algorithm is desired to retrieve as many as inliers from initial feature matches, giving a rise in both precision and recall. Towards this rule, we deploy the experiments on three benchmarks respectively addressing shape retrieval, 3D object recognition and point cloud registration scenarios. The variety in application context brings a rich category of nuisances including noise, varying point densities, clutter, occlusion and partial overlaps. It also results to different ratios of inliers and correspondence distributions for comprehensive evaluation. Based on the quantitative outcomes, we give a summarization of the merits/demerits of the evaluated algorithms from both performance and efficiency perspectives. Jiaqi Yang 0002, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001 |
3DV | 1 |
| 2017 | Rotational contour signatures for both real-valued and binary feature representations of 3D local shape
Jiaqi Yang 0002, Qian Zhang 0046, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2017 | Multi-attribute statistics histograms for accurate and robust pairwise registration of range images
Jiaqi Yang 0002, Qian Zhang 0046, Zhiguo Cao 0001 |
Neurocomputing | 1 |
| 2017 | TOLDI: An effective and robust approach for 3D local shape description
Jiaqi Yang 0002, Qian Zhang 0046, Yang Xiao 0007, Zhiguo Cao 0001 |
Pattern Recognit. | 1 |
| 2017 | The effect of spatial information characterization on 3D local feature descriptors: A quantitative evaluation
Jiaqi Yang 0002, Qian Zhang 0046, Zhiguo Cao 0001 |
Pattern Recognit. | 1 |
| 2016 | Rotational contour signatures for robust local surface descriptionabstractThis paper presents a novel local surface descriptor called rotational contour signatures (RCS) for 3D rigid objects. RCS comprises several signatures that characterize the 2D contour information derived from 3D-to-2D projection of the local surface. The inspiration of our encoding technique comes from that, viewing towards an object, its contour is an effective and robust cue for representing its shape. In order to achieve a comprehensive geometry encoding, the local surface is continually rotated in a predefined local reference frame (LRF) so that multi-view information is obtained. Experiments on two publicly available datasets demonstrate the effectiveness and robustness of the proposed descriptor. Further, comparisons with five state-of-the-art descriptors show the superiority of our RCS descriptor. Jiaqi Yang 0002, Qian Zhang 0046, Ke Xian, Yang Xiao 0007, Zhiguo Cao 0001 |
ICIP | 1 |
| 2016 | A fast and robust local descriptor for 3D point cloud registration
Jiaqi Yang 0002, Zhiguo Cao 0001, Qian Zhang 0046 |
Inf. Sci. | 1 |