Tong Wei 0002

dblp:49/933-2 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-6053-8732ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › multi-view geometry › epipolar geometry estimation
fundamental matrix estimation
1.322023
Generalized Differentiable RANSAC · ICCV 2023
Adaptive Reordering Sampler with Neurally Guided MAGSAC · ICCV 2023
Computer vision › 3D vision
robust estimation
1.322023
Generalized Differentiable RANSAC · ICCV 2023
Adaptive Reordering Sampler with Neurally Guided MAGSAC · ICCV 2023
Computer vision › 3D vision › multi-view geometry
two-view geometry
1.322023
Generalized Differentiable RANSAC · ICCV 2023
Adaptive Reordering Sampler with Neurally Guided MAGSAC · ICCV 2023
Computer vision › 3D vision › multi-view geometry › epipolar geometry estimation
essential matrix estimation
0.712023
Adaptive Reordering Sampler with Neurally Guided MAGSAC · ICCV 2023
Computer vision › 3D vision
point cloud registration
0.712023
Generalized Differentiable RANSAC · ICCV 2023

Methods — techniques the papers use, named apart from their topics

superpoint · 0.7sampling distribution learning · 0.7feature detection and matching networks · 0.7differentiable relaxation · 0.7deep network guidance · 0.7bayesian inlier probability update · 0.7SIFT · 0.7MAGSAC++ · 0.7
YearPublicationVenuePosition
2025 Breaking the Frame: Visual Place Recognition by Overlap Prediction
abstract
Visual place recognition methods struggle with occlusion and partial visual overlaps. We propose a novel visual place recognition approach based on overlap prediction, called VOP, shifting from traditional reliance on global image similarities and local features to image overlap prediction. VOP proceeds co-visible image sections by obtaining patch-level embeddings using a Vision Transformer backbone and establishing patch-to-patch correspondences without requiring expensive feature detection and matching. Our approach uses a voting mechanism to assess overlap scores for potential database images. It provides a nuanced image retrieval metric in challenging scenarios. Experimental results show that VOP leads to more accurate relative pose estimation and localization results on the retrieved image pairs than state-of-the-art base-lines on a number of large-scale, real-world indoor and outdoor benchmarks. The code is available at https://github.com/weitong8591/vop.git.
Tong Wei 0002, Philipp Lindenberger, Jiri Matas, Daniel Barath
WACV1
2023 Adaptive Reordering Sampler with Neurally Guided MAGSAC
abstract
We propose a new sampler for robust estimators that always selects the sample with the highest probability of consisting only of inliers. After every unsuccessful iteration, the inlier probabilities are updated in a principled way via a Bayesian approach. The probabilities obtained by the deep network are used as prior (so-called neural guidance) inside the sampler. Moreover, we introduce a new loss that exploits, in a geometrically justifiable manner, the orientation and scale that can be estimated for any type of feature, e.g., SIFT or SuperPoint, to estimate two-view geometry. The new loss helps to learn higher-order information about the underlying scene geometry. Benefiting from the new sampler and the proposed loss, we combine the neural guidance with the state-of-the-art MAGSAC++. Adaptive Reordering Sampler with Neurally Guided MAGSAC (ARS-MAGSAC) is superior to the state-of-the-art in terms of accuracy and run-time on the PhotoTourism and KITTI datasets for essential and fundamental matrix estimation. The code and trained models are available at https://github.com/weitong8591/ars_magsac.
Tong Wei 0002, Jiri Matas, Daniel Barath
ICCV1
2023 Generalized Differentiable RANSAC
abstract
We propose ▽-RANSAC, a generalized differentiable RANSAC that allows learning the entire randomized robust estimation pipeline. The proposed approach enables the use of relaxation techniques for estimating the gradients in the sampling distribution, which are then propagated through a differentiable solver. The trainable quality function marginalizes over the scores from all the models estimated within ▽-RANSAC to guide the network learning accurate and useful inlier probabilities or to train feature detection and matching networks. Our method directly maximizes the probability of drawing a good hypothesis, allowing us to learn better sampling distributions. We test ▽-RANSAC on various real-world scenarios on fundamental and essential matrix estimation, and 3D point cloud registration, outdoors and indoors, with handcrafted and learning-based features. It is superior to the state-of-the-art in terms of accuracy while running at a similar speed to its less accurate alternatives. The code and trained models are available at https://github.com/weitong8591/differentiable_ransac.
Tong Wei 0002, Alexander Shekhovtsov 0001, Jiri Matas, Daniel Barath
ICCV1
2021 Wavelet Multi-Level Attention Capsule Network for Texture Classification
abstract
Texture classification is one of the essential problems in computer vision. Due to the powerful feature extraction ability, convolutional neural network (CNN) based texture classification methods have attracted extensive attention in recent years. However, there are still some challenges, such as the extraction of multi-level texture features and their relationships. To address these problems, this letter proposes the wavelet multi-level attention capsule network (WMACapsNet), which integrates multi-scale wavelet decomposition and multi-level attention blocks into the capsule network. Specifically, multi-scale spectral features in frequency domain are extracted by multi-level wavelet transform; and then the self-attention block explores the dependencies of capsule features within each scale; finally, the cross-attention block refines capsule features and their relationships with attention mechanism across different scales. The proposed WMACapsNet provides an efficient way to explore spatial domain features, frequency domain features and their dependencies, useful for most texture classification tasks. Experimental results on several texture datasets show that the proposed WMACapsNet outperforms the state-of-the-art texture classification methods not only in accuracy but also in robustness.
Zhiyong Tao, Tong Wei 0002
IEEE Signal Process. Lett.2