Vojtech Panek

dblp:324/8183 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-0601-7682ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 100%
Computer graphics and multimedia
3 papers
Geometric modeling and processing · 68% Virtual and augmented reality · 32%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
visual localization
3.242026
A Guide to Structureless Visual Localization · Int. J. Comput. Vis. 2026
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization · Int. J. Comput. Vis. 2026
Visual Localization using Imperfect 3D Models from the Internet · CVPR 2023
Geometric modeling and processing
3d reconstruction
2.022026
A Guide to Structureless Visual Localization · Int. J. Comput. Vis. 2026
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization · Int. J. Comput. Vis. 2026
Computer vision › 3D vision
camera pose estimation
1.722026
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization · Int. J. Comput. Vis. 2026
Visual Localization using Imperfect 3D Models from the Internet · CVPR 2023
Virtual and augmented reality › tracking
camera pose estimation
1.012026
A Guide to Structureless Visual Localization · Int. J. Comput. Vis. 2026
Geometric modeling and processing › surface reconstruction
mesh reconstruction
0.212022
MeshLoc: Mesh-Based Visual Localization · ECCV (22) 2022

Methods — techniques the papers use, named apart from their topics

relative pose estimation · 2.0pose regression · 2.0geometric reasoning · 2.02d-3d matching · 2.02d-2d matching · 2.0structure-based localization · 1.1mesh-based localization · 1.1structure from motion · 0.7
YearPublicationVenuePosition
2026 Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
abstract
Abstract Visual localization is the problem of estimating the camera pose of a given query image within a known scene. Most state-of-the-art localization approaches follow a structure-based paradigm and use 2D-3D matches between pixels in a query image and 3D points in the scene for pose estimation. These approaches assume an accurate 3D model of the scene, which might not always be available, especially if only relatively few images are available to compute the scene representation. In contrast, structure-less methods only use 2D-2D matches and do not require any 3D scene model. However, they are also less accurate than structure-based methods. Although some prior works proposed to combine structure-based and structure-less pose estimation strategies, their practical relevance has not been shown. We analyze combining structure-based and structure-less strategies while exploring how to select between poses obtained from 2D-2D and 2D-3D matches, respectively. We show that combining both strategies improves localization performance in multiple practically relevant scenarios. In particular, the combined strategy allows to gracefully handle degradations in 3D scene model quality.
Vojtech Panek, Torsten Sattler, Zuzana Kukelova
Int. J. Comput. Vis.1
2026 A Guide to Structureless Visual Localization
abstract
Visual localization algorithms, i.e., methods that estimate the camera pose of a query image in a known scene, are core components of many applications, including self-driving cars and augmented / mixed reality systems. State-of-the-art visual localization algorithms are structure-based, i.e., they store a 3D model of the scene and use 2D-3D correspondences between the query image and 3D points in the model for camera pose estimation. While such approaches are highly accurate, they are also rather inflexible when it comes to adjusting the underlying 3D model after changes in the scene. Structureless localization approaches represent the scene as a database of images with known poses and thus offer a much more flexible representation that can be easily updated by adding or removing images. Although there is a large amount of literature on structure-based approaches, there is significantly less work on structureless methods. Hence, this paper is dedicated to providing the, to the best of our knowledge, first comprehensive discussion and comparison of structureless methods. Extensive experiments show that approaches that use a higher degree of classical geometric reasoning generally achieve higher pose accuracy. In particular, approaches based on classical absolute or semi-generalized relative pose estimation outperform very recent methods based on pose regression by a wide margin. Compared with state-of-the-art structure-based approaches, the flexibility of structureless methods comes at the cost of (slightly) lower pose accuracy, indicating an interesting direction for future work.
Vojtech Panek, Qunjie Zhou, Yaqing Ding 0001, Sérgio Agostinho, Zuzana Kukelova, Torsten Sattler, Laura Leal-Taixé
Int. J. Comput. Vis.1
2023 Visual Localization using Imperfect 3D Models from the Internet
abstract
Visual localization is a core component in many applications, including augmented reality (AR). Localization algorithms compute the camera pose of a query image w.r.t. a scene representation, which is typically built from images. This often requires capturing and storing large amounts of data, followed by running Structure-from-Motion (SfM) algorithms. An interesting, and underexplored, source of data for building scene representations are 3D models that are readily available on the Internet, e.g., hand-drawn CAD models, 3D models generated from building footprints, or from aerial images. These models allow to perform visual localization right away without the time-consuming scene capturing and model building steps. Yet, it also comes with challenges as the available 3D models are often imperfect reflections of reality. E.g., the models might only have generic or no textures at all, might only provide a simple approximation of the scene geometry, or might be stretched. This paper studies how the imperfections of these models affect localization accuracy. We create a new benchmark for this task and provide a detailed experimental evaluation based on multiple 3D models per scene. We show that 3D models from the Internet show promise as an easy-to-obtain scene representation. At the same time, there is significant room for improvement for visual localization pipelines. To foster research on this interesting and challenging task, we release our benchmark at v-pnk.github.io/cadloc.
Vojtech Panek, Zuzana Kukelova, Torsten Sattler
CVPR1
2022 MeshLoc: Mesh-Based Visual Localization
Vojtech Panek, Zuzana Kukelova, Torsten Sattler
ECCV (22)1