Yaqing Ding 0001

dblp:241/5519-1 · DBLP profile ↗
← Back
27ranked-venue papers
16as first author
21since 2021 · last 2026
0000-0002-7448-6686ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 16 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 first-author · 13 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Homography Decomposition Revisited
abstract
Abstract Homography refers to a specific type of transformation that relates two images of the same planar surface taken from different perspectives. Recovering motion parameters from a homography matrix is a classic problem in computer vision. It is important to derive a fast and stable solution to homography decomposition, since it forms a critical component of many vision systems, e . g ., in Structure-from-Motion and visual localization. The current state-of-the-art solvers can be categorized into two types of methods, the numerical procedures based on singular value decomposition (SVD), and the closed-form solution. The SVD-based methods are stable but time-consuming, while the existing closed-form solution is faster but less stable. In this paper, we discuss the homography decomposition problem from a different viewpoint. In contrast to the existing methods which focus on the properties of the homography matrix, we propose a new method that uses three random point correspondences to obtain the motion parameters in closed form. The proposed method is conceptually simple, easy to understand and implement, and has a good geometrical interpretation. This solution can be seen as an alternative to the existing closed-form solution. We also discuss the configurations where the closed-form solutions might be unstable and present a framework for homography decomposition taking into account both the efficiency and stability.
Yaqing Ding 0001, Jian Yang 0003, Zuzana Kukelova
Int. J. Comput. Vis.1
2026 Are Minimal Radial Distortion Solvers Really Necessary for Relative Pose Estimation?
abstract
Estimating the relative pose between two cameras is a fundamental step in many applications such as Structure-from-Motion. The common approach to relative pose estimation is to apply a minimal solver inside a RANSAC loop. Highly efficient solvers exist for pinhole cameras. Yet, (nearly) all cameras exhibit radial distortion. Not modeling radial distortion leads to (significantly) worse results. However, minimal radial distortion solvers are significantly more complex than pinhole solvers, both in terms of run-time and implementation efforts. This paper compares radial distortion solvers with two simple-to-implement approaches that do not use minimal radial distortion solvers: The first approach combines an efficient pinhole solver with sampled radial undistortion parameters, where the sampled parameters are used for undistortion prior to applying the pinhole solver. The second approach uses a state-of-the-art neural network to estimate the distortion parameters rather than sampling them from a set of potential values. Extensive experiments on multiple datasets, and different camera setups, show that complex minimal radial distortion solvers are not necessary in practice. We discuss under which conditions a simple sampling of radial undistortion parameters is preferable over calibrating cameras using a learning-based prior approach. Code and newly created benchmark for relative pose estimation under radial distortion are available at https://github.com/kocurvik/rdnet.
Viktor Kocur, Charalambos Tzamos, Yaqing Ding 0001, Zuzana Berger Haladová, Torsten Sattler, Zuzana Kukelova
Int. J. Comput. Vis.3
2026 A Guide to Structureless Visual Localization
abstract
Visual localization algorithms, i.e., methods that estimate the camera pose of a query image in a known scene, are core components of many applications, including self-driving cars and augmented / mixed reality systems. State-of-the-art visual localization algorithms are structure-based, i.e., they store a 3D model of the scene and use 2D-3D correspondences between the query image and 3D points in the model for camera pose estimation. While such approaches are highly accurate, they are also rather inflexible when it comes to adjusting the underlying 3D model after changes in the scene. Structureless localization approaches represent the scene as a database of images with known poses and thus offer a much more flexible representation that can be easily updated by adding or removing images. Although there is a large amount of literature on structure-based approaches, there is significantly less work on structureless methods. Hence, this paper is dedicated to providing the, to the best of our knowledge, first comprehensive discussion and comparison of structureless methods. Extensive experiments show that approaches that use a higher degree of classical geometric reasoning generally achieve higher pose accuracy. In particular, approaches based on classical absolute or semi-generalized relative pose estimation outperform very recent methods based on pose regression by a wide margin. Compared with state-of-the-art structure-based approaches, the flexibility of structureless methods comes at the cost of (slightly) lower pose accuracy, indicating an interesting direction for future work.
Vojtech Panek, Qunjie Zhou, Yaqing Ding 0001, Sérgio Agostinho, Zuzana Kukelova, Torsten Sattler, Laura Leal-Taixé
Int. J. Comput. Vis.3
2025 Three-view Focal Length Recovery From Homographies
abstract
In this paper, we propose a novel approach for recovering focal lengths from three-view homographies. By examining the consistency of normal vectors between two homographies, we derive new explicit constraints between the focal lengths and homographies using an elimination technique. We demonstrate that three-view homographies provide two additional constraints, enabling the recovery of one or two focal lengths. We discuss four possible cases, including three cameras having an unknown equal focal length, three cameras having two different unknown focal lengths, three cameras where one focal length is known, and the other two cameras have equal or different unknown focal lengths. All the problems can be converted into solving polynomials in one or two unknowns, which can be efficiently solved using Sturm sequence or hidden variable technique. Evaluation using both synthetic and real data shows that the proposed solvers are both faster and more accurate than methods relying on existing two-view solvers. The code and data are available on https://github.com/kocurvik/hf.
Yaqing Ding 0001, Viktor Kocur, Zuzana Berger Haladová, Qianliang Wu, Shen Cai, Jian Yang 0003, Zuzana Kukelova
CVPR1
2025 Practical Solutions to the Relative Pose of Three Calibrated Cameras
abstract
We study the challenging problem of estimating the relative pose of three calibrated cameras from four point correspondences. We propose novel efficient solutions to this problem that are based on the simple idea of using four correspondences to estimate an approximate geometry of the first two views. We model this geometry either as an affine or a fully perspective geometry estimated using one additional approximate correspondence. We generate such an approximate correspondence using a very simple and efficient strategy, where the new point is the mean point of three corresponding input points. The new solvers are efficient and easy to implement, since they are based on existing efficient minimal solvers, i.e., the 4-point affine fundamental matrix, the well-known 5-point relative pose solver, and the P3P solver. Extensive experiments on real data show that the proposed solvers, when properly coupled with local optimization, achieve state-of-the-art results, with the novel solver based on approximate mean-point correspondences being more robust and accurate than the affine-based solver.
Charalambos Tzamos, Viktor Kocur, Yaqing Ding 0001, Daniel Barath, Zuzana Berger Haladová, Torsten Sattler, Zuzana Kukelova
CVPR3
2025 RePoseD: Efficient Relative Pose Estimation With Known Depth Information
Yaqing Ding 0001, Viktor Kocur, Václav Vávra, Zuzana Berger Haladová, Jian Yang 0003, Torsten Sattler, Zuzana Kukelova
ICCV1
2024 Fast Relative Pose Estimation using Relative Depth
abstract
In this paper, we revisit the problem of estimating the relative pose from a sparse set of point-correspondences. For each point-correspondence we also estimate the relative depth, i.e. the relative distance to the scene point in the two images. This yields an additional constraint, allowing us to use fewer matches in RANSAC to generate the pose candidates. In the paper we propose two novel minimal solvers: one for general motion and one for the case of known vertical direction. To obtain the relative depth estimates, we explore using scale estimates obtained from a keypoint detector as well as a neural network that directly predicts the relative depth for a pair of patches. We show in experiments that while our estimates are more noisy compared to the purely point-based solvers, the smaller sample size leads to a significantly reduced runtime in settings with high outlier ratios.
Jonathan Astermark, Yaqing Ding 0001, Viktor Larsson, Anders Heyden
3DV2
2024 Noisy One-Point Homographies are Surprisingly Good
abstract
Two-view homography estimation is a classic and fundamental problem in computer vision. While conceptually simple, the problem quickly becomes challenging when multiple planes are visible in the image pair. Even with correct matches, each individual plane (homography) might have a very low number of inliers when comparing to the set of all correspondences. In practice, this requires a large number of RANSAC iterations to generate a good model hypothesis. The current state-of-the-art methods therefore seek to reduce the sample size, from four point correspondences originally, by including additional information such as keypoint orientation/angles or local affine information. In this work, we continue in this direction and propose a novel one-point solver that leverages different approximate constraints derived from the same auxiliary information. In experiments we obtain state-of-the-art results, with execution time speed-ups, on large benchmark datasets and show that it is more beneficial for the solver to be sample efficient compared to generating more accurate homographies.
Yaqing Ding 0001, Jonathan Astermark, Magnus Oskarsson, Viktor Larsson
CVPR1
2024 Fundamental Matrix Estimation Using Relative Depths
Yaqing Ding 0001, Václav Vávra, Snehal Bhayani, Qianliang Wu, Jian Yang 0003, Zuzana Kukelova
ECCV (71)1
2024 Diff-Reg: Diffusion Model in Doubly Stochastic Matrix Space for Registration Problem
Qianliang Wu, Haobo Jiang, Lei Luo 0001, Jun Li 0027, Yaqing Ding 0001, Jin Xie 0001, Jian Yang 0003
ECCV (65)5
2024 SGNet: Salient Geometric Network for Point Cloud Registration
abstract
Point Cloud Registration (PCR) is a critical and challenging task in computer vision and robotics. One of the primary difficulties in PCR is identifying salient and meaningful points that exhibit consistent semantic and geometric properties across different scans. Previous methods have encountered challenges with ambiguous matching due to the similarity among patch blocks throughout the entire point cloud and the lack of consideration for efficient global geometric consistency. To address these issues, we propose a new framework that includes several novel techniques. Firstly, we introduce a semantic-aware geometric encoder that combines object-level and patch-level semantic information. This encoder significantly improves registration recall by reducing ambiguity in patch-level superpoint matching. Additionally, we incorporate a prior knowledge approach that utilizes an intrinsic shape signature to identify salient points. This enables us to extract the most salient super points and meaningful dense points in the scene. Secondly, we introduce an innovative transformer that encodes High-Order (HO) geometric features. These features are crucial for identifying salient points within initial overlap regions while considering global high-order geometric consistency. We introduce an anchor node selection strategy to optimize this high-order transformer further. By encoding inter-frame triangle or polyhedron consistency features based on these anchor nodes, we can effectively learn high-order geometric features of salient super points. These high-order features are then propagated to dense points and utilized by a Sinkhorn matching module to identify critical correspondences for successful registration. The experiments conducted on the 3DMatch/3DLoMatch and KITTI datasets demonstrate the effectiveness of our method.
Qianliang Wu, Yaqing Ding 0001, Lei Luo 0001, Haobo Jiang, Shuo Gu, Chuanwei Zhou, Jin Xie 0001, Jian Yang 0003
IROS2
2023 Revisiting the P3P Problem
abstract
One of the classical multi-view geometry problems is the so called P3P problem, where the absolute pose of a calibrated camera is determined from three 2D-to-3D correspondences. Since these solvers form a critical component of many vision systems (e.g. in localization and Structure-from-Motion), there have been significant effort in developing faster and more stable algorithms. While the current state-of-the-art solvers are both extremely fast and stable, there still exist configurations where they break down. In this paper we algebraically formulate the problem as finding the intersection of two conics. With this formulation we are able to analytically characterize the real roots of the polynomial system and employ a tailored solution strategy for each problem instance. The result is a fast and stable solver, that is able to correctly solve cases where competing methods might fail. Our experimental evaluation shows that we outperform the current state-of-the-art methods both in terms of speed and success rate.
Yaqing Ding 0001, Jian Yang 0003, Viktor Larsson, Carl Olsson, Kalle Åström
CVPR1
2023 Minimal Solutions to Generalized Three-View Relative Pose Problem
abstract
For a generalized (or non-central) camera model, the minimal problem for two views of six points has efficient solvers. However, minimal problems of three views with four points and three views of six lines have not yet been explored and solved, despite the efforts from the computer vision community. This paper develops the formulations of these two minimal problems and shows how state-of-the-art GPU implementations of Homotopy Continuation solver can be used effectively. The proposed methods are evaluated on both synthetic and real datasets, demonstrating that they are fast, accurate and that they improve on structure from motion estimations, when employed in an hypothesis and test setting.
Yaqing Ding 0001, Chiang-Heng Chien, Viktor Larsson, Kalle Åström, Benjamin B. Kimia
ICCV1
2023 Graph Matching Optimization Network for Point Cloud Registration
abstract
Point Cloud Registration is a fundamental and challenging problem in 3D computer vision. Recent works often utilize geometric structure features in downsampled points (patches) to seek correspondences, then propagate these sparse patch correspondences to the dense level in the corresponding patches' neighborhood. However, they neglect the explicit global scale rigid constraint at the dense level point matching. We claim that the explicit isometry-preserving constraint in the dense level on a global scale is also important for improving feature representation in the training stage. To this end, we propose a Graph Matching Optimization based Network (GMONet for short), which utilizes the graph-matching optimizer to explicitly exert the isometry preserving constraints in the point feature training to improve the point feature representation. Specifically, we exploit a partial graph-matching optimizer to enhance the super point (i.e., down-sampled key points) features and a full graph-matching optimizer to improve the dense level point features in the overlap region. Meanwhile, we leverage the inexact proximal point method and the mini-batch sampling technique to accelerate these two graph-matching optimizers. Given high discriminative point features in the evaluation stage, we utilize the RANSAC approach to estimate the transformation between the scanned pairs. The proposed method has been evaluated on the 3DMatch/3DLoMatch and the KITTI datasets. The experimental results show that our method performs competitively compared to state-of-the-art baselines.
Qianliang Wu, Yaqi Shen, Haobo Jiang, Guofeng Mei, Yaqing Ding 0001, Lei Luo 0001, Jin Xie 0001, Jian Yang 0003
IROS5
2022 Relative Pose from a Calibrated and an Uncalibrated Smartphone Image
abstract
In this paper, we propose a new minimal and a non-minimal solver for estimating the relative camera pose together with the unknown focal length of the second camera. This configuration has a number of practical benefits, e.g., when processing large-scale datasets. Moreover, it is resistant to the typical degenerate cases of the traditional six-point algorithm. The minimal solver requires four point correspondences and exploits the gravity direction that the built-in IMU of recent smart devices recover. We also propose a linear solver that enables estimating the pose from a larger-than-minimal sample extremely efficiently which then can be improved by, e.g., bundle adjustment. The methods are tested on 35654 image pairs from publicly available real-world and new datasets. When combined with a recent robust estimator, they lead to results superior to the traditional solvers in terms of rotation, translation and focal length accuracy, while being notably faster.
Yaqing Ding 0001, Daniel Barath, Jian Yang 0003, Zuzana Kukelova
CVPR1
2022 Globally Optimal Relative Pose Estimation for Multi-Camera Systems with Known Gravity Direction
abstract
Multiple-camera systems have been widely used in self-driving cars, robots, and smartphones. In addition, they are typically also equipped with IMUs (inertial measurement units). Using the gravity direction extracted from the IMU data, the y-axis of the body frame of the multi-camera system can be aligned with this common direction, reducing the original three degree-of-freedom(DOF) relative rotation to a single DOF one. This paper presents a novel globally optimal solver to compute the relative pose of a generalized camera. Existing optimal solvers based on LM (Levenberg-Marquardt) method or SDP (semidefinite program) are either iterative or have high computational complexity. Our proposed optimal solver is based on minimizing the algebraic residual objective function. According to our derivation, using the least-squares algorithm, the original optimization problem can be converted into a system of two polynomials with only two variables. The proposed solvers have been tested on synthetic data and the KITTI benchmark. The experimental results show that the proposed methods have competitive robustness and accuracy compared with the existing state-of-the-art solvers.
Qianliang Wu, Yaqing Ding 0001, Xinlei Qi, Jin Xie 0001, Jian Yang 0003
ICRA2
2022 Homography-Based Minimal-Case Relative Pose Estimation With Known Gravity Direction
abstract
In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with known gravity direction. This case is relevant to smart phones, tablets, and other camera-IMU (Inertial measurement unit) systems which have accelerometers to measure the gravity vector. We explore the rank-1 constraint on the difference between the euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and semi-calibrated (unknown focal length) problems. Based on the hidden variable technique, we convert the problems to the polynomial eigenvalue problems, and derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with the existing 6- and 7-point solvers, including results with smart phone images.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Image Stitching with Locally Shared Rotation Axis
abstract
We consider the problem of stitching image sequences with cameras undergoing pure rotational motion. We leverage the assumption of a locally constant rotation axis, i.e., neighboring frames have a shared but unknown rotation axis. This assumption holds in many common image capturing scenarios, e.g., panoramic sweeping motions. Using this additional constraint, we develop techniques for three-view camera rotation estimation; a minimal solver for the two-view estimation with a known rotation axis; and a globally optimal robust estimator for the two-view case. We show on publicly available datasets that the proposed methods lead to camera rotation estimation superior to the state-of-the-art in terms of accuracy with comparable run-time. The source code will be made available.
Daniel Barath, Yaqing Ding 0001, Zuzana Kukelova, Viktor Larsson
3DV2
2021 Globally Optimal Relative Pose Estimation With Gravity Prior
abstract
Smartphones, tablets and camera systems used, e.g., in cars and UAVs, are typically equipped with IMUs (inertial measurement units) that can measure the gravity vector accurately. Using this additional information, the y-axes of the cameras can be aligned, reducing their relative orientation to a single degree-of-freedom. With this assumption, we propose a novel globally optimal solver, minimizing the algebraic error in the least squares sense, to estimate the relative pose in the over-determined case. Based on the epipolar constraint, we convert the optimization problem into solving two polynomials with only two unknowns. Also, a fast solver is proposed using the first-order approximation of the rotation. The proposed solvers are compared with the state-of-the-art ones on four real-world datasets with approx. 50000 image pairs in total. Moreover, we collected a dataset, by a smartphone, consisting of 10933 image pairs, gravity directions and ground truth 3D reconstructions. The source code and dataset are available at https://github.com/yaqding/opt_pose_gravity
Yaqing Ding 0001, Daniel Barath, Jian Yang 0003, Hui Kong 0001, Zuzana Kukelova
CVPR1
2021 Minimal Solutions for Panoramic Stitching Given Gravity Prior
abstract
When capturing panoramas, people tend to align their cameras with the vertical axis, i.e., the direction of gravity. Moreover, modern devices, e.g. smartphones and tablets, are equipped with an IMU (Inertial Measurement Unit) that can measure the gravity vector accurately. Using this prior, the y-axes of the cameras can be aligned or assumed to be already aligned, reducing the relative orientation to 1-DOF (degree of freedom). Exploiting this assumption, we propose new minimal solutions to panoramic stitching of images taken by cameras with coinciding optical centers, i.e. undergoing pure rotation. We consider six practical camera configurations, from fully calibrated ones up to a camera with unknown fixed or varying focal length and with or without radial distortion. The solvers are tested both on synthetic scenes, on more than 500k real image pairs from the Sun360 dataset, and from scenes captured by us using two smartphones equipped with IMUs. The new solvers have similar or better accuracy than the state-of-the-art ones and outperform them in terms of processing time.
Yaqing Ding 0001, Daniel Barath, Zuzana Kukelova
ICCV1
2021 A general elimination strategy for camera motion estimation
abstract
Camera motion estimation, such as relative pose estimation and absolute pose estimation, are fundamental problems in computer vision and robotics. To obtain the motion parameters, classical methods rely on studying the properties of the geometric matrices, e.g., rotation matrix, essential matrix, homography matrix. The well known five-point algorithm was successfully derived using the singular constraint and trace constraints on the essential matrix. However, finding all the algebraic constraints is not always trivial for some recent problems. In this paper, we propose a simple and general technique to find complete algebraic constraints so that we can derive efficient algorithms. We show that using the quaternion to formulate the rotation matrix we can eliminate any unknowns from the original equations and obtain constraints on the rest of the unknowns based on Gröbner basis. We demonstrate that this approach can be applied to almost all the camera motion estimation and show its improvement compared to the existing methods. Further more, based on this elimination technique, we exploit new constraints for the relative pose estimation with gravity prior, and derive a new globally optimal algorithm to this problem. We compare our algorithm with the state-of-the-art methods on both synthetic and real-world data, and show the benefits including accuracy and efficiency.
Yaqing Ding 0001, Yingna Su, Cheng-Zhong Xu 0001, Jian Yang 0003, Hui Kong 0001
ICRA1
2020 Homography-Based Egomotion Estimation Using Gravity and SIFT Features
Yaqing Ding 0001, Daniel Barath, Zuzana Kukelova
ACCV (1)1
2020 Minimal Solutions to Relative Pose Estimation From Two Views Sharing a Common Direction With Unknown Focal Length
abstract
We propose minimal solutions to relative pose estimation problem from two views sharing a common direction with unknown focal length. This is relevant for cameras equipped with an IMU (inertial measurement unit), e.g., smart phones, tablets. Similar to the 6-point algorithm for two cameras with unknown but equal focal lengths and 7-point algorithm for two cameras with different and unknown focal lengths, we derive new 4- and 5-point algorithms for these two cases, respectively. The proposed algorithms can cope with coplanar points, which is a degenerate configuration for these 6- and 7-point counterparts. We present a detailed analysis and comparisons with the state of the art. Experimental results on both synthetic data and real images from a smart phone demonstrate the usefulness of the proposed algorithms.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
CVPR1
2020 A two-step approach to Lidar-Camera calibration
abstract
Autonomous vehicles and robots are typically equipped with Lidar and camera. Hence, calibrating the Lidar-camera system is of extreme importance for ego-motion estimation and scene understanding. In this paper, we propose a two-step approach (coarse + fine) for the external calibration between a camera and a multiple-line Lidar. First, a new closed-form solution is proposed to obtain the initial calibration parameters. We compare our solution with the state-of-the-art SVD-based algorithm, and show the benefits of both the efficiency and stability. With the initial calibration parameters, the ICP-based calibration framework is used to register the point clouds which extracted from the camera and Lidar coordinate frames, respectively. Our method has been applied to two Lidar-camera systems: an HDL-64E Lidar-camera system, and a VLP-16 Lidar-camera system. Experimental results demonstrate that our method achieves promising performance and higher accuracy than two open-source methods.
Yingna Su, Yaqing Ding 0001, Jian Yang 0003, Hui Kong 0001
ICPR2
2020 An efficient solution to the relative pose estimation with a common direction
abstract
In this paper, we propose an efficient solution to the calibrated camera motion estimation with a common direction. This case is relevant to smart phones, tablets, and other camera-IMU (Inertial measurement unit) systems, which have accelerometers to measure the gravity direction. We can align one of the axes of the camera with this common direction so that the relative rotation between the views reduces to only 1DOF (degree of freedom). This allows us to use only three point correspondences for relative pose estimation. Unlike previous work, we derive new constraints on the simplified essential matrix using an elimination strategy based on Gröbner basis. In this case, computing the coefficients of these constraints require less computation and we only need to solve a polynomial eigenvalue problem. We show detailed analyses and comparisons against the existing 3-point algorithms, with satisfactory results obtained.
Yaqing Ding 0001, Jian Yang 0003, Hui Kong 0001
ICRA1
2019 An Efficient Solution to the Homography-Based Relative Pose Problem With a Common Reference Direction
abstract
In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with a common reference direction. We explore the rank-1 constraint on the difference between the Euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and partially calibrated (unknown focal length) problems. We derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with existing 6 and 7-point solvers, including results with smart phone images.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICCV1
2018 Visual Odometry for Indoor Mobile Robot by Recognizing Local Manhattan Structures
Zhixing Hou, Yaqing Ding 0001, Ying Wang 0007, Hui Kong 0001
ACCV (5)2