Ji Zhao 0001

dblp:70/3758-1 · DBLP profile ↗
← Back
62ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-0150-4601ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 8 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 4 first-author · 10 since 2021Systems, architecture and hardware · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2026 Affine Correspondences Between Multi-Camera Systems for Relative Pose Estimation
abstract
We present a novel method to compute the relative pose of multi-camera systems using two affine correspondences (ACs). Existing solutions to the multi-camera relative pose estimation are either restricted to special cases of motion, have too high computational complexity, or require too many point correspondences (PCs). Thus, these solvers impede an efficient or accurate relative pose estimation when applying RANSAC as a robust estimator. This paper shows that the 6DOF relative pose estimation problem using ACs permits a feasible minimal solution, when exploiting the geometric constraints between ACs and multi-camera systems using a special parameterization. We present a problem formulation based on two ACs that encompass two common types of ACs across two views, i.e., inter-camera and intra-camera. Moreover, we exploit a unified and versatile framework for generating 6DOF solvers. Building upon this foundation, we use this framework to address two categories of practical scenarios. First, for the more challenging 7DOF relative pose estimation problem-where the scale transformation of multi-camera systems is unknown-we propose 7DOF solvers to compute the relative pose and scale using three ACs. Second, leveraging inertial measurement units (IMUs), we introduce several minimal solvers for constrained relative pose estimation problems. These include 5DOF solvers with known relative rotation angle, and 4DOF solver with known vertical direction. Experiments on both virtual and real multi-camera systems prove that the proposed solvers are more efficient than the state-of-the-art algorithms, while resulting in a better relative pose accuracy.
Banglei Guan, Ji Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 A Complete Solution to Generalized Relative Pose Estimation From Affine Correspondences
abstract
In recent years, affine correspondences (ACs) have emerged as widely adopted alternative to point correspondences (PCs) in geometric problems in computer vision. An AC is composed of a PC across two different views plus an affine transformation between the small patches around this PC. Prior studies have shown that a single affine correspondence (AC) generally yields three independent constraints for estimating relative pose. This work addresses relative pose estimation in multi-perspective camera systems, a relevant problem given their prevalence in modern technologies such as autonomous vehicles and augmented reality. More specifically, we introduce the first comprehensive suite of minimal solvers for 6DoF relative pose estimation across multiple cameras using only two ACs, which is notably valuable for robust model fitting scenarios. We analyze all possible configurations of two ACs in two views, and present minimal solvers covering all identified minimal cases. We make use of the hidden variable technique to eliminate the translation parameters, and represent rotation using either Cayley parameters or quaternions. We furthermore introduce novel constraints on the generalized relative pose problem that are beneficial in deriving more compact solvers with fewer solutions. Comprehensive experiments on synthetic and real-world data show that the proposed affine correspondence-based solvers are highly effective and computationally efficient.
Banglei Guan, Ji Zhao 0001, Laurent Kneip
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 A Geometric Framework for Absolute Pose and Velocity Estimation With Event Cameras
abstract
Despite the rapid advancements in event-based motion estimation, current geometric methods primarily focus on velocity estimation. However, absolute pose estimation, which is equally crucial for key applications such as robotic navigation and augmented reality, remains relatively underexplored. Consequently, the simultaneous recovery of absolute pose and velocity from event streams remains an open and challenging problem. To address this gap, we propose a geometric framework for absolute pose and velocity estimation by leveraging 3D lines in the scene and the events they trigger. At the core of the framework lie two key geometric constraints: the orthogonality between a 3D line and the normal vector of its corresponding event plane, and the collinearity of an event with the 2D projection of its associated line. Based on these constraints, we present both linear and polynomial solvers for absolute pose estimation. The former enables efficient computation, while the latter provides a globally optimal solution for rotation. For velocity estimation, we develop an efficient linear solver and a more accurate optimization-based solver to recover both angular and linear velocities. Notably, our methods require a minimum of three event-line correspondences to determine the 6-DoF absolute pose or velocities independently. Extensive experiments in simulation and on real-world datasets demonstrate that our methods achieve state-of-the-art performance, with significant improvements in accuracy and computational efficiency compared to existing methods. The demo code is publicly available at https://github.com/Zibin6/EventPoseVelocity.
Zibin Liu, Shunkun Liang, Banglei Guan, Yang Shang, Ji Zhao 0001
IEEE Trans. Image Process.6
2025 Full-DoF Egomotion Estimation for Event Cameras Using Geometric Solvers
abstract
For event cameras, current sparse geometric solvers for egomotion estimation assume that the rotational displacements are known, such as those provided by an IMU. Thus, they can only recover the translational motion parameters. Recovering full-DoF motion parameters using a sparse geometric solver is a more challenging task, and has not yet been investigated. In this paper, we propose several solvers to estimate both rotational and translational velocities within a unified framework. Our method leverages event manifolds induced by line segments. The problem formulations are based on either an incidence relation for lines or a novel coplanarity relation for normal vectors. We demonstrate the possibility of recovering full-DoF egomotion parameters for both angular and linear velocities without requiring extra sensor measurements or motion priors. To achieve efficient optimization, we exploit the Adam framework with a first-order approximation of rotations for quick initialization. Experiments on both synthetic and real-world data demonstrate the effectiveness of our method. The code is available at https://github.com/jizhaox/relpose-event.
Ji Zhao 0001, Banglei Guan, Zibin Liu, Laurent Kneip
CVPR1
2025 Six-Point Method for Multi-Camera Systems with Reduced Solution Space
abstract
Relative pose estimation using point correspondences (PC) is a widely used technique. A minimal configuration of six PCs is required for two views of generalized cameras. In this paper, we present several minimal solvers that use six PCs to compute the 6DOF relative pose of multi-camera systems, including a minimal solver for the generalized camera and two minimal solvers for the practical configuration of two-camera rigs. The equation construction is based on the decoupling of rotation and translation. Rotation is represented by Cayley or quaternion parametrization, and translation can be eliminated by using the hidden variable technique. Ray bundle constraints are found and proven when a subset of PCs relate the same cameras across two views. This is the key to reducing the number of solutions and generating numerically stable solvers. Moreover, all configurations of six-point problems for multi-camera systems are enumerated by the Pólya enumeration theorem. Extensive experiments demonstrate the superior accuracy and efficiency of our solvers compared to state-of-the-art six-point methods. The code is available at https://github.com/jizhaox/relpose-6pt .
Banglei Guan, Ji Zhao 0001, Saibal Mitra, Laurent Kneip
Int. J. Comput. Vis.2
2024 Six-Point Method for Multi-camera Systems with Reduced Solution Space
Banglei Guan, Ji Zhao 0001, Laurent Kneip
ECCV (55)2
2024 Leveraging Enhanced Queries of Point Sets for Vectorized Map Construction
Zihao Liu 0018, Xiaoyu Zhang 0017, Guangwei Liu, Ji Zhao 0001, Ningyi Xu
ECCV (57)4
2024 Enhancing Vectorized Map Perception with Historical Rasterized Maps
Xiaoyu Zhang 0017, Guangwei Liu, Zihao Liu 0018, Ningyi Xu, Yun-Hui Liu 0001, Ji Zhao 0001
ECCV (17)6
2023 Minimal Solvers for Relative Pose Estimation of Multi-Camera Systems using Affine Correspondences
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer
Int. J. Comput. Vis.2
2023 Hong Kong World: Leveraging Structural Regularity for Line-Based SLAM
abstract
Manhattan and Atlanta worlds hold for the structured scenes with only vertical and horizontal dominant directions (DDs). To describe the scenes with additional sloping DDs, a mixture of independent Manhattan worlds seems plausible, but may lead to unaligned and unrelated DDs. By contrast, we propose a novel structural model called Hong Kong world. It is more general than Manhattan and Atlanta worlds since it can represent the environments with slopes, e.g., a city with hilly terrain, a house with sloping roof, and a loft apartment with staircase. Moreover, it is more compact and accurate than a mixture of independent Manhattan worlds by enforcing the orthogonality constraints between not only vertical and horizontal DDs, but also horizontal and sloping DDs. We further leverage the structural regularity of Hong Kong world for the line-based SLAM. Our SLAM method is reliable thanks to three technical novelties. First, we estimate DDs/vanishing points in Hong Kong world in a semi-searching way. We use a new consensus voting strategy for search, instead of traditional branch and bound. This method is the first one that can simultaneously determine the number of DDs, and achieve quasi-global optimality in terms of the number of inliers. Second, we compute the camera pose by exploiting the spatial relations between DDs in Hong Kong world. This method generates concise polynomials, and thus is more accurate and efficient than existing approaches designed for unstructured scenes. Third, we refine the estimated DDs in Hong Kong world by a novel filter-based method. Then we use these refined DDs to optimize the camera poses and 3D lines, leading to higher accuracy and robustness than existing optimization algorithms. In addition, we establish the first dataset of sequential images in Hong Kong world. Experiments showed that our approach outperforms state-of-the-art methods in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Pyojin Kim, Kyungdon Joo, Zhenjun Zhao, Yun-Hui Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Affine Correspondences Between Multi-camera Systems for 6DOF Relative Pose Estimation
Banglei Guan, Ji Zhao 0001
ECCV (32)2
2022 LiDAR-Aided Visual-Inertial Localization with Semantic Maps
abstract
Accurate and robust localization is an essential task for autonomous driving systems. In this paper, we propose a novel 3D LiDAR-aided visual-inertial localization method. Our method fully explores the complementarity of visual and LiDAR observations. On the one hand, the association between semantic features in images and a given semantic map provides constraints for the absolute pose. On the other hand, LiDAR odometry (LO) can provide an accurate and robust 6DOF relative pose. The Error State Kalman Filter (ESKF) framework is exploited to estimate the vehicle pose relative to the semantic map, which fuses the global constraints between the image and the semantic map, the relative pose from the LO, and the raw IMU data. The method achieves centimeter-level localization accuracy in a variety of challenging scenarios. We validate the robustness and accuracy of our method in real-world scenes over 50 km. The experimental results show that the proposed method is able to achieve an average lateral accuracy of 0.059 m and longitudinal accuracy of 0.158 m, which demonstrates the practicality of the proposed system in autonomous driving applications.
Liangliang Pan, Ji Zhao 0001
IROS3
2022 Relative Pose Estimation for Multi-Camera Systems from Point Correspondences with Scale Ratio
abstract
The use of multi-camera systems is becoming more common in self-driving cars, micro aerial vehicles or augmented reality headsets. In order to perform 3D geometric tasks, the accuracy and efficiency of relative pose estimation algorithms are very important for the multi-camera systems, and is catching significant research attention these days. The point coordinates of point correspondences (PCs) obtained from feature matching strategies have been widely used for relative pose estimation. This paper exploits known scale ratios besides the point coordinates, which are also intrinsically provided by scale invariant feature detectors (e.g., SIFT). Two-view geometry of scale ratio associated with the extracted features is derived for multi-camera systems. Thanks to the constraints provided by the scale ratio across two views, the number of PCs needed for relative pose estimation is reduced from 6 to 3. Requiring fewer PCs makes RANSAC-like randomized robust estimation significantly faster. For different point correspondence layouts, four minimal solvers are proposed for typical two-camera rigs. Extensive experiments demonstrate that our solvers have better accuracy than the state-of-the-art ones and outperform them in terms of processing time.
Banglei Guan, Ji Zhao 0001
ACM Multimedia2
2022 Quasi-Globally Optimal and Near/True Real-Time Vanishing Point Estimation in Manhattan World
abstract
Image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim to cluster them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and camera frame. To estimate three degrees of freedom (DOF) of this rotation, state-of-the-art methods are based on either data sampling or parameter search. However, they fail to guarantee high accuracy and efficiency simultaneously. In contrast, we propose a set of approaches that hybridize these two strategies. We first constrain two or one DOF of the rotation by two or one sampled image line. Then we search for the remaining one or two DOF based on branch and bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search achieves quasi-global optimality. Specifically, it guarantees to retrieve the maximum number of inliers on the condition that two or one DOF is constrained. Our hybridization of two-line sampling and one-DOF search can estimate VPs in real time. Our hybridization of one-line sampling and two-DOF search can estimate VPs in near real time. Experiments on both synthetic and real-world datasets demonstrated that our approaches outperform state-of-the-art methods in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Yun-Hui Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 An Efficient Solution to Non-Minimal Case Essential Matrix Estimation
abstract
Finding relative pose between two calibrated images is a fundamental task in computer vision. Given five point correspondences, the classical five-point methods can be used to calculate the essential matrix efficiently. For the case of N ( ) inlier point correspondences, which is called N-point problem, existing methods are either inefficient or prone to local minima. In this paper, we propose a certifiably globally optimal and efficient solver for the N-point problem. First we formulate the problem as a quadratically constrained quadratic program (QCQP). Then a certifiably globally optimal solution to this problem is obtained by semidefinite relaxation. This allows us to obtain certifiably globally optimal solutions to the original non-convex QCQPs in polynomial time. The theoretical guarantees of the semidefinite relaxation are also provided, including tightness and local stability. To deal with outliers, we propose a robust N-point method using M-estimators. Though global optimality cannot be guaranteed for the overall robust framework, the proposed robust N-point method can achieve good performance when the outlier ratio is not high. Extensive experiments on synthetic and real-world datasets demonstrated that our N-point method is 2 ∼ 3 orders of magnitude faster than state-of-the-art methods. Moreover, our robust N-point method outperforms state-of-the-art methods in terms of robustness and accuracy.
Ji Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Relative Pose Estimation With a Single Affine Correspondence
abstract
In this article, we present four cases of minimal solutions for two-view relative pose estimation by exploiting the affine transformation between feature points, and we demonstrate efficient solvers for these cases. It is shown that under the planar motion assumption or with knowledge of a vertical direction, a single affine correspondence is sufficient to recover the relative camera pose. The four cases considered are two-view planar relative motion for calibrated cameras as a closed-form and least-squares solutions, a closed-form solution for unknown focal length, and the case of a known vertical direction. These algorithms can be used efficiently for outlier detection within a RANSAC loop and for initial motion estimation. All the methods are evaluated on both synthetic data and real-world datasets. The experimental results demonstrate that our methods outperform comparable state-of-the-art methods in accuracy with the benefit of a reduced number of needed RANSAC iterations. The source code is released at https://github.com/jizhaox/relative_pose_from_affine.
Banglei Guan, Ji Zhao 0001, Friedrich Fraundorfer
IEEE Trans. Cybern.2
2021 Hybrid Rotation Averaging: A Fast and Robust Rotation Averaging Approach
abstract
We address rotation averaging (RA) and its application to real-world 3D reconstruction. Local optimisation based approaches are the de facto choice, though they only guarantee a local optimum. Global optimisers ensure global optimality in low noise conditions, but they are inefficient and may easily deviate under the influence of outliers or elevated noise levels. We push the envelope of rotation averaging by leveraging the advantages of a global RA method and a local RA method. Combined with a fast view graph filtering as preprocessing, the proposed hybrid approach is robust to outliers. We further apply the proposed hybrid rotation averaging approach to incremental Structure from Motion (SfM), the accuracy and robustness of SfM are both improved by adding the resulting global rotations as regularisers to bundle adjustment. Overall, we demonstrate high practicality of the proposed method as bad camera poses are effectively corrected and drift is reduced.
Ji Zhao 0001, Laurent Kneip
CVPR2
2021 Learning To Identify Correct 2D-2D Line Correspondences on Sphere
abstract
Given a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both structured and unstructured scenes. Instead of geometric constraint, we leverage the spatial regularity on sphere. Specifically, we propose to map line correspondences into vectors tangent to sphere. We use these vectors to encode both angular and positional variations of image lines, which is more reliable and concise than directly using inclinations, midpoints or endpoints of image lines. Neighboring vectors mapped from correct matches exhibit a spatial regularity called local trend consistency, regardless of the type of scenes. To encode this regularity, we design a neural network and also propose a novel loss function that enforces the smoothness constraint of vector field. In addition, we establish a large real-world dataset for image line matching. Experiments showed that our approach outperforms state-of-the-art ones in terms of accuracy, efficiency and robustness, and also leads to high generalization.
Haoang Li, Kai Chen 0028, Ji Zhao 0001, Jiangliu Wang, Pyojin Kim, Zhe Liu 0022, Yun-Hui Liu 0001
CVPR3
2021 Minimal Cases for Computing the Generalized Relative Pose using Affine Correspondences
abstract
We propose three novel solvers for estimating the relative pose of a multi-camera system from affine correspondences (ACs). A new constraint is derived interpreting the relationship of ACs and the generalized camera model. Using the constraint, we demonstrate efficient solvers for two types of motions assumed. Considering that the cameras undergo planar motion, we propose a minimal solution using a single AC and a solver with two ACs to overcome the degenerate case. Also, we propose a minimal solution using two ACs with known vertical direction, e.g., from an IMU. Since the proposed methods require significantly fewer correspondences than state-of-the-art algorithms, they can be efficiently used within RANSAC for outlier removal and initial motion estimation. The solvers are tested both on synthetic data and on real-world scenes from the KITTI odometry benchmark. It is shown that the accuracy of the estimated poses is superior to the state-of-the-art techniques.
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer
ICCV2
2021 Efficient Recovery of Multi-Camera Motion from Two Affine Correspondences
abstract
We propose an efficient method to estimate the relative pose of a multi-camera system from a minimum of two affine correspondences (ACs). Our solution is novel as it computes the 6DOF relative pose by utilizing a first-order rotation approximation. We directly derive a single polynomial based on the constraint between ACs and the generalized camera model. Then a closed-form solution is found analytically and it produces an accurate relative pose estimation efficiently. Benefiting from the low number of exploited correspondences and the speed of the solver, it speeds up robust estimators, e.g. RANSAC, significantly. The proposed method is evaluated both on synthetic data and real-world image sequences from the KITTI benchmark. It is shown that the proposed solver is superior to the state-of-the-art algorithms in terms of accuracy.
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer
ICRA2
2021 Structure Reconstruction Using Ray-Point-Ray Features: Representation and Camera Pose Estimation
abstract
Straight line features have been increasingly utilized in visual SLAM and 3D reconstruction systems. The straight lines’ parameterization, parallel constraint, and coplanar constraint are studied in many recent works. In this paper, we explore the novel intersection constraint of straight lines for structure reconstruction. First, a minimum parameterized representation of ray-point-ray (RPR) structures is proposed to represent the intersection of two straight lines in the 3D space. Second, an efficient solver is designed for the camera pose estimation, which leverages the perpendicularity and intersection of straight lines. Third, we build a stereo visual odometry based on RPR features and evaluate it on the simulation and real datasets. The experimental results verify that the intersection constraints from RPR can effectively improve the accuracy and efficiency of line-based SLAM and reconstruction system.
Yijia He, Xiao Liu 0042, Ji Zhao 0001
ICRA4
2021 Tightly-Coupled Multi-Sensor Fusion for Localization with LiDAR Feature Maps
abstract
Robust and accurate pose estimation in long-term localization is crucial to autonomous driving. In this paper, we dealt with absolute localization with a LiDAR feature map and multi-sensor measurements. We proposed a tightly-coupled fusion method with fixed-lag smoothing. A sliding window of recently maintained states is estimated by minimizing a joint cost function. This cost function includes residuals of global LiDAR registration and relative kinematic constraints from an IMU and wheel encoders. In addition, we enhance the robustness of our method by improving LiDAR registration. To achieve this goal, LiDAR feature maps with a hybrid of geometric and normal distribution features are constructed and exploited. The effectiveness of the proposed method is verified in several challenging test sequences over 200km. The experimental results demonstrate that the proposed method achieves accurate localization and high robustness in challenging scenarios even when the LiDAR observation is degraded.
Liangliang Pan, Kaijin Ji, Ji Zhao 0001
ICRA3
2020 A Certifiably Globally Optimal Solution to Generalized Essential Matrix Estimation
abstract
We present a convex optimization approach for generalized essential matrix (GEM) estimation. The six-point minimal solver for the GEM has poor numerical stability and applies only for a minimal number of points. Existing non-minimal solvers for GEM estimation rely on either local optimization or relinearization techniques, which impedes high accuracy in common scenarios. Our proposed non-minimal solver minimizes the sum of squared residuals by reformulating the problem as a quadratically constrained quadratic program. The globally optimal solution is thus obtained by a semidefinite relaxation. The algorithm retrieves certifiably globally optimal solutions to the original non-convex problem in polynomial time. We also provide the necessary and sufficient conditions to recover the optimal GEM from the relaxed problems. The improved performance is demonstrated over experiments on both synthetic and real multi-camera systems.
Ji Zhao 0001, Wanting Xu, Laurent Kneip
CVPR1
2020 Minimal Solutions for Relative Pose With a Single Affine Correspondence
abstract
In this paper we present four cases of minimal solutions for two-view relative pose estimation by exploiting the affine transformation between feature points and we demonstrate efficient solvers for these cases. It is shown, that under the planar motion assumption or with knowledge of a vertical direction, a single affine correspondence is sufficient to recover the relative camera pose. The four cases considered are two-view planar relative motion for calibrated cameras as a closed-form and a least-squares solution, a closed-form solution for unknown focal length and the case of a known vertical direction. These algorithms can be used efficiently for outlier detection within a RANSAC loop and for initial motion estimation. All the methods are evaluated on both synthetic data and real-world datasets from the KITTI benchmark. The experimental results demonstrate that our methods outperform comparable state-of-the-art methods in accuracy with the benefit of a reduced number of needed RANSAC iterations.
Banglei Guan, Ji Zhao 0001, Friedrich Fraundorfer
CVPR2
2020 Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World
Haoang Li, Pyojin Kim, Ji Zhao 0001, Kyungdon Joo, Zhe Liu 0022, Yun-Hui Liu 0001
ECCV (22)3
2020 Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual Odometry
abstract
Given a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In contrast, we propose a novel approach based on the robust "L2-minimizing estimate" (L2E) loss. We first define a novel cost function by integrating the projection constraint into the L2E loss. Then to efficiently obtain the global minimum of this function, we propose a hybrid strategy of a local optimizer and branch-and-bound. For branch-and-bound, we derive effective function bounds. Our approach can handle high outlier ratios, leading to high robustness. It can run reliably regardless of whether the initial pose is appropriate, providing high generality. Moreover, given a decent initial pose, it is suitable for real-time applications. Experiments on synthetic and real-world datasets showed that our approach outperforms state-of-the-art methods in terms of robustness and/or efficiency.
Haoang Li, Wen Chen 0021, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001
ICRA3
2020 Minimal Case Relative Pose Computation Using Ray-Point-Ray Features
abstract
Corners are popular features for relative pose computation with 2D-2D point correspondences. Stable corners may be formed by two 3D rays sharing a common starting point. We call such elements ray-point-ray (RPR) structures. Besides a local invariant keypoint given by the lines' intersection, their reprojection also defines a corner orientation and an inscribed angle in the image plane. The present paper investigates such RPR features, and aims at answering the fundamental question of what additional constraints can be formed from correspondences between RPR features in two views. In particular, we show that knowing the value of the inscribed angle between the two 3D rays poses additional constraints on the relative orientation. Using the latter enables the solution of the relative pose problem with as few as 3 correspondences across the two images. We provide a detailed analysis of all minimal cases distinguishing between 90-degree RPR-structures and structures with an arbitrary, known inscribed angle. We furthermore investigate the special cases of a known directional correspondence and planar motion, the latter being solvable with only a single RPR correspondence. We complete the exposition by outlining an image processing technique for robust RPR-feature extraction. Our results suggest high practicality in man-made environments, where 90-degree RPR-structures naturally occur.
Ji Zhao 0001, Laurent Kneip, Yijia He, Jiayi Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Robust Estimation of Absolute Camera Pose via Intersection Constraint and Flow Consensus
abstract
Estimating the absolute camera pose requires 3D-to-2D correspondences of points and/or lines. However, in practice, these correspondences are inevitably corrupted by outliers, which affects the pose estimation. Existing outlier removal strategies for robust pose estimation have some limitations. They are only applicable to points, rely on prior pose information, or fail to handle high outlier ratios. By contrast, we propose a general and accurate outlier removal strategy. It can be integrated with various existing pose estimation methods originally vulnerable to outliers, and is applicable to points, lines, and the combination of both. Moreover, it does not rely on any prior pose information. Our strategy has a nested structure composed of the outer and inner modules. First, our outer module leverages our intersection constraint, i.e., the projection rays or planes defined by inliers intersect at the camera center. Our outer module alternately computes the inlier probabilities of correspondences and estimates the camera pose. It can run reliably and efficiently under high outlier ratios. Second, our inner module exploits our flow consensus. The 2D displacement vectors or 3D directed arcs generated by inliers exhibit a common directional regularity, i.e., follow a dominant trend of flow. Our inner module refines the inlier probabilities obtained at each iteration of our outer module. This refinement improves the accuracy and facilitates the convergence of our outer module. Experiments on both synthetic data and real-world images have shown that our method outperforms state-of-the-art approaches in terms of accuracy and robustness.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Yun-Hui Liu 0001
IEEE Trans. Image Process.2
2019 An Evaluation of Feature Matchers for Fundamental Matrix Estimation
Jiawang Bian, Yu-Huan Wu, Ji Zhao 0001, Yun Liu 0011, Le Zhang 0001, Ming-Ming Cheng, Ian D. Reid 0001
BMVC3
2019 Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan World
abstract
The image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The VP estimation can be reformulated as computing the rotation between the Manhattan frame and the camera frame. To compute this rotation, state-of-the-art methods are based on either data sampling or parameter search, and they fail to guarantee the accuracy and efficiency simultaneously. In contrast, we propose to hybridize these two strategies. We first compute two degrees of freedom (DOF) of the above rotation by two sampled image lines, and then search for the optimal third DOF based on the branch-and-bound. Our sampling accelerates our search by reducing the search space and simplifying the bound computation. Our search is not sensitive to noise and achieves quasi-global optimality in terms of maximizing the number of inliers. Experiments on synthetic and real-world images showed that our method outperforms state-of-the-art approaches in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Zhe Liu 0022, Yun-Hui Liu 0001
ICCV2
2019 Leveraging Structural Regularity of Atlanta World for Monocular SLAM
abstract
A wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocular SLAM. First, we robustly cluster image lines. Based on these clusters, we compute the local Atlanta frames in the camera frame by solving polynomial equations. Our method provides the global optimum and satisfies inherent geometric constraints. Second, we define the posterior probabilities to refine the initial clusters and Atlanta frames alternately by the maximum a posteriori estimation. Third, based on multiple local Atlanta frames, we compute the global Atlanta frames in the world frame using Kalman filtering. We optimize rotations by the global alignment and then refine translations and 3D line-based map under the directional constraints. Experiments on both synthesized and real data have demonstrated that our approach outperforms state-of-the-art methods.
Haoang Li, Yazhou Xing, Ji Zhao 0001, Jean-Charles Bazin, Zhe Liu 0022, Yun-Hui Liu 0001
ICRA3
2019 Line-based Absolute and Relative Camera Pose Estimation in Structured Environments
abstract
3D lines in structured environments encode particular regularity like parallelism and orthogonality. We leverage this structural regularity to estimate the absolute and relative camera poses. We decouple the rotation and translation, and propose a novel rotation estimation method. We decompose the absolute and relative rotations and reformulate the problem as computing the rotation from the Manhattan frame to the camera frame. To compute this rotation, we propose an accurate and efficient two-step method. We first estimate its two degrees of freedom (DOF) by two image lines, and then estimate its third DOF by another image line. For these lines, we assume their associated 3D lines are mutually orthogonal, or two 3D lines are parallel to each other and orthogonal to the third. Thanks to our two-step DOF estimation, our absolute and relative pose estimation methods are accurate and efficient. Moreover, our relative pose estimation method relies on weaker assumptions or less correspondences than existing approaches. We also propose a novel strategy to reject outliers and identify dominant directions of the scene. We integrate it into our pose estimation methods, and show that it is more robust than RANSAC. Experiments on synthetic and real-world datasets demonstrated that our methods outperform state-of-the-art approaches.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Wen Chen 0021, Kai Chen 0028, Yun-Hui Liu 0001
IROS2
2019 Locality Preserving Matching
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Xiaojie Guo 0001
Int. J. Comput. Vis.2
2019 LMR: Learning a Two-Class Classifier for Mismatch Removal
abstract
Feature matching, which refers to establishing reliable correspondence between two sets of features, is a critical prerequisite in a wide spectrum of vision-based tasks. Existing attempts typically involve the mismatch removal from a set of putative matches based on estimating the underlying image transformation. However, the transformation could vary with different data. Thus, a pre-defined transformation model is often demanded, which severely limits the applicability. From a novel perspective, this paper casts the mismatch removal into a two-class classification problem, learning a general classifier to determine the correctness of an arbitrary putative match, termed as Learning for Mismatch Removal (LMR). The classifier is trained based on a general match representation associated with each putative match through exploiting the consensus of local neighborhood structures based on a multiple K -nearest neighbors strategy. With only ten training image pairs involving about 8000 putative matches, the learned classifier can generate promising matching results in linearithmic time complexity on arbitrary testing data. The generality and robustness of our approach are verified under several representative supervised learning techniques as well as on different training and testing data. Extensive experiments on feature matching, visual homing, and near-duplicate image retrieval are conducted to reveal the superiority of our LMR over the state-of-the-art competitors.
Jiayi Ma 0001, Xingyu Jiang 0005, Junjun Jiang, Ji Zhao 0001, Xiaojie Guo 0001
IEEE Trans. Image Process.4
2019 Nonrigid Point Set Registration With Robust Transformation Learning Under Manifold Regularization
abstract
This paper solves the problem of nonrigid point set registration by designing a robust transformation learning scheme. The principle is to iteratively establish point correspondences and learn the nonrigid transformation between two given sets of points. In particular, the local feature descriptors are used to search the correspondences and some unknown outliers will be inevitably introduced. To precisely learn the underlying transformation from noisy correspondences, we cast the point set registration into a semisupervised learning problem, where a set of indicator variables is adopted to help distinguish outliers in a mixture model. To exploit the intrinsic structure of a point set, we constrain the transformation with manifold regularization which plays a role of prior knowledge. Moreover, the transformation is modeled in the reproducing kernel Hilbert space, and a sparsity-induced approximation is utilized to boost efficiency. We apply the proposed method to learning motion flows between image pairs of similar scenes for visual homing, which is a specific type of mobile robot navigation. Extensive experiments on several publicly available data sets reveal the superiority of the proposed method over state-of-the-art competitors, particularly in the context of the degenerated data.
Jiayi Ma 0001, Jia Wu 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Quan Z. Sheng
IEEE Trans. Neural Networks Learn. Syst.3
2018 Visual Homing via Guided Locality Preserving Matching
abstract
This study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoramic images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from hundreds of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost true matches without sacrifice in accuracy. To apply our GLPM to the visual homing problem, we develop a method for dense motion flow estimation from sparse feature matches based on Tikhonov regularization. Moreover, the focus-of-contraction/focus-of-expansion is derived to determine homing directions. The effectiveness of our method is demonstrated on a panoramic database in both feature matching and visual homing.
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Yu Zhou 0016, Zheng Wang 0007, Xiaojie Guo 0001
ICRA2
2018 Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector Field
abstract
Estimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect outliers by leveraging the fact that only inliers comply with two effective consensuses, i.e., 3D ray bundle consensus and 2D vector field consensus. Our strategy has a nested structure. First, the outer module utilizes the 3D ray bundle consensus. We define the likelihood based on the probabilistic mixture model and maximize it by the expectation-maximization (EM) algorithm. The inlier probability of each correspondence and the camera pose are determined alternately. Second, the inner module exploits the 2D vector field consensus to refine the probabilities obtained by the outer module. The refinement based on the Bayesian rule facilitates the convergence of the outer module and improves the accuracy of the entire framework. Our strategy can be integrated into various existing camera pose estimation methods which are originally vulnerable to outliers. Experiments on both synthesized data and real images have shown that our approach outperforms state-of-the-art outlier rejection methods in terms of accuracy and robustness.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Jian Yao 0002
IROS2
2018 Guided Locality Preserving Feature Matching for Remote Sensing Image Registration
abstract
Feature matching, which refers to establishing reliable correspondences between two sets of feature points, is a critical prerequisite in feature-based image registration. This paper proposes a simple yet surprisingly effective approach, termed as guided locality preserving matching, for robust feature matching of remote sensing images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost the true matches without sacrifice in accuracy. Experiments on various real remote sensing image pairs demonstrate the generality of our method for handling both rigid and nonrigid image deformations, and it is more than two orders of magnitude faster than the state-of-the-art methods with better accuracy, making it practical for real-time applications.
Jiayi Ma 0001, Junjun Jiang, Huabing Zhou, Ji Zhao 0001, Xiaojie Guo 0001
IEEE Trans. Geosci. Remote. Sens.4
2017 Non-Rigid Point Set Registration with Robust Transformation Estimation under Manifold Regularization
abstract
In this paper, we propose a robust transformation estimation method based on manifold regularization for non-rigid point set registration. The method iteratively recovers the point correspondence and estimates the spatial transformation between two point sets. The correspondence is established based on existing local feature descriptors which typically results in a number of outliers. To achieve an accurate estimate of the transformation from such putative point correspondence, we formulate the registration problem by a mixture model with a set of latent variables introduced to identify outliers, and a prior involving manifold regularization is imposed on the transformation to capture the underlying intrinsic geometry of the input data. The non-rigid transformation is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on both 2D and 3D data demonstrate that our method can yield superior results compared to other state-of-the-arts, especially in case of badly degraded data.
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou
AAAI2
2017 Locality Preserving Matching
abstract
Seeking reliable correspondences between two feature sets is a fundamental and important task in computer vision. This paper attempts to remove mismatches from given putative image feature correspondences. To achieve the goal, an efficient approach, termed as locality preserving matching (LPM), is designed, the principle of which is to maintain the local neighborhood structures of those potential true matches. We formulate the problem into a mathematical model, and derive a closed-form solution with linearithmic time and linear space complexities. More specifically, our method can accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. Experiments on various real image pairs for general feature matching, as well as for visual homing and image retrieval demonstrate the generality of our method for handling different types of image deformations, and it is more than two orders of magnitude faster than state-of-the-art methods in the same range of or better accuracy.
Jiayi Ma 0001, Ji Zhao 0001, Hanqi Guo 0002, Junjun Jiang, Huabing Zhou, Yuan Gao 0015
IJCAI2
2017 Visual homing by robust interpolation for sparse motion flow
abstract
In this paper, we propose a visual homing method by mismatch removal and robust interpolation of sparse motion flows. First, a smoothness prior is proposed and verified by using synthetic and real panoramic images. Then a mismatch removal method is developed for panoramic images based on such smoothness prior. Finally, an interpolation function for the motion flow is obtained simultaneously as a byproduct of mismatch removal, and the focus-of-expansion or focus-of-contraction is used to determine homing directions. The proposed visual homing can be used alone. Also the mismatch removal method can be used as a pre-processing method, and integrated with any other visual homing method which depends on precious keypoints matching. The effectiveness of our method is demonstrated by a panoramic dataset for visual homing.
Ji Zhao 0001, Jiayi Ma 0001
IROS1
2017 Instance Annotation via Optimal BoW for Weakly Supervised Object Localization
abstract
In this paper, we aim at irregular-shape object localization under weak supervision. With over-segmentation, this task can be transformed into multiple-instance context. However, most multiple-instance learning methods only emphasize single most positive instance in a positive bag to optimize bag-level classification, and leads to imprecise or incomplete localization. To address this issue, we propose a scheme for instance annotation, where all of the positive instances are detected by labeling each instance in each positive bag. Inspired by the successful application of bag-of-words (BoW) to feature representation, we leverage it at instance-level to model the distributions of the positive class and negative class, and then incorporate the BoW learning and instance labeling in a single optimization formulation. We also demonstrate that the scheme is well suited to weakly supervised object localization of irregular-shape. Experimental results validate the effectiveness both for the problem of generic instance annotation and for the application of weakly supervised object localization compared to some existing methods.
Liantao Wang, Deyu Meng, Xuelei Hu, Jianfeng Lu 0003, Ji Zhao 0001
IEEE Trans. Cybern.5
2016 Object localization by density-based spatial clustering
abstract
Region search is widely used for object localization in computer vision area. After projecting the score of an image classifier into an image plane, region search aims to find regions that precisely localize desired objects. The popular region search methods, such as efficient subwindow search and efficient region search, usually find regions with maximal score. In this paper, we observe that there is a large score density around a desired object usually. Based on this observation, we proposed a region search method by density-based spatial clustering. The resulted regions of this method can guarantee that their density is above a threshold. Besides, this method has linear time complexity for regularly sampling feature points, which is useful for popular bag-of-words feature representation. We demonstrate its superiority for synthetic data and real image dataset for weakly-supervised localization task.
Ya Lu, Ji Zhao 0001, Jiayi Ma 0001
VCIP2
2016 Nonrigid Feature Matching for Remote Sensing Images via Probabilistic Inference With Global and Local Regularizations
abstract
In this letter, we propose a probabilistic method for the feature matching of remote sensing images which undergo nonrigid transformations. We start by creating a set of putative correspondences based on the feature similarity and then focus on removing outliers from the putative set and estimating the transformation as well. This is formulated as a maximum likelihood estimation of a Bayesian model with latent variables indicating whether matches in the putative set are inliers or outliers. We impose nonparametric global geometrical constraints on the correspondence using Tikhonov regularizers in a reproducing kernel Hilbert space. We also introduce a local geometrical constraint to preserve local structures among neighboring feature points. The problem is solved by using the expectation-maximization algorithm, and the closed-form solution of the transformation is derived in the maximization step. Moreover, a fast implementation based on sparse approximation is given which reduces the method computation complexity to linearithmic without performance sacrifice. Extensive experiments on real remote sensing images demonstrate accurate results of the proposed method which outperforms current state-of-the-art methods, particularly in case of severe outliers.
Huabing Zhou, Jiayi Ma 0001, Changcai Yang, Renfeng Liu, Ji Zhao 0001
IEEE Geosci. Remote. Sens. Lett.6
2016 Feature and Region Selection for Visual Learning
abstract
Visual learning problems, such as object classification and action recognition, are typically approached using extensions of the popular bag-of-words (BoWs) model. Despite its great success, it is unclear what visual features the BoW model is learning. Which regions in the image or video are used to discriminate among classes? Which are the most discriminative visual words? Answering these questions is fundamental for understanding existing BoW models and inspiring better models for visual recognition. To answer these questions, this paper presents a method for feature selection and region selection in the visual BoW model. This allows for an intermediate visualization of the features and regions that are important for visual learning. The main idea is to assign latent weights to the features or regions, and jointly optimize these latent variables with the parameters of a classifier (e.g., support vector machine). There are four main benefits of our approach: 1) our approach accommodates non-linear additive kernels, such as the popular χ(2) and intersection kernel; 2) our approach is able to handle both regions in images and spatio-temporal regions in videos in a unified way; 3) the feature selection problem is convex, and both problems can be solved using a scalable reduced gradient method; and 4) we point out strong connections with multiple kernel learning and multiple instance learning approaches. Experimental results in the PASCAL VOC 2007, MSR Action Dataset II and YouTube illustrate the benefits of our approach.
Ji Zhao 0001, Liantao Wang, Ricardo Silveira Cabral, Fernando De la Torre
IEEE Trans. Image Process.1
2016 Non-Rigid Point Set Registration by Preserving Global and Local Structures
abstract
In previous work on point registration, the input point sets are often represented using Gaussian mixture models and the registration is then addressed through a probabilistic approach, which aims to exploit global relationships on the point sets. For non-rigid shapes, however, the local structures among neighboring points are also strong and stable and thus helpful in recovering the point correspondence. In this paper, we formulate point registration as the estimation of a mixture of densities, where local features, such as shape context, are used to assign the membership probabilities of the mixture model. This enables us to preserve both global and local structures during matching. The transformation between the two point sets is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on both synthesized and real data show the robustness of our approach under various types of distortions, such as deformation, noise, outliers, rotation, and occlusion. It greatly outperforms the state-of-the-art methods, especially when the data is badly degraded.
Jiayi Ma 0001, Ji Zhao 0001, Alan L. Yuille
IEEE Trans. Image Process.2
2015 Density-based region search with arbitrary shape for object localisation
abstract
Region search is widely used for object localisation in computer vision. After projecting the score of an image classifier into an image plane, region search aims to find regions that precisely localise desired objects. The recently proposed region search methods, such as efficient subwindow search and efficient region search, usually find regions with maximal score. For some classifiers and scenarios, the projected scores are nearly all positive or very noisy, then maximising the score of a region results in localising nearly the entire images as objects, or causes localisation results unstable. In this study, the authors observe that the projected scores with large magnitudes are mainly concentrated on or around objects. On the basis of this observation, they propose a region search method for object localisation, named level set maximum‐weight connected subgraph (LS‐MWCS). It localises objects by searching regions by graph mode‐seeking rather than the maximal score. The score density by localised region can be controlled by a parameter flexibly. They also prove an interesting property of the proposed LS‐MWCS, which guarantees that the region with desired density can be found. Moreover, the LS‐MWCS can be efficiently solved by the belief propagation scheme. The effectiveness of the author's method is validated on the problem of weakly‐supervised object localisation. Quantitative results on synthetic and real data demonstrate the superiorities of their method compared to other state‐of‐the‐art methods.
Ji Zhao 0001, Deyu Meng, Jiayi Ma 0001
IET Comput. Vis.1
2015 FastMMD: Ensemble of Circular Discrepancy for Efficient Two-Sample Test
abstract
The maximum mean discrepancy (MMD) is a recently proposed test statistic for the two-sample test. Its quadratic time complexity, however, greatly hampers its availability to large-scale applications. To accelerate the MMD calculation, in this study we propose an efficient method called FastMMD. The core idea of FastMMD is to equivalently transform the MMD with shift-invariant kernels into the amplitude expectation of a linear combination of sinusoid components based on Bochner's theorem and Fourier transform (Rahimi & Recht, 2007). Taking advantage of sampling the Fourier transform, FastMMD decreases the time complexity for MMD calculation from O(N(2)d) to O(LN d), where N and d are the size and dimension of the sample set, respectively. Here, L is the number of basis functions for approximating kernels that determines the approximation accuracy. For kernels that are spherically invariant, the computation can be further accelerated to O(LN log d) by using the Fastfood technique (Le, Sarlós, & Smola, 2013). The uniform convergence of our method has also been theoretically proved in both unbiased and biased estimates. We also provide a geometric explanation for our method, ensemble of circular discrepancy, which helps us understand the insight of MMD and we hope will lead to more extensive metrics for assessing the two-sample test task. Experimental results substantiate that the accuracy of FastMMD is similar to that of MMD and with faster computation and lower variance than existing MMD approximation methods.
Ji Zhao 0001, Deyu Meng
Neural Comput.1
2015 Non-rigid visible and infrared face registration via regularized Gaussian fields criterion
Jiayi Ma 0001, Ji Zhao 0001, Yong Ma 0001, Jinwen Tian
Pattern Recognit.2
2015 Image Feature Matching via Progressive Vector Field Consensus
abstract
In this letter, we propose a simple yet effective approach, named Progressive Vector Field Consensus (PVFC), for addressing the problem of finding more true feature correspondences between images. The key idea is to progressively perform feature matching based on Vector Field Consensus, and hence greatly boost the number of true matches as well as avoid false matches. More specifically, it uses matching results on a small putative correspondence set with high inlier ratio to guide the matching on a large putative correspondence set which probably covers the whole true correspondences. We model the transformation between images in a reproducing kernel Hilbert space, and a sparse approximation is applied to the transformation to avoid high computational complexity. Our results quantitatively show that our PVFC outperforms state-of-the-art methods, both in accuracy and in efficiency. Moreover, the progressive framework is general and can be applied to other cases for robust estimation.
Jiayi Ma 0001, Yong Ma 0001, Ji Zhao 0001, Jinwen Tian
IEEE Signal Process. Lett.3
2015 Robust Feature Matching for Remote Sensing Image Registration via Locally Linear Transforming
abstract
Feature matching, which refers to establishing reliable correspondence between two sets of features (particularly point features), is a critical prerequisite in feature-based registration. In this paper, we propose a flexible and general algorithm, which is called locally linear transforming (LLT), for both rigid and nonrigid feature matching of remote sensing images. We start by creating a set of putative correspondences based on the feature similarity and then focus on removing outliers from the putative set and estimating the transformation as well. We formulate this as a maximum-likelihood estimation of a Bayesian model with hidden/latent variables indicating whether matches in the putative set are outliers or inliers. To ensure the well-posedness of the problem, we develop a local geometrical constraint that can preserve local structures among neighboring feature points, and it is also robust to a large number of outliers. The problem is solved by using the expectation-maximization algorithm (EM), and the closed-form solutions of both rigid and nonrigid transformations are derived in the maximization step. In the nonrigid case, we model the transformation between images in a reproducing kernel Hilbert space (RKHS), and a sparse approximation is applied to the transformation that reduces the method computation complexity to linearithmic. Extensive experiments on real remote sensing images demonstrate accurate results of LLT, which outperforms current state-of-the-art methods, particularly in the case of severe outliers (even up to 80%).
Jiayi Ma 0001, Huabing Zhou, Ji Zhao 0001, Yuan Gao 0015, Junjun Jiang, Jinwen Tian
IEEE Trans. Geosci. Remote. Sens.3
2014 Visual saliency detection using feature activity weighted decorrelation cues
abstract
In this paper, a novel model based on feature activity weighted decorrelation cues is proposed for visual saliency detection in natural images. It consists of two parts: the feature decorrelation and feature information-activity. For the first part, Laplacian sparse coding and low-rank decomposition are used to extract decorrelated features from the scenes. For the second part, Incremental Coding Length is applied to measure the information-activity contained in features, which is then employed to weight the decorrelated features. Finally, visual saliency is estimated through a max pooling strategy. Experimental results on a publicly available benchmark demonstrate the effectiveness of our proposed model with good performance against the state-of-the-art methods.
Shengxiang Qi, Jin-Gang Yu, Ji Zhao 0001, Jie Ma 0003, Jinwen Tian
ICIP3
2014 Weakly supervised object localization via maximal entropy random walk
abstract
In this paper, we investigate the problem of weakly supervised object localization in images. For such a problem, the goal is to predict the locations of objects in test images while the labels of the training images are given at image-level. That means a label only indicates whether an image contains objects or not, but does not provide the exact locations of the objects. We propose to address this problem using Maximal Entropy Random Walk (MERW). Specifically, we first train a linear SVM classifier with the weakly labeled data. Based on bag-of-words feature representation, the response of a region to the linear SVM classifier can be formulated as the sum of the feature-weights within the region. For a test image, by properly constructing a graph on the feature-points, the stationary distribution of a MERW can indicate the region with the densest positive feature-weights, and thus provides a probabilistic object localization. Experiments compared with state-of-the-art methods on two datasets validate the performance of our method.
Liantao Wang, Ji Zhao 0001, Xuelei Hu, Jianfeng Lu 0003
ICIP2
2014 A robust and outlier-adaptive method for non-rigid point registration
Yuan Gao 0015, Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Dazhi Zhang
Pattern Anal. Appl.3
2014 Maximal Entropy Random Walk for Region-Based Visual Saliency
abstract
Visual saliency is attracting more and more research attention since it is beneficial to many computer vision applications. In this paper, we propose a novel bottom-up saliency model for detecting salient objects in natural images. First, inspired by the recent advance in the realm of statistical thermodynamics, we adopt a novel mathematical model, namely, the maximal entropy random walk (MERW) to measure saliency. We analyze the rationality and superiority of MERW for modeling visual saliency. Then, based on the MERW model, we establish a generic framework for saliency detection. Different from the vast majority of existing saliency models, our method is built on a purely region-based strategy, which is able to yield high-resolution saliency maps with well preserved object shapes and uniformly highlighted salient regions. In the proposed framework, the input image is first over-segmented into superpixels, which are taken as the primary units for subsequent procedures, and regional features are extracted. Then, saliency is measured according to two principles, i.e., uniqueness and visual organization, both implemented in a unified approach, i.e., the MERW model based on graph representation. Intensive experimental results on publicly available datasets demonstrate that our method outperforms the state-of-the-art saliency models.
Jin-Gang Yu, Ji Zhao 0001, Jinwen Tian, Yihua Tan
IEEE Trans. Cybern.2
2014 Robust Point Matching via Vector Field Consensus
abstract
In this paper, we propose an efficient algorithm, called vector field consensus, for establishing robust point correspondences between two sets of points. Our algorithm starts by creating a set of putative correspondences which can contain a very large number of false correspondences, or outliers, in addition to a limited number of true correspondences (inliers). Next, we solve for correspondence by interpolating a vector field between the two point sets, which involves estimating a consensus of inlier points whose matching follows a nonparametric geometrical constraint. We formulate this a maximum a posteriori (MAP) estimation of a Bayesian model with hidden/latent variables indicating whether matches in the putative set are outliers or inliers. We impose nonparametric geometrical constraints on the correspondence, as a prior distribution, using Tikhonov regularizers in a reproducing kernel Hilbert space. MAP estimation is performed by the EM algorithm which by also estimating the variance of the prior model (initialized to a large value) is able to obtain good estimates very quickly (e.g., avoiding many of the local minima inherent in this formulation). We illustrate this method on data sets in 2D and 3D and demonstrate that it is robust to a very large number of outliers (even up to 90%). We also show that in the special case where there is an underlying parametric geometrical model (e.g., the epipolar line constraint) that we obtain better results than standard alternatives like RANSAC if a large number of outliers are present. This suggests a two-stage strategy, where we use our nonparametric model to reduce the size of the putative set and then apply a parametric variant of our approach to estimate the geometric parameters. Our algorithm is computationally efficient and we provide code for others to use it. In addition, our approach is general and can be applied to other problems, such as learning with a badly corrupted training data set.
Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Alan L. Yuille, Zhuowen Tu
IEEE Trans. Image Process.2
2013 A Cyclic Weighted Median Method for L1 Low-Rank Matrix Factorization with Missing Entries
abstract
A challenging problem in machine learning, information retrieval and computer vision research is how to recover a low-rank representation of the given data in the presence of outliers and missing entries. The L1-norm low-rank matrix factorization (LRMF) has been a popular approach to solving this problem. However, L1-norm LRMF is difficult to achieve due to its non-convexity and non-smoothness, and existing methods are often inefficient and fail to converge to a desired solution. In this paper we propose a novel cyclic weighted median (CWM) method, which is intrinsically a coordinate decent algorithm, for L1-norm LRMF. The CWM method minimizes the objective by solving a sequence of scalar minimization sub-problems, each of which is convex and can be easily solved by the weighted median filter. The extensive experimental results validate that the CWM method outperforms state-of-the-arts in terms of both accuracy and computational efficiency.
Deyu Meng, Zongben Xu, Lei Zhang 0006, Ji Zhao 0001
AAAI4
2013 Robust Estimation of Nonrigid Transformation for Point Set Registration
abstract
We present a new point matching algorithm for robust nonrigid registration. The method iteratively recovers the point correspondence and estimates the transformation between two point sets. In the first step of the iteration, feature descriptors such as shape context are used to establish rough correspondence. In the second step, we estimate the transformation using a robust estimator called L_2E. This is the main novelty of our approach and it enables us to deal with the noise and outliers which arise in the correspondence step. The transformation is specified in a functional space, more specifically a reproducing kernel Hilbert space. We apply our method to nonrigid sparse image feature correspondence on 2D images and 3D surfaces. Our results quantitatively show that our approach outperforms state-of-the-art methods, particularly when there are a large number of outliers. Moreover, our method of robustly estimating transformations from correspondences is general and has many other applications.
Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Zhuowen Tu, Alan L. Yuille
CVPR2
2013 Regularized vector field learning with sparse approximation for mismatch removal
Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Xiang Bai, Zhuowen Tu
Pattern Recognit.2
2013 Nonrigid Image Deformation Using Moving Regularized Least Squares
abstract
This letter presents an image deformation method based on Moving Regularized Least Squares optimization. The user controls the deformation by simply choosing a set of point handles in the input image, and also the target positions that the source point handles should be deformed to. The deformation function in our method is nonrigid and specified in a functional space, more specifically a reproducing kernel Hilbert space. The proposed method possesses three characteristics: 1) it is able to create detail-preserving and intuitive deformations; 2) the solution of the deformation function has a simple closed-form; 3) it is extremely computationally efficient which can be performed in real-time (less than 0.1 milliseconds per frame for an image of size 500 ×500). We compare our method to a state-of-the-art method which is modeled by rigid transformations; the qualitative and quantitative results demonstrate the benefits of using the nonrigid formulation in aspects of both accuracy and efficiency. Moreover, the proposed method is general and it can be applied to other applications for interpolation.
Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian
IEEE Signal Process. Lett.2
2012 Mismatch removal via coherent spatial mapping
abstract
We propose a method for removing mismatches from given putative point correspondences in image pairs. Our algorithm aims to recover the underlying coherent spatial mapping which related to inliers. The thin-plate spline (TPS) is chosen to parameterize the coherent spatial mapping, and we formulate the solution of it as a maximum likelihood problem. The mismatches could be successfully removed after the EM algorithm, which we used for solving the problem, converges. The quantitative results on various experimental data demonstrate that our method outperforms many state-of-the-art methods. Moreover, the proposed method is also able to handle the case that image pairs contain non-rigid motions.
Jiayi Ma 0001, Ji Zhao 0001, Yu Zhou 0016, Jinwen Tian
ICIP2
2011 A robust method for vector field learning with application to mismatch removing
abstract
We propose a method for vector field learning with outliers, called vector field consensus (VFC). It could distinguish inliers from outliers and learn a vector field fitting for the inliers simultaneously. A prior is taken to force the smoothness of the field, which is based on the Tiknonov regularization in vector-valued reproducing kernel Hilbert space. Under a Bayesian framework, we associate each sample with a latent variable which indicates whether it is an inlier, and then formulate the problem as maximum a posteriori problem and use Expectation Maximization algorithm to solve it. The proposed method possesses two characteristics: 1) robust to outliers, and being able to tolerate 90% outliers and even more, 2) computationally efficient. As an application, we apply VFC to solve the problem of mismatch removing. The results demonstrate that our method outperforms many state-of-the-art methods, and it is very robust.
Ji Zhao 0001, Jiayi Ma 0001, Jinwen Tian, Jie Ma 0003, Dazhi Zhang
CVPR1