EDBT 2026 Demo / reviewers in the wild / expert
Lipu Zhou
dblp:120/0668
· DBLP profile ↗
18ranked-venue papers
15as first author
7since 2021 · last 2025
0000-0003-2148-393XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 13 first-author · 6 since 2021Systems, architecture and hardware · 7 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Novel Iterative Solution to the Perspective-$n$-Point Problem via Cost Function ApproximationabstractThe Perspective-$n$-Point (P$n$P) problem, which estimates the camera pose through$N$2D/3D point correspondences, has been extensively studied. Although minimizing the reprojection cost is regarded as the gold standard for solving the P$n$P problem, this cost lacks an analytic solution, leading previous works to focus on developing simpler costs. State-of-the-art P$n$P solutions are generally considered to be close to the gold-standard solution. However, this perception is based on limited experimental setups. Our extensive evaluations show that these solutions generally deviate from the gold-standard solution as the depth range of 3D points increases. This paper investigates two noise models of the P$n$P problem and provides a unified, accurate, and efficient solution. The main contributions of this paper are threefold. First, we propose an efficient initialization method that compresses$ 2N$constraints to three quadratic equations for rotation using principal component analysis (PCA). Second, we prove that our initialization algorithm provides a solution to the P3P problem, making it applicable to the full range$N \geq 3$of the P$n$P problem. Third, we propose a novel iterative algorithm that approximates reprojection residuals using second-order polynomials and determines the optimal step size analytically, ensuring fast convergence. Extensive experiments on synthetic and real data demonstrate that our algorithm outperforms state-of-the-art methods in terms of accuracy and robustness, while achieving comparable efficiency. Lipu Zhou, Zhenzhong Wei, Xu Wang 0056 |
IEEE Trans. Robotics | 1 |
| 2023 | Efficient Second-Order Plane AdjustmentabstractPlanes are generally used in 3D reconstruction for depth sensors, such as RGB-D cameras and LiDARs. This paper focuses on the problem of estimating the optimal planes and sensor poses to minimize the point-to-plane distance. The resulting least-squares problem is referred to as plane adjustment (PA) in the literature, which is the counterpart of bundle adjustment (BA) in visual reconstruction. Iterative methods are adopted to solve these least-squares problems. Typically, Newton's method is rarely used for a large-scale least-squares problem, due to the high computational complexity of the Hessian matrix. Instead, methods using an approximation of the Hessian matrix, such as the Levenberg-Marquardt (LM) method, are generally adopted. This paper adopts the Newton's method to efficiently solve the PA problem. Specifically, given poses, the optimal plane have a close-form solution. Thus we can eliminate planes from the cost function, which significantly reduces the number of variables. Furthermore, as the optimal planes are functions of poses, this method actually ensures that the optimal planes for the current estimated poses can be obtained at each iteration, which benefits the convergence. The difficulty lies in how to efficiently compute the Hessian matrix and the gradient of the resulting cost. This paper provides an efficient solution. Empirical evaluation shows that our algorithm outperforms the state-of-the-art algorithms. Lipu Zhou |
CVPR | 1 |
| 2023 | Ada3D : Exploiting the Spatial Redundancy with Adaptive Inference for Efficient 3D Object DetectionabstractVoxel-based methods have achieved state-of-the-art performance for 3D object detection in autonomous driving. However, their significant computational and memory costs pose a challenge for their application to resource-constrained vehicles. One reason for this high resource consumption is the presence of a large number of redundant background points in Lidar point clouds, resulting in spatial redundancy in both 3D voxel and BEV map representations. To address this issue, we propose an adaptive inference framework called Ada3D, which focuses on reducing the spatial redundancy to compress the model’s computational and memory cost. Ada3D adaptively filters the redundant input, guided by a lightweight importance predictor and the unique properties of the Lidar point cloud. Additionally, we maintain the BEV features’ intrinsic sparsity by introducing the Sparsity Preserving Batch Normalization. With Ada3D, we achieve 40% reduction for 3D voxels and decrease the density of 2D BEV feature maps from 100% to 20% without sacrificing accuracy. Ada3D reduces the model computational and memory cost by 5×, and achieves 1.52× / 1.45× end-to-end GPU latency and 1.5× / 4.5× GPU peak memory optimization for the 3D and 2D backbone respectively. Tianchen Zhao, Xuefei Ning, Ke Hong, Zhongyuan Qiu, Pu Lu, Yali Zhao, Linfeng Zhang 0001, Lipu Zhou, Guohao Dai 0001, Huazhong Yang, Yu Wang 0002 |
ICCV | 8 |
| 2023 | Efficient Visual-Inertial Navigation with Point-Plane MapabstractAccurate and real-time global pose estimation relative to a global prior map is indispensable in many applications, such as logistics with micro aerial vehicles and Augmented Reality. Supposed that a pure sparse 3D point map can provide a structureless representation of the environment, then generating a point-plane prior map can further model the environment topology and offer global constraints for an accurate localization. To implement this, we propose a filter-based, large-scale visual-inertial odometry system, termed PPM-VIO, which utilizes a point-plane map to correct the cumulative drift. Our system, detecting coplanar information from sparse point clouds with semantic information, achieves accurate online plane matching via geometric constraints, semantic constraints, and descriptor constraints. To improve the localization performance, we effectively integrate and formulate the global planar measurements and points measurements in a filter-based estimator. The effectiveness of the proposed method is extensively validated on real-world datasets collected in different scenarios. Experimental results demonstrate that, rather than using the point map alone, leveraging the plane information in the prior map can yield better trajectory estimates and broaden the effective scope of the prior map in different scenes. Kefei Ren, Lipu Zhou, Xiaoming Lang, Yinian Mao, Guoquan Huang 0003 |
ICRA | 4 |
| 2023 | Efficient Bundle Adjustment for Coplanar Points and LinesabstractBundle adjustment (BA) is a well-studied fundamental problem in the robotics and vision community. In man-made environments, coplanar points and lines are ubiquitous. However, the number of works on bundle adjustment with coplanar points and lines is relatively small. This paper focuses on this special BA problem, referred to as$\pi-\mathbf{BA}$. For a point or a line on a plane, we derive a new constraint to describe the relationship among two poses and the plane, called$\pi$-constraint. We distribute$\pi$-constraints into different groups. Each group is called a$\pi$-factor. We prove that, with some simple preprocessing, the computational complexity associated with a$\pi$-factor in the Levenberg-Marquardt (LM) algorithm is$O(1)$, independent of the number of$\pi$-constraints packed into the$\pi$-factor. In$\pi-\mathbf{BA}, \pi$-factors replace original reprojection errors. One problem is how to divide$\pi$-constraints into$\pi$-factors. Different strategies may result in different numbers of$\pi$-factors, which in turn affects the efficiency. It is difficult to get the optimal division. We present a greedy algorithm to overcome this problem. Experimental results verify that our algorithm can significantly accelerate the computation. Lipu Zhou, Jiacheng Liu 0008, Fengguang Zhai, Pan Ai, Kefei Ren, Yinian Mao, Guoquan Huang 0003, Ziyang Meng 0001, Michael Kaess |
ICRA | 1 |
| 2022 | EDPLVO: Efficient Direct Point-Line Visual OdometryabstractThis paper introduces an efficient direct visual odometry (VO) algorithm using points and lines. Pixels on lines are generally adopted in direct methods. However, the original photometric error is only defined for points. It seems difficult to extend it to lines. In previous works, the collinear constraints for points on lines are either ignored [1] or introduce heavy computational load into the resulting optimization system [2]. This paper extends the photometric error for lines. We prove that the 3D points of the points on a 2D line are determined by the inverse depths of the endpoints of the 2D line, and derive a closed-form solution for this problem. This property can significantly reduce the number of variables to speed up the optimization, and can make the collinear constraint exactly satisfied. Furthermore, we introduce a two-step method to further accelerate the optimization, and prove the convergence of this method. The experimental results show that our algorithm outperforms the state-of-the-art direct VO algorithms. Lipu Zhou, Guoquan Huang 0003, Yinian Mao, Shengze Wang 0002, Michael Kaess |
ICRA | 1 |
| 2021 | π-LSAM: LiDAR Smoothing and Mapping With PlanesabstractThis paper introduces a real-time dense planar LiDAR SLAM system, named π-LSAM, for the indoor environment. The widely used LiDAR odometry and mapping (LOAM) framework [1] does not include bundle adjustment (BA) and generates a low fidelity tracking pose. This paper seeks to overcome these drawbacks for the indoor environment. Specifically, we use the plane as the landmark, and introduce plane adjustment (PA) as our back-end to jointly optimize planes and keyframe poses. We present the π-factor to significantly reduce the computational complexity of PA. In addition, we introduce an efficient loop detection algorithm based on the RANSAC framework using planes. In the front-end, our algorithm performs global registration in real time. To achieve this performance, we maintain the local-to-global point-to-plane correspondences scan by scan, so that we only need a small local KD-tree to establish the data association between a LiDAR scan and the global planes, rather than a large global KD-tree used in previous works. With this local-to-global data association, our algorithm directly identifies planes in a LiDAR scan, and yields an accurate and globally consistent pose. Experimental results show that our algorithm significantly outperforms the state-of-the-art LOAM variant, LeGO-LOAM [2], and our algorithm achieves real time. Lipu Zhou, Shengze Wang 0002, Michael Kaess |
ICRA | 1 |
| 2020 | A Fast and Accurate Solution for Pose Estimation from 3D CorrespondencesabstractEstimating pose from given 3D correspondences, including point-to-point, point-to-line and point-to-plane correspondences, is a fundamental task in computer vision with many applications. We present a fast and accurate solution for the least-squares problem of this task. Previous works mainly focus on studying the way to find the global minimizer of the least-squares problem. However, existing works that show the ability to achieve the global minimizer are still unsuitable for real-time applications. Furthermore, as one of contributions of this paper, we prove that there exist ambiguous configurations for any number of lines and planes. These configurations have several solutions in theory, which makes the correct solution may come from a local minimizer when the data are with noise. Previous works based on convex optimization which is unable to find local minimizers do not work in the ambiguous configuration. Our algorithm is efficient and able to reveal local minimizers. We employ the Cayley-Gibbs-Rodriguez (CGR) parameterization of the rotation to derive a general rational cost for the three cases of 3D correspondences. The main contribution of this paper is to solve the first-order optimality conditions of the least-squares problem, which are of a complicated rational form. The central idea of our algorithm is to introduce some intermediate unknowns to simplify the problem. Extensive experimental results show that our algorithm is more stable than previous algorithms when the number N of correspondences is small. Besides, when N is large, our algorithm achieves the same accuracy as the state-of-the-art algorithm [1], but our algorithm is about 7 times faster than [1] in real applications. Lipu Zhou, Shengze Wang 0002, Michael Kaess |
ICRA | 1 |
| 2020 | An Efficient Planar Bundle Adjustment AlgorithmabstractThis paper presents an efficient algorithm for the least-squares problem using the point-to-plane cost, which aims to jointly optimize depth sensor poses and plane parameters for 3D reconstruction. We call this least-squares problem Planar Bundle Adjustment (PBA), due to the similarity between this problem and the original Bundle Adjustment (BA) in visual reconstruction. As planes ubiquitously exist in the man-made environment, they are generally used as landmarks in SLAM algorithms for various depth sensors. PBA is important to reduce drift and improve the quality of the map. However, directly adopting the well-established BA framework in visual reconstruction will result in a very inefficient solution for PBA. This is because a 3D point only has one observation at a camera pose. In contrast, a depth sensor can record hundreds of points in a plane at a time, which results in a very large nonlinear least-squares problem even for a small-scale space. The main contribution of this paper is an efficient solution for the PBA problem using the point-to-plane cost. We introduce a reduced Jacobian matrix and a reduced residual vector, and prove that they can replace the original Jacobian matrix and residual vector in the generally adopted Levenberg-Marquardt (LM) algorithm. This significantly reduces the computational cost. Besides, when planes are combined with other features for 3D reconstruction, the reduced Jacobian matrix and residual vector can also replace the corresponding parts derived from planes. Our experimental results show that our algorithm can significantly reduce the computational time compared to the solution using the traditional BA framework. In addition, our algorithm is faster, more accurate, and more robust to initialization errors compared to the start-of-the-art solution using the plane-to-plane cost [3]. Lipu Zhou, Daniel Koppel, Hui Ju, Frank Steinbrücker, Michael Kaess |
ISMAR | 1 |
| 2019 | A Robust and Efficient Algorithm for the PnL Problem Using Algebraic Distance to Approximate the Reprojection DistanceabstractThis paper proposes a novel algorithm to solve the pose estimation problem from 2D/3D line correspondences, known as the Perspective-n-Line (PnL) problem. It is widely known that minimizing the geometric distance generally results in more accurate results than minimizing an algebraic distance. However, the rational form of the reprojection distance of the line yields a complicated cost function, which makes solving the first-order optimality conditions infeasible. Furthermore, iterative algorithms based on the reprojection distance are time-consuming for a large-scale problem. In contrast to previous works which minimize a cost function based on an algebraic distance that may not approximate the reprojection distance of the line, we design two simple algebraic distances to gradually approximate the reprojection distance. This speeds up the computation, and maintains the robustness of the geometric distance. The two algebraic distances result in two polynomial cost functions, which can be efficiently solved. We directly solve the first-order optimality conditions of the first problem with a novel hidden variable method. This algorithm makes use of the specific structure of the resulting polynomial system, therefore it is more stable than the general Gröbner basis polynomial solver. Then, we minimize the second polynomial cost function by the damped Newton iteration, starting from the solution of the first cost function. Experimental results show that the first step of our algorithm is already superior to the state-of-the-art algorithms in terms of accuracy and applicability, and faster than the algorithms based on Gröbner basis polynomial solver. The second step yields comparable results to the results from minimizing the reprojection distance, but is much more efficient. For speed, our algorithm is applicable to real-time applications. Lipu Zhou, Montiel Abello, Michael Kaess |
AAAI | 1 |
| 2019 | An Efficient and Accurate Algorithm for the Perspecitve-n-Point ProblemabstractIn this paper, we address the problem of pose estimation from N 2D/3D point correspondences, known as the Perspective-n-Point (PnP) problem. Although many solutions have been proposed, it is hard to optimize both computational complexity and accuracy at the same time. In this paper, we propose an accurate and simultaneously efficient solution to the PnP problem. Previous PnP algorithms generally involve two sets of unknowns including the depth of each pixel and the pose of the camera. Our formulation does not involve the depth of each pixel. By introducing some intermediate variables, this formulation leads to a fourth degree polynomial cost function with 3 unknowns that only involves the rotation. In contrast to previous works, we do not address this minimization problem by solving the first-order optimality conditions using the off-the-shelf Gröbner basis method, as the Gröbner basis method may encounter numeric problems. Instead, we present a method based on linear system null space analysis to provide a robust initial estimation for a Newton iteration. Experimental results demonstrate that our algorithm is comparable to the start-of-the-art algorithms in terms of accuracy, and the speed of our algorithm is among the fastest algorithms. Lipu Zhou, Michael Kaess |
IROS | 1 |
| 2018 | A Stable Algebraic Camera Pose Estimation for Minimal Configurations of 2D/3D Point and Line Correspondences
Lipu Zhou, Jiamin Ye, Michael Kaess |
ACCV (4) | 1 |
| 2018 | Automatic Extrinsic Calibration of a Camera and a 3D LiDAR Using Line and Plane CorrespondencesabstractIn this paper, we address the problem of extrinsic calibration of a camera and a 3D Light Detection and Ranging (LiDAR) sensor using a checkerboard. Unlike previous works which require at least three checkerboard poses, our algorithm reduces the minimal number of poses to one by combining 3D line and plane correspondences. Besides, we prove that parallel planar targets with parallel boundaries provide the same constraints in our algorithm. This allows us to place the checkerboard close to the LiDAR so that the laser points better approximate the target boundary without loss of generality. Moreover, we present an algorithm to estimate the similarity transformation between the LiDAR and the camera for the applications where only the correspondences between laser points and pixels are concerned. Using a similarity transformation can simplify the calibration process since the physical size of the checkerboard is not needed. Meanwhile, estimating the scale can yield a more accurate result due to the inevitable measurement errors of the checkerboard size and the LiDAR intrinsic scale factor that transforms the LiDAR measurement to the metric measurement. Our algorithm is validated through simulations and experiments. Compared to the plane-only algorithms, our algorithm can obtain more accurate result by fewer number of poses. This is beneficial to the large-scale commercial application. Lipu Zhou, Zimo Li, Michael Kaess |
IROS | 1 |
| 2018 | Detection and Recognition of Traffic Planar Objects Using Colorized Laser Scan and Perspective Distortion RectificationabstractReliable detection and recognition of planar objects including traffic sign, street sign, and road surface in dynamic cluttered natural scenes are a big challenge for self-driving cars. In this paper, we propose a comprehensive method for planar object detection and recognition. First, the data association of LIDAR and camera is set up to acquire colorized laser scans, which simultaneously contain both color and geometrical information. Second, we combine three color spaces of RGB, HSV, and CIE L*a*b* with laser reflectivity as an aggregation-based feature vector. Third, the 3-D geometrical characteristics of planar objects that contain planarity, size, and aspect ratio are exploited to further reduce false alarm. Fourth, in order to increase robustness to any viewpoint variation, we present a new virtual camera-based rectification method to synthesize fronto-parallel views of refined object descriptors in 3-D space. Finally, experimental results achieved under a variety of challenging conditions show that integration of color space aggregation and laser reflectivity is superior to individuals. Specifically, the proposed perspective distortion rectification method remarkably eliminates false recognition error by 45.5%. Overall, the detection rate of our comprehensive method has up to 95.87% and the recognition rate even reaches 95.07% for traffic signs ranging within 100 m, with about 33.25 ms average running time per frame. Zhidong Deng, Lipu Zhou |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Visual Relocalization Using Long-Short Term Memory Fully Convolutional NetworkabstractThis paper tackles the problem of camera relocalization using a single image. We formulate this problem as a regression problem and directly learn the mapping between am image and its pose by a new Deep Neural Network (DNN) architecture in an end-to-end manner. The main contribution of this work is the proposed network, called Long-Short Term Memory Fully Convolutional Network (LSTMFCN), which consists of a Fully Convolutional Network (FCN) as the feature extractor and a Long-Short Term Memory (LSTM) as the pooling layer to aggregate information across the image. In contrast to the previous DNN-based relocalization algorithms that only consider a small patch of the image, the new network has a much larger receptive field. This can avoid the aperture problem and can make it more robust to partial occlusion and moving objects. Besides, we adopt the shortcut connection to fuse features from different layers, and introduce the Error of Average Pose (EAP) into the cost function. Moreover, we show that our algorithm can be viewed as a keyframe-based relocalization algorithm, if we treat the training samples as keyframes. But unlike the traditional keyframe-based algorithms whose computational time and storage will increase as the size of the scene enlarges, our algorithm has constant computational time and storage. We investigate different network structures and parameter settings, and compare our algorithm with the previous algorithms by experiments. The experimental results show that our algorithm significantly outperforms the state-of-the-art DNN-based algorithm and achieves real time. Lipu Zhou |
ICTAI | 1 |
| 2015 | Learn to Solve Algebra Word Problems Using Quadratic ProgrammingabstractThis paper presents a new algorithm to automatically solve algebra word problems.Our algorithm solves a word problem via analyzing a hypothesis space containing all possible equation systems generated by assigning the numbers in the word problem into a set of equation system templates extracted from the training data.To obtain a robust decision surface, we train a log-linear model to make the margin between the correct assignments and the false ones as large as possible.This results in a quadratic programming (QP) problem which can be efficiently solved.Experimental results show that our algorithm achieves 79.7% accuracy, about 10% higher than the state-of-the-art baseline (Kushman et al., 2014). Lipu Zhou, Shuaixiang Dai |
EMNLP | 1 |
| 2013 | Fusing laser point cloud and visual image at data level using a new reconstruction algorithmabstractCamera and LIDAR provide complementary information for robots to perceive the environment. In this paper, we present a system to fuse laser point cloud and visual information at the data level. Generally, cameras and LIDARs mounted on the unmanned ground vehicle have different viewports. Some objects which are visible to a LIDAR may become invisible to a camera. This will result in false depth assignment for the visual image and incorrect colorization for laser points. The inputs of the system are a color image and the corresponding LIDAR data. Coordinates of 3D laser points are first transformed into the camera coordinate system. Points outside the camera viewing volume are clipped. A new algorithm is proposed to recreate the underlying object surface of the potentially visible laser points as quadrangle mesh by exploiting the structure of the LIDAR as a priori. False edge is eliminated by constraining the angle between the laser scan trace and the radial direction of a given laser point, and quadrangles with non-consistent normal are pruned. In addition, the missing laser points are solved to avoid large holes in the reconstructed mesh. At last z-buffer algorithm is used to work for occlusion reasoning. Experimental results show that our algorithm outperforms the previous one. It can assign correct depth information to the visual image and provide the exact color to each laser point which is visible to the camera. Lipu Zhou |
Intelligent Vehicles Symposium | 1 |
| 2012 | Extrinsic calibration of a camera and a lidar based on decoupling the rotation from the translationabstractIn this paper, we propose a novel robust algorithm for the extrinsic calibration of a camera and a lidar. This algorithm utilizes checkerboard as a calibration object. Since the interaction between the estimation errors of the plane parameters obtained from checkerboard images downgrades the quality of extrinsic calibration results, a new geometric constraint is presented to decouple the rotation from the translation so as to reduce the effect of such an interaction. Weights that represent uncertainty of the unit normal vector to the checkerboard plane are introduced to totally evaluate the quality of each pair of image and lidar scan. Furthermore, we analyze the configuration of checkerboard pose and give a formula that is used to assess the configuration. We compare the proposed algorithm with the previous ones. Simulation and experimental results show that our algorithm is able to achieve more accurate extrinsic parameters than the existing algorithms. Meanwhile, we also design experiments to validate the effectiveness and efficiency of the presented weight and the assessment formula. Lipu Zhou, Zhidong Deng |
Intelligent Vehicles Symposium | 1 |