EDBT 2026 Demo / reviewers in the wild / expert
Xiang Gao 0009
dblp:14/3881-9
· DBLP profile ↗
27ranked-venue papers
12as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incremental rotation averaging revisited
Xiang Gao 0009, Hainan Cui, Yangdong Liu, Shuhan Shen |
Pattern Recognit. | 1 |
| 2025 | BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with TransformerabstractIn this paper, we present BWFormer, a novel Transformerbased model for building wireframe reconstruction from airborne LiDAR point cloud. The problem is solved in a ground-up manner here by detecting the building corners in 2D, lifting and connecting them in 3D space afterwards with additional data augmentation. Due to the 2.5D characteristic of the airborne LiDAR point cloud, we simplify the problem by projecting the points on the ground plane to produce a 2D height map. With the height map, a heat map is first generated with pixel-wise corner likelihood to predict the possible 2D corners. Then, 3D corners are predicted by a Transformer-based network with extra height embedding initialization. This 2D-to-3D corner detection strategy reduces the search space significantly. To recover the topological connections among the corners, edges are finally predicted from the height map with the proposed edge attention mechanism, which extracts holistic features and preserves local details simultaneously. In addition, due to the limited datasets in the field and the irregularity of the point clouds, a conditional latent diffusion model for LiDAR scanning simulation is utilized for data augmentation. BW-Former surpasses other state-of-the-art methods, especially in reconstruction completeness. Our code is available at: https : //github.com/3dv-casia/BWformer/. Lingjie Zhu, Hanqiao Ye, Shangfeng Huang, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen |
CVPR | 5 |
| 2025 | BPN: Building Pointer Network for Satellite Imagery Building Contour ExtractionabstractExtracting structured building contours from satellite imagery plays an important role in many geospatial tasks. However, it still remains a challenge due to the high cost of manual labeling, and models trained on simple polygons show poor generalization on buildings with more complex shapes. To deal with this, we propose a novel neural network called building pointer network (BPN) in this letter, which builds upon a recurrent neural network (RNN) architecture that integrates visual and geometric signals with an input-focused attention mechanism, making it more general for various shape complexity. Given an RGB satellite image, the model first uses a convolutional neural network (CNN) to obtain the set of key points for each building. Then, the coordinates of the key points and their image features are fused and fed into the RNN which ultimately predicts the index of the building corners sequentially. Results show that our method has good generalization ability for building data with complex shapes, provided that a dataset with relatively simple shapes is used as the training set. Lingjie Zhu, Zexiao Xie, Xiang Gao 0009, Shuhan Shen |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | PolyRoom: Room-Aware Transformer for Floorplan Reconstruction
Lingjie Zhu, Hanqiao Ye, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen |
ECCV (50) | 5 |
| 2024 | IRAv3+: Hierarchical Incremental Rotation Averaging via Multiple Connected Dominating SetsabstractFocusing on the difficulty of absolute rotation globalization of large-scale rotation averaging problem, a novel hierarchical pipeline, termed as IRAv3+, based on multiple Connected Dominating Sets (CDSs) is proposed in this paper. Specifically, the proposed method not only obtains the graph clusters for local rotation averaging like other cluster-based methods, but also generate a subset via connected dominating set extraction, which is served as a reference for rotation globalization. To facilitate the rotation globalization, two key techniques are proposed: 1) to provide a more reliable global reference, instead of a single CDS, multiple CDSs are randomly selected and united; 2) to give a more accurate local-to-global alignment estimation, instead of using the relative rotation measurements of the sharing edges between local clusters and global reference, the absolute rotations of common vertices between them are involved. Experiments on the 1DSfM dataset demonstrate the effectiveness of the proposed IRAv3+ and its advantages over the existing cluster-based rotation averaging methods and other state of the arts. Xiang Gao 0009, Hainan Cui, Wantao Huang, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Recent Advances in Conventional and Deep Learning-Based Depth Completion: A SurveyabstractDepth completion aims to recover pixelwise depth from incomplete and noisy depth measurements with or without the guidance of a reference RGB image. This task attracted considerable research interest due to its importance in various computer vision-based applications, such as scene understanding, autonomous driving, 3-D reconstruction, object detection, pose estimation, trajectory prediction, and so on. As the system input, an incomplete depth map is usually generated by projecting the 3-D points collected by ranging sensors, such as LiDAR in outdoor environments, or obtained directly from RGB-D cameras in indoor areas. However, even if a high-end LiDAR is employed, the obtained depth maps are still very sparse and noisy, especially in the regions near the object boundaries, which makes the depth completion task a challenging problem. To address this issue, a few years ago, conventional image processing-based techniques were employed to fill the holes and remove the noise from the relatively dense depth maps obtained by RGB-D cameras, while deep learning-based methods have recently become increasingly popular and inspiring results have been achieved, especially for the challenging situation of LiDAR-image-based depth completion. This article systematically reviews and summarizes the works related to the topic of depth completion in terms of input modalities, data fusion strategies, loss functions, and experimental settings, especially for the key techniques proposed in deep learning-based multiple input methods. On this basis, we conclude by presenting the current status of depth completion and discussing several prospects for its future research directions. Zexiao Xie, Xiaoxuan Yu, Xiang Gao 0009, Kunqian Li, Shuhan Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | IRAv3: Hierarchical Incremental Rotation Averaging on the FlyabstractWe present IRAv3, which is built upon the state-of-the-art rotation averaging method, IRA++, to push this fundamental task in 3D computer vision one step further. The key observation of this letter lies in that during IRA++, the community detection-based Epipolar-geometry Graph (EG) clustering is preemptive and permanent, which is not relevant to the follow-up rotation averaging task and limits the upper bound of absolute rotation estimation accuracy. In this letter, however, the EG clustering is performed along with the cluster-wise absolute rotation estimation, i.e. instead of pre-determination, the affiliation of each vertex to which EG cluster is determined “on the fly”, and the EG clustering finishes until all the vertices find the clusters they belong to, together with their absolute rotations estimated (in the local coordinate systems of the clusters they attached). By this way, a rotation averaging-targeted and -friendly EG clustering is obtained, which facilitates the rotation averaging task in turn. Experiments on both 1DSfM and KITTI odometry datasets demonstrate the effectiveness of our proposed IRAv3 on large-scale rotation averaging problems and its advantages over its previous works (IRA and IRA++) and other state of the arts. Xiang Gao 0009, Hainan Cui, Zexiao Xie, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Beyond Single Reference for Training: Underwater Image Enhancement via Comparative LearningabstractDue to the wavelength-dependent light absorption and scattering, the raw underwater images are usually inevitably degraded. Underwater image enhancement (UIE) is of great importance for underwater observation and operation. Data-driven methods, such as deep learning-based UIE approaches, tend to be more applicable to real underwater scenarios. However, the training of deep models is limited by the extreme scarcity of underwater images with enhancement references, resulting in their poor performance in dynamic and diverse underwater scenes. As an alternative, enhancement reference achieved by volunteer voting alleviate the sample shortage to some extent. Since such artificially acquired references are not veritable ground truth, they are far from complete and accurate to provide correct and rich supervision for the enhancement model training. Beyond training with single reference, we propose the first comparative learning framework for UIE problem, namely CLUIE-Net, to learn from multiple candidates of enhancement reference. This new strategy also supports semi-supervised learning mode. Besides, we propose a regional quality-superiority discriminative network (RQSD-Net) as an embedded quality discriminator for the CLUIE-Net. Comprehensive experiments demonstrate the effectiveness of RQSD-Net and the comparative learning strategy for UIE problem. The code, models and new dataset RQSD-UI are available at: https://justwj.github.io/CLUIE-Net.html/. Kunqian Li, Qi Qi 0008, Xiang Gao 0009, Liqin Zhou 0001, Dalei Song |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | MCSfM: Multi-Camera-Based Incremental Structure-From-MotionabstractFully perceiving the surrounding world is a vital capability for autonomous robots. To achieve this goal, a multi-camera system is usually equipped on the data collecting platform and the structure from motion (SfM) technology is used for scene reconstruction. However, although incremental SfM achieves high-precision modeling, it is inefficient and prone to scene drift in large-scale reconstruction tasks. In this paper, we propose a tailored incremental SfM framework for multi-camera systems, where the internal relative poses between cameras can not only be calibrated automatically but also serve as an additional constraint to improve the system robustness. Previous multi-camera based modeling work has mainly focused on stereo setups or multi-camera systems with known calibration information, but we allow arbitrary configurations and only require images as input. First, one camera is selected as the reference camera, and the other cameras in the multi-camera system are denoted as non-reference cameras. Based on the pose relationship between the reference and non-reference camera, the non-reference camera pose can be derived from the reference camera pose and internal relative poses. Then, a two-stage multi-camera based camera registration module is proposed, where the internal relative poses are computed first by local motion averaging, and then the rigid units are registered incrementally. Finally, a multi-camera based bundle adjustment is put forth to iteratively refine the reference camera and the internal relative poses. Experiments demonstrate that our system achieves higher accuracy and robustness on benchmark data compared to the state-of-the-art SfM and SLAM (simultaneous localization and mapping) methods. Hainan Cui, Xiang Gao 0009, Shuhan Shen |
IEEE Trans. Image Process. | 2 |
| 2023 | Vehicle-Borne Multi-Sensor Temporal-Spatial Pose Globalization via Cross-Domain Data AssociationabstractLarge-scale urban scene 3D mapping has urgent demands and wide applications in many areas, where sensor pose globalization remains its fundamental problem and critical step. As the street-view images and vehicle-borne Light Detection And Ranging (LiDAR) points contain complementary advantages in urban scene 3D mapping, it is desirable to make the most of both to facilitate this task. Most existing methods make strong assumptions of strict synchronization, and even further, exact calibration between the vehicle-borne cameras and LiDARs, which are hard to guarantee in practice. To deal with this, we propose a novel pipeline for vehicle-borne camera and LiDAR temporal and spatial pose globalization with the guidance of Global Navigation Satellite System/Inertial Measurement Unit (GNSS/IMU), where both of the assumptions on strict synchronization and exact calibration are loosened. Specifically, the global poses of both cameras and LiDARs are first initialized by leveraging GNSS/IMU signals and multi-sensor pre-calibrations, and then refined by a global optimization scheme. To perform the global pose optimization, image-based, LiDAR-based, and cross-domain data association and constraint construction are conducted. Among them, the cross-domain ones, which are achieved by LiDAR point projection, image feature back-projection, and spatial point association, provide key clues for associating these two kinds of data with significant differences. Comprehensive experiments on both of a self-collected and the KITTI Odometry datasets demonstrate the effectiveness of our proposed method on multi-sensor pose globalization for large-scale urban scene 3D mapping. Xiang Gao 0009, Dongdong Tao, Yuqian Liu, Zexiao Xie, Shuhan Shen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | IRA++: Distributed Incremental Rotation AveragingabstractBy observing that the recently presented Incremental Rotation Averaging (IRA) suffers from drifting and efficiency problems in large-scale situations, it is upgraded in this work to possess stronger scalability in both accuracy and efficiency based on the thought of divide and conquer. This upgraded version is termed as IRA++. Specifically, the original Epipolar-geometry Graph (EG) is clustered into several sub-graphs and inner-rotation averaging is distributedly performed in each of them with IRA at first. Then, the relative rotation between each pair of inner-sub-EG coordinate systems is distributedly estimated by a voting-based single rotation averaging method. Subsequently, IRA-based inter-rotation averaging is performed to obtain the absolute rotation of each inner-sub-EG coordinate system. And finally, the absolute rotations of all the cameras in the original EG are globally aligned and optimized to get the final rotation averaging result. Comprehensive evaluations on the 1DSfM, Campus, and San Francisco datasets demonstrate the advantages of our proposed IRA++ over IRA and several other state-of-the-art rotation averaging methods in both efficiency and accuracy, especially the accuracy in noise-polluted and efficiency in large-scale situations. Xiang Gao 0009, Lingjie Zhu, Hainan Cui, Zexiao Xie, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Incremental Translation AveragingabstractTranslation averaging is known to be more difficult than rotation averaging due to scale ambiguity, estimation sensitivity, and solution uncertainty. Existing approaches have exposed their limitations in terms of accuracy, robustness, simplicity, or efficiency. To tackle this tough problem, a simple yet effective translation averaging pipeline, termed as Incremental Translation Averaging (ITA), is proposed in this paper. It combines the advantages of high accuracy and robustness in incremental parameter estimation pipeline and the advantages of high simplicity and efficiency in global motion averaging approach. Unlike the traditional translation averaging methods which estimate all the absolute camera locations simultaneously and suffer from inaccuracy in parameter estimation and incompleteness in scene reconstruction, our ITA computes them novelly in an incremental way with higher accuracy and robustness. Thanks to the introduction of incremental parameter estimation thought into the translation averaging pipeline, 1) our ITA is robust to measurement outliers and accurate in parameter estimation; and 2) our ITA is simple and efficient because of its less dependency on complicated optimization, carefully-designed preprocessing, or additional information. Comprehensive evaluations on the 1DSfM dataset demonstrate the effectiveness of our ITA and its advantages over several state-of-the-art translation averaging approaches. Xiang Gao 0009, Lingjie Zhu, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Robust Camera Translation Estimation via Rank EnforcementabstractCamera translation averaging, aiming to recover the global camera locations from a given set of camera translation directions, is a challenging problem for Structure from Motion (SfM) in the field of computer vision, largely due to the fact that the given relative translation directions from a set of noisy essential matrices are generally of low accuracy. To tackle this problem, we first reveal a novel but a simple property of the camera translation matrix consisting of all the pairwise camera translations among an arbitrary set of cameras that the rank of this translation matrix is always smaller or equal to 4. Then, by explicitly enforcing this rank property, a novel translation estimation method for computing global camera locations is proposed, called TERE. Moreover, to further improve the performances of the explored TERE in the two aspects of accuracy and speed, an iterative batch-based translation estimation method is proposed, called B-TERE, where a small-scale batch of cameras is selected without replacement from the given set of cameras according to a simple camera selection strategy at each iterative step, and the locations of the selected cameras are estimated by the proposed TERE accordingly. Extensive experimental results on various datasets demonstrate that our proposed methods could achieve better performances in comparison to several state-of-the-art methods. Qiulei Dong, Xiang Gao 0009, Hainan Cui, Zhanyi Hu |
IEEE Trans. Cybern. | 2 |
| 2022 | SGUIE-Net: Semantic Attention Guided Underwater Image Enhancement With Multi-Scale PerceptionabstractDue to the wavelength-dependent light attenuation, refraction and scattering, underwater images usually suffer from color distortion and blurred details. However, due to the limited number of paired underwater images with undistorted images as reference, training deep enhancement models for diverse degradation types is quite difficult. To boost the performance of data-driven approaches, it is essential to establish more effective learning mechanisms that mine richer supervised information from limited training sample resources. In this paper, we propose a novel underwater image enhancement network, called SGUIE-Net, in which we introduce semantic information as high-level guidance via region-wise enhancement feature learning. Accordingly, we propose semantic region-wise enhancement module to better learn local enhancement features for semantic regions with multi-scale perception. After using them as complementary features and feeding them to the main branch, which extracts the global enhancement features on the original image scale, the fused features bring semantically consistent and visually superior enhancements. Extensive experiments on the publicly available datasets and our proposed dataset demonstrate the impressive performance of SGUIE-Net. The code and proposed dataset are available at https://trentqq.github.io/SGUIE-Net.html. Qi Qi 0008, Kunqian Li, Haiyong Zheng, Xiang Gao 0009, Guojia Hou, Kun Sun 0002 |
IEEE Trans. Image Process. | 4 |
| 2021 | Incremental Rotation Averaging
Xiang Gao 0009, Lingjie Zhu, Zexiao Xie, Hongmin Liu 0001, Shuhan Shen |
Int. J. Comput. Vis. | 1 |
| 2021 | Urban Scene LOD Vectorized Modeling From Photogrammetry MeshesabstractUrban scene modeling is a challenging task for the photogrammetry and computer vision community due to its large scale, structural complexity, and topological delicacy. This paper presents an efficient multistep modeling framework for large-scale urban scenes from aerial images. It takes aerial images and a textured 3D mesh model generated by an image-based modeling system as the input and outputs compact polygon models with semantics at different levels of detail (LODs). Based on the key observation that urban buildings usually have piecewise planar rooftops and vertical walls, we propose a segment-based modeling method, which consists of three major stages: scene segmentation, roof contour extraction, and building modeling. By combining the deep neural network predictions with geometric constraints of the 3D mesh, the scene is first segmented into three classes. Then, for each building mesh, the 2D line segments are detected and used to slice the ground into polygon cells, followed by assigning each cell a roof plane via a MRF optimization. Finally, the LOD model is obtained by extruding cells to their corresponding planes. Compared with direct modeling in 3D space, we transform the mesh into a uniform 2D image grid representation and most of the modeling work is performed in 2D space, which has the advantages of low computational complexity and high robustness. In addition, our method doesn't require any global prior, such as the Manhattan or Atlanta world assumption, making it flexible to model scenes with different characteristics and complexity. Experiments on both single buildings and large-scale urban scenes demonstrate that by combining 2D photometric with 3D geometric information, the proposed algorithm is robust and efficient in urban scene LOD vectorized modeling compared with the state-of-the-art approaches. Jiali Han, Lingjie Zhu, Xiang Gao 0009, Zhanyi Hu, Liyang Zhou, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Image Process. | 3 |
| 2020 | Hierarchical RANSAC-Based Rotation AveragingabstractIn this letter, we present a novel rotation averaging pipeline, which is performed in a hierarchical manner. Unlike the traditional rotation averaging methods which focus on designing robust loss function to get rid of the impacts of the relative rotation outliers, here the outliers are detected and filtered by leveraging the well-known robust model estimation procedure, RANdom SAmple Consensus (RANSAC). During the RANSAC process, the minimal set is randomly sampled by random tree spanning on the Epipolar-geometry Graph (EG). As the RANSAC estimation result is sensitive to the size of minimal set, the EG is clustered into several sub-graphs, and the inner- and inter-cluster RANSAC-based rotation averaging are performed hierarchically. In addition, both random generation and optimal selection of the minimal set are performed in a weighted manner to make the rotation averaging pipeline more robust. Ablation studies and comparison experiments on the 1DSfM and San Francisco (SNF) datasets demonstrate the effectiveness of our proposed method. Xiang Gao 0009, Jiazheng Luo, Kunqian Li, Zexiao Xie |
IEEE Signal Process. Lett. | 1 |
| 2020 | Complete Scene Reconstruction by Merging Images and Laser ScansabstractImage based modeling and laser scanning are two commonly used approaches in large-scale architectural scene reconstruction nowadays. In order to generate a complete scene reconstruction, an effective way is to completely cover the scene using ground and aerial images, supplemented by laser scanning on certain regions with low textures and complicated structures. Thus, the key issue is to accurately calibrate cameras and register laser scans in a unified framework. To this end, we proposed a three-step pipeline for complete scene reconstruction by merging images and laser scans. First, images are captured around the architecture in a multiview and multiscale way and are feed into a structure-from-motion (SfM) pipeline to generate SfM points. Then, based on the SfM result, the laser scanning locations are automatically planned by considering textural richness, structural complexity of the scene and spatial layout of the laser scans. Finally, the images and laser scans are accurately merged in a coarse-to-fine manner. Experimental evaluations on two ancient Chinese architecture datasets demonstrate the effectiveness of our proposed complete scene reconstruction pipeline. Xiang Gao 0009, Shuhan Shen, Lingjie Zhu, Tianxin Shi, Zhiheng Wang 0001, Zhanyi Hu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Visual Localization Using Sparse Semantic 3D MapabstractAccurate and robust visual localization under a wide range of viewing condition variations including season and illumination changes, as well as weather and day-night variations, is the key component for many computer vision and robotics applications. Under these conditions, most traditional methods would fail to locate the camera. In this paper we present a visual localization algorithm that combines structure-based method and image-based method with semantic information. Given semantic information about the query and database images, the retrieved images are scored according to the semantic consistency of the 3D model and the query image. Then the semantic matching score is used as weight for RANSAC's sampling and the pose is solved by a standard PnP solver. Experiments on the challenging long-term visual localization benchmark dataset demonstrate that our method has significant improvement compared with the state-of-the-arts. Tianxin Shi, Shuhan Shen, Xiang Gao 0009, Lingjie Zhu |
ICIP | 3 |
| 2019 | Ground and aerial meta-data integration for localization and reconstruction: A review
Xiang Gao 0009, Shuhan Shen, Zhanyi Hu, Zhiheng Wang 0001 |
Pattern Recognit. Lett. | 1 |
| 2019 | Multi-source data-based 3D digital preservation of largescale ancient chinese architecture: A case reportabstractThe 3D digitalization and documentation of ancient Chinese architecture is challenging because of architectural complexity and structural delicacy. To generate complete and detailed models of this architecture, it is better to acquire, process, and fuse multi-source data instead of single-source data. In this paper, we describe our work on 3D digital preservation of ancient Chinese architecture based on multisource data. We first briefly introduce two surveyed ancient Chinese temples, Foguang Temple and Nanchan Temple. Then, we report the data acquisition equipment we used and the multi-source data we acquired. Finally, we provide an overview of several applications we conducted based on the acquired data, including ground and aerial image fusion, image and LiDAR (light detection and ranging) data fusion, and architectural scene surface reconstruction and semantic modeling. We believe that it is necessary to involve multi-source data for the 3D digital preservation of ancient Chinese architecture, and that the work in this paper will serve as a heuristic guideline for the related research communities. Xiang Gao 0009, Hainan Cui, Lingjie Zhu, Tianxin Shi, Shuhan Shen |
Virtual Real. Intell. Hardw. | 1 |
| 2018 | Large Scale Urban Scene Modeling from MVS Meshes
Lingjie Zhu, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ECCV (11) | 3 |
| 2018 | Accurate and efficient ground-to-aerial model alignment
Xiang Gao 0009, Lihua Hu, Hainan Cui, Shuhan Shen, Zhanyi Hu |
Pattern Recognit. | 1 |
| 2017 | Batched Incremental Structure-from-MotionabstractThe incremental Structure-from-Motion (SfM) technique has advanced in both robustness and accuracy, but the efficiency and scalability remain its key challenges. In this paper, we propose a novel batched incremental SfM technique to tackle these problems in a unified framework, where two iteration loops are contained. The inner loop is a tracks triangulation loop, where a novel tracks selection method is proposed to find a compact subset of tracks for the bundle adjustment (BA). The outer loop is a camera registration loop, where a batch of cameras are simultaneously added to alleviate the drifting risk and reduce the running times of BA. By the tracks selection and batched camera registration, we find these two iteration loops converge fast. Extensive experiments demonstrate that our new SfM system performs similarly or better than many of the state-of-the-art SfM systems in terms of camera calibration accuracy, while is more efficient, robust and scalable for large-scale scene reconstruction. Hainan Cui, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
3DV | 3 |
| 2017 | HSfM: Hybrid Structure-from-MotionabstractStructure-from-Motion (SfM) methods can be broadly categorized as incremental or global according to their ways to estimate initial camera poses. While incremental system has advanced in robustness and accuracy, the efficiency remains its key challenge. To solve this problem, global reconstruction system simultaneously estimates all camera poses from the epipolar geometry graph, but it is usually sensitive to outliers. In this work, we propose a new hybrid SfM method to tackle the issues of efficiency, accuracy and robustness in a unified framework. More specifically, we propose an adaptive community-based rotation averaging method first to estimate camera rotations in a global manner. Then, based on these estimated camera rotations, camera centers are computed in an incremental way. Extensive experiments show that our hybrid method performs similarly or better than many of the state-of-the-art global SfM approaches, in terms of computational efficiency, while achieves similar reconstruction accuracy and robustness with two other state-of-the-art incremental SfM approaches. Hainan Cui, Xiang Gao 0009, Shuhan Shen, Zhanyi Hu |
CVPR | 2 |
| 2017 | CSFM: Community-based structure from motionabstractStructure-from-Motion approaches could be broadly divided into two classes: incremental and global. While incremental manner is robust to outliers, it suffers from error accumulation and heavy computation load. The global manner has the advantage of simultaneously estimating all camera poses, but it is usually sensitive to epipolar geometry outliers. In this paper, we propose an adaptive community-based SfM (CSfM) method which takes both robustness and efficiency into consideration. First, the epipolar geometry graph is partitioned into separate communities. Then, the reconstruction problem is solved for each community in parallel. Finally, the reconstruction results are merged by a novel global similarity averaging method, which solves three convex L1 optimization problems. Experimental results show that our method performs better than many of the state-of-the-art global SfM approaches in terms of computational efficiency, while achieves similar or better reconstruction accuracy and robustness than many of the state-of-the-art incremental SfM approaches. Hainan Cui, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ICIP | 3 |
| 2017 | Accurate mesh-based alignment for ground and aerial multi-view stereo modelsabstractWe propose a method for accurate alignment of ground and aerial multi-view stereo (MVS) models. We achieve this goal by reconstructing the surface meshes from MVS point clouds generated by aerial and ground images respectively, and then iteratively removing the gap between them. The key issue is how to establish reliable correspondences between two meshes. To address this issue, we introduce a new set called the skeleton facet set (SFS) to represent the locally smooth part on the mesh, and then compute the transformation matrix by comparing the depths of the facets in SFS between aerial and ground models. Experimental results show that the proposed method is able to yield accurate alignment results and is robust to noise as well. Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ICIP | 3 |