VLDB 2026 Research / reviewers in the wild / expert
Shuhan Shen
dblp:84/5969
· DBLP profile ↗
81ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0002-8704-7914ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 8 first-author · 25 since 2021Artificial intelligence and machine learning · 45 · 6 first-author · 25 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MESA: Effective Matching Redundancy Reduction by Semantic Area SegmentationabstractMatching redundancy, which refers to fine-grained feature comparison between irrelevant image areas, is a prevalent limitation in current feature matching approaches. It leads to unnecessary and error-prone computations, ultimately diminishing matching accuracy. To reduce matching redundancy, we propose MESA and DMESA, both leveraging advanced image understanding of Segment Anything Model (SAM) to establish semantic area matches prior to point matching. These informative area matches, then, can undergo effective internal feature comparison, facilitating precise inside-area point matching. Specifically, MESA adopts a sparse matching framework, while DMESA applies a dense one. Both of them first obtain candidate areas from SAM results through a novel Area Graph (AG). In MESA, matching the candidates is formulated as a graph energy minimization and solved by graphical models derived from AG. In contrast, DMESA performs area matching by generating dense matching distributions on the entire image, aiming at enhancing efficiency. The distributions are produced from off-the-shelf patch matching, modeled as the Gaussian Mixture Model, and refined via the Expectation Maximization. With less repetitive computation, DMESA showcases an area matching speed improvement of nearly five times compared to MESA, while maintaining competitive accuracy. Our methods are extensively evaluated on four different tasks across six datasets, encompassing both indoor and outdoor scenes. The results suggest that our method achieves notable accuracy improvements for nine baselines of point matching in most cases. Furthermore, our methods exhibit promise generalization and improved robustness against image resolution. Yesheng Zhang, Shuhan Shen, Xu Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Incremental rotation averaging revisited
Xiang Gao 0009, Hainan Cui, Yangdong Liu, Shuhan Shen |
Pattern Recognit. | 4 |
| 2026 | MC-MVSNet: When multi-view stereo meets monocular cues
Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
Pattern Recognit. | 5 |
| 2025 | BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with TransformerabstractIn this paper, we present BWFormer, a novel Transformerbased model for building wireframe reconstruction from airborne LiDAR point cloud. The problem is solved in a ground-up manner here by detecting the building corners in 2D, lifting and connecting them in 3D space afterwards with additional data augmentation. Due to the 2.5D characteristic of the airborne LiDAR point cloud, we simplify the problem by projecting the points on the ground plane to produce a 2D height map. With the height map, a heat map is first generated with pixel-wise corner likelihood to predict the possible 2D corners. Then, 3D corners are predicted by a Transformer-based network with extra height embedding initialization. This 2D-to-3D corner detection strategy reduces the search space significantly. To recover the topological connections among the corners, edges are finally predicted from the height map with the proposed edge attention mechanism, which extracts holistic features and preserves local details simultaneously. In addition, due to the limited datasets in the field and the irregularity of the point clouds, a conditional latent diffusion model for LiDAR scanning simulation is utilized for data augmentation. BW-Former surpasses other state-of-the-art methods, especially in reconstruction completeness. Our code is available at: https : //github.com/3dv-casia/BWformer/. Lingjie Zhu, Hanqiao Ye, Shangfeng Huang, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen |
CVPR | 7 |
| 2025 | CoMatcher: Multi-View Collaborative Feature MatchingabstractThis paper proposes a multi-view collaborative matching strategy for reliable track construction in complex scenarios. We observe that the pairwise matching paradigms applied to image set matching often result in ambiguous estimation when the selected independent pairs exhibit significant occlusions or extreme viewpoint changes. This challenge primarily stems from the inherent uncertainty in interpreting intricate 3D structures based on limited two-view observations, as the 3D-to-2D projection leads to significant information loss. To address this, we introduce CoMatcher, a deep multi-view matcher to (i) leverage complementary context cues from different views to form a holistic 3D scene understanding and (ii) utilize cross-view projection consistency to infer a reliable global solution. Building on CoMatcher, we develop a groupwise framework that fully exploits cross-view relationships for large-scale matching tasks. Extensive experiments on various complex scenarios demonstrate the superiority of our method over the mainstream two-view matching paradigm. Zimin Xia, Mingyue Dong, Shuhan Shen, Linwei Yue, Xianwei Zheng |
CVPR | 4 |
| 2025 | MGSfM: Multi-Camera Geometry Driven Global Structure-from-MotionabstractMulti-camera systems are increasingly vital in the environmental perception of autonomous vehicles and robotics. Their physical configuration offers inherent fixed relative pose constraints that benefit Structure-from-Motion (SfM). However, traditional global SfM systems struggle with robustness due to their optimization framework. We propose a novel global motion averaging framework for multi-camera systems, featuring two core components: a decoupled rotation averaging module and a hybrid translation averaging module. Our rotation averaging employs a hierarchical strategy by first estimating relative rotations within rigid camera units and then computing global rigid unit rotations. To enhance the robustness of translation averaging, we incorporate both camera-to-camera and camera-to-point constraints to initialize camera positions and 3D points with a convex distance-based objective function and refine them with an unbiased non-bilinear angle-based objective function. Experiments on large-scale datasets show that our system matches or exceeds incremental SfM accuracy while significantly improving efficiency. Our framework outperforms existing global SfM methods, establishing itself as a robust solution for real-world multi-camera SfM applications. The code is available at https://github.com/3dv-casia/MGSfM/. Peilin Tao, Hainan Cui, Diantao Tu, Shuhan Shen |
ICCV | 4 |
| 2025 | Quadratic Gaussian Splatting: High Quality Surface Reconstruction with Second-Order Geometric Primitives
Hanqing Jiang, Liyang Zhou, Xiaojun Xiang, Shuhan Shen |
ICCV | 6 |
| 2025 | NeuralPlane: Structured 3D Reconstruction in Planar Primitives with Neural Fieldsabstract3D maps assembled from planar primitives are compact and expressive in representing man-made environments. In this paper, we present **NeuralPlane**, a novel approach that explores **neural** fields for multi-view 3D **plane** reconstruction. Our method is centered upon the core idea of distilling geometric and semantic cues from inconsistent 2D plane observations into a unified 3D neural representation, which unlocks the full leverage of plane attributes. It is accomplished through several key designs, including: 1) a monocular module that generates geometrically smooth and semantically meaningful segments known as 2D plane observations, 2) a plane-guided training procedure that implicitly learns accurate 3D geometry from the multi-view plane observations, and 3) a self-supervised feature field termed *Neural Coplanarity Field* that enables the modeling of scene semantics alongside the geometry. Without relying on prior plane annotations, our method achieves high-fidelity reconstruction comprising planar primitives that are not only crisp but also well-aligned with the semantic content. Comprehensive experiments on ScanNetv2 and ScanNet++ demonstrate the superiority of our method in both geometry and semantics. Hanqiao Ye, Yangdong Liu, Shuhan Shen |
ICLR | 4 |
| 2025 | Uncertainty Aware Multiple View Stereo Network with Accurate SupervisionabstractLearning-based multiple view stereo has gained significant attention recently. However, most methods rely on direct network supervision using provided ground-truth depth, which poses three inherent problems: resolution-dependent ground-truth artifacts, excessively challenging training examples (with relatively featureless textures), and use of less-viewed reference pixels for supervision, all of which hinder network optimization. To alleviate these problems, we propose an accurate network supervision paradigm that includes a ground-truth mask, an entropy mask, and a consistency mask, which provide more accurate supervision signals to aid network optimization. Furthermore, we introduce UANet, an uncertainty aware multi-view stereo network, which adaptively determines a pixel-wise search range using a dynamic range sampler (DRS) built upon estimation confidence and learned uncertainty. Experimental results on recent MVS datasets demonstrate the effectiveness of our method. Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
Comput. Vis. Media | 5 |
| 2025 | BPN: Building Pointer Network for Satellite Imagery Building Contour ExtractionabstractExtracting structured building contours from satellite imagery plays an important role in many geospatial tasks. However, it still remains a challenge due to the high cost of manual labeling, and models trained on simple polygons show poor generalization on buildings with more complex shapes. To deal with this, we propose a novel neural network called building pointer network (BPN) in this letter, which builds upon a recurrent neural network (RNN) architecture that integrates visual and geometric signals with an input-focused attention mechanism, making it more general for various shape complexity. Given an RGB satellite image, the model first uses a convolutional neural network (CNN) to obtain the set of key points for each building. Then, the coordinates of the key points and their image features are fused and fed into the RNN which ultimately predicts the index of the building corners sequentially. Results show that our method has good generalization ability for building data with complex shapes, provided that a dataset with relatively simple shapes is used as the training set. Lingjie Zhu, Zexiao Xie, Xiang Gao 0009, Shuhan Shen |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Fast and Interpretable 2D Homography Decomposition: Similarity-Kernel-Similarity and Affine-Core-Affine TransformationsabstractIn this article, we present two fast and interpretable decomposition methods for 2D homography, which are named Similarity-Kernel-Similarity (SKS) and Affine-Core-Affine (ACA) transformations respectively. Under the minimal 4-point configuration, two similarity transformations in SKS are computed by two anchor points on source and target planes, respectively. Then, the other two point correspondences can be exploited to compute the middle kernel transformation with only four parameters. Furthermore, ACA uses three anchor points to compute the source and the target affine transformations, followed by computation of the middle core transformation utilizing the other one point correspondence. ACA can compute a homography up to a scale with only 85 floating-point operations (FLOPs), without even any division operations. Therefore, as a plug-in module, ACA facilitates various traditional feature-based Random Sample Consensus (RANSAC) pipelines, as well as deep homography pipelines estimating 4-point offsets. In addition to the advantages of geometric parameterization and computational efficiency, SKS and ACA can express each element of homography by a polynomial of input coordinates (7th degree to 9th degree), extend the existing essential Similarity-Affine-Projective (SAP) decomposition and calculate 2D affine transformations in a unified way. Shen Cai, Zhanhao Wu, Lingxi Guo, Jiachun Wang, Junchi Yan, Shuhan Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | IncreLM: Incremental 3D Line Mapping
Xulong Bai, Hainan Cui, Shuhan Shen |
BMVC | 3 |
| 2024 | Unsigned Orthogonal Distance Fields: An Accurate Neural Implicit Representation for Diverse 3D ShapesabstractNeural implicit representation of geometric shapes has witnessed considerable advancements in recent years. However, common distance field based implicit represen-tations, specifically signed distance field (SDF) for water-tight shapes or unsigned distance field (UDF) for arbitrary shapes, routinely suffer from degradation of reconstruction accuracy when converting to explicit surface points and meshes. In this paper, we introduce a novel neural implicit representation based on unsigned orthogonal distance fields (UODFs). In UODFs, the minimal unsigned distance from any spatial point to the shape surface is de-fined solely in one orthogonal direction, contrasting with the multi-directional determination made by SDF and UDF. Consequently, every point in the 3D UODFs can directly access its closest surface points along three orthogonal di-rections. This distinctive feature leverages the accurate re-construction of surface points without interpolation errors. We verify the effectiveness of UODFs through a range of re-construction examples, extending from simple watertight or non-watertight shapes to complex shapes that include hol-lows, internal or assembling structures. Long Wan, Nayu Ding, Shuhan Shen, Shen Cai, Lin Gao 0004 |
CVPR | 5 |
| 2024 | Revisiting Global Translation Estimation with Feature TracksabstractGlobal translation estimation is a highly challenging step in the global structure from motion (SfM) algorithm. Many existing methods rely solely on relative translations, leading to inaccuracies in low parallax scenes and degradation under collinear camera motion. While recent approaches aim to address these issues by incorporating feature tracks into objective functions, they are often sensitive to outliers. In this paper, we first revisit global translation estimation methods with feature tracks and categorize them into explicit and implicit methods. Then, we highlight the superiority of the objective function based on the cross-product distance metric and propose a novel explicit global translation estimation framework that integrates both relative translations and feature tracks as input. To enhance the accuracy of input observations, we re-estimate relative translations with the coplanarity constraint of the epipolar plane and propose a simple yet effective strategy to select reliable feature tracks. Finally, we demonstrate the effectiveness of our approach through experiments on urban image sequences and unordered Internet images, showcasing its superior accuracy and robustness compared to many state-of-the-art techniques. Peilin Tao, Hainan Cui, Mengqi Rong, Shuhan Shen |
CVPR | 4 |
| 2024 | PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesabstractScaled relative pose estimation, i.e., estimating relative rotation and scaled relative translation between two images, has always been a major challenge in global Structure-from-Motion (SfM). This difficulty arises because the two-view relative translation computed by traditional geometric vision methods, e.g. the five-point algorithm, is scaleless. Many researchers have proposed diverse translation averaging methods to solve this problem. Instead of solving the problem in the motion averaging phase, we focus on estimating scaled relative pose with the help of panoramic cameras and deep neural networks. In this paper, a novel network, namely PanoPose, is proposed to estimate the relative motion in a fully self-supervised manner and a global SfM pipeline is built for panorama images. The proposed PanoPose comprises a depth-net and a pose-net, with self-supervision achieved by reconstructing the reference image from its neighboring images based on the estimated depth and relative pose. To maintain precise pose estimation under large viewing angle differences, we randomly rotate the panoramic images and pre-train the posenet with images before and after the rotation. To enhance scale accuracy, a fusion block is introduced to incorporate depth information into pose estimation. Extensive experiments on panoramic SfM datasets demonstrate the effectiveness of PanoPose compared with state-of-the-arts. Diantao Tu, Hainan Cui, Xianwei Zheng, Shuhan Shen |
CVPR | 4 |
| 2024 | Consistent 3D Line Mapping
Xulong Bai, Hainan Cui, Shuhan Shen |
ECCV (60) | 3 |
| 2024 | Easing 3D Pattern Reasoning with Side-View Features for Semantic Scene Completion
Linxi Huan, Mingyue Dong, Linwei Yue, Shuhan Shen, Xianwei Zheng |
ECCV (69) | 4 |
| 2024 | PolyRoom: Room-Aware Transformer for Floorplan Reconstruction
Lingjie Zhu, Hanqiao Ye, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen |
ECCV (50) | 7 |
| 2024 | BEV2PR: BEV-Enhanced Visual Place Recognition with Structural CuesabstractIn this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) For the methods relying on LiDAR sensors, the integration of LiDAR in robotic systems has led to increased expenses, while the alignment of data between different sensors is also a major challenge. 2) Other image-/camera-based methods, involving integrating RGB images and their derived variants (e.g., pseudo depth images, pseudo 3D point clouds), exhibit several limitations, such as the failure to effectively exploit the explicit spatial relationships between different objects. To tackle the above issues, we design a new BEV-enhanced VPR framework, namely BEV2PR, generating a composite descriptor with both visual cues and spatial awareness based on a single camera. The key points lie in: 1) We use BEV features as an explicit source of structural knowledge in constructing global features. 2) The lower layers of the pretrained backbone from BEV generation are shared for visual and structural streams in VPR, facilitating the learning of fine-grained local features in the visual stream. 3) The complementary visual and structural features can jointly enhance VPR performance. Our BEV2PR framework enables consistent performance improvements over several popular aggregation modules for RGB global features. The experiments on our collected VPR-NuScenes dataset demonstrate an absolute gain of 2.47% on Recall@1 for the strong Conv-AP baseline to achieve the best performance in our setting, and notably, a 18.06% gain on the hard set. The code and dataset will be available at https://github.com/FudongGe/BEV2PR. Fudong Ge, Shuhan Shen, Weiming Hu 0004 |
IROS | 3 |
| 2024 | IRAv3+: Hierarchical Incremental Rotation Averaging via Multiple Connected Dominating SetsabstractFocusing on the difficulty of absolute rotation globalization of large-scale rotation averaging problem, a novel hierarchical pipeline, termed as IRAv3+, based on multiple Connected Dominating Sets (CDSs) is proposed in this paper. Specifically, the proposed method not only obtains the graph clusters for local rotation averaging like other cluster-based methods, but also generate a subset via connected dominating set extraction, which is served as a reference for rotation globalization. To facilitate the rotation globalization, two key techniques are proposed: 1) to provide a more reliable global reference, instead of a single CDS, multiple CDSs are randomly selected and united; 2) to give a more accurate local-to-global alignment estimation, instead of using the relative rotation measurements of the sharing edges between local clusters and global reference, the absolute rotations of common vertices between them are involved. Experiments on the 1DSfM dataset demonstrate the effectiveness of the proposed IRAv3+ and its advantages over the existing cluster-based rotation averaging methods and other state of the arts. Xiang Gao 0009, Hainan Cui, Wantao Huang, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Neural Reflectance Decomposition Under Dynamic Point LightabstractDecomposing a scene into its 3D geometry, surface material textures, and illumination is a challenging but important problem in computer vision and graphics. While recent neural implicit representation based works have shown tremendous advantages, existing methods are not applicable to images illuminated by a single dynamic point light. We propose an entirely self-supervised end-to-end neural implicit representation based reflectance decomposition algorithm for objects under a dynamic point light. Our method adopts a staged training framework to estimate the geometry, light source position, and surface material textures through volume rendering, self-shadow inverse rendering, and physical model based surface rendering respectively. This scheme allows accurate recovery of the surface material textures which are coupled to the dynamic light, improving the reflectance decomposition capability. For evaluation, we collect a new dataset of several synthetic and real world objects illuminated by a moving point light. Experiments show that our method achieves superior reflectance decomposition performance compared to state-of-the-art methods, and the recovered elements can be deployed in existing graphics pipelines to perform relighting, material editing, and scene composition. Yuandong Li, Qinglei Hu, Zhenchao Ouyang, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Recent Advances in Conventional and Deep Learning-Based Depth Completion: A SurveyabstractDepth completion aims to recover pixelwise depth from incomplete and noisy depth measurements with or without the guidance of a reference RGB image. This task attracted considerable research interest due to its importance in various computer vision-based applications, such as scene understanding, autonomous driving, 3-D reconstruction, object detection, pose estimation, trajectory prediction, and so on. As the system input, an incomplete depth map is usually generated by projecting the 3-D points collected by ranging sensors, such as LiDAR in outdoor environments, or obtained directly from RGB-D cameras in indoor areas. However, even if a high-end LiDAR is employed, the obtained depth maps are still very sparse and noisy, especially in the regions near the object boundaries, which makes the depth completion task a challenging problem. To address this issue, a few years ago, conventional image processing-based techniques were employed to fill the holes and remove the noise from the relatively dense depth maps obtained by RGB-D cameras, while deep learning-based methods have recently become increasingly popular and inspiring results have been achieved, especially for the challenging situation of LiDAR-image-based depth completion. This article systematically reviews and summarizes the works related to the topic of depth completion in terms of input modalities, data fusion strategies, loss functions, and experimental settings, especially for the key techniques proposed in deep learning-based multiple input methods. On this basis, we conclude by presenting the current status of depth completion and discussing several prospects for its future research directions. Zexiao Xie, Xiaoxuan Yu, Xiang Gao 0009, Kunqian Li, Shuhan Shen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Shape Anchor Guided Holistic Indoor Scene UnderstandingabstractThis paper proposes a shape anchor guided learning strategy (AncLearn) for robust holistic indoor scene under-standing. We observe that the search space constructed by current methods for proposal feature grouping and instance point sampling often introduces massive noise to instance detection and mesh reconstruction. Accordingly, we develop AncLearn to generate anchors that dynamically fit instance surfaces to (i) unmix noise and target-related features for offering reliable proposals at the detection stage, and (ii) reduce outliers in object point sampling for directly providing well-structured geometry priors without segmentation during reconstruction. We embed AncLearn into a reconstruction-from-detection learning system (AncRec) to generate high-quality semantic scene models in a purely instance-oriented manner. Experiments conducted on the challenging ScanNetv2 dataset demonstrate that our shape anchor-based method consistently achieves state-of-the-art performance in terms of 3D object detection, layout estimation, and shape reconstruction. The code will be available at https://github.com/Geo-Tell/AncRec. Mingyue Dong, Linxi Huan, Hanjiang Xiong, Shuhan Shen, Xianwei Zheng |
ICCV | 4 |
| 2023 | IRAv3: Hierarchical Incremental Rotation Averaging on the FlyabstractWe present IRAv3, which is built upon the state-of-the-art rotation averaging method, IRA++, to push this fundamental task in 3D computer vision one step further. The key observation of this letter lies in that during IRA++, the community detection-based Epipolar-geometry Graph (EG) clustering is preemptive and permanent, which is not relevant to the follow-up rotation averaging task and limits the upper bound of absolute rotation estimation accuracy. In this letter, however, the EG clustering is performed along with the cluster-wise absolute rotation estimation, i.e. instead of pre-determination, the affiliation of each vertex to which EG cluster is determined “on the fly”, and the EG clustering finishes until all the vertices find the clusters they belong to, together with their absolute rotations estimated (in the local coordinate systems of the clusters they attached). By this way, a rotation averaging-targeted and -friendly EG clustering is obtained, which facilitates the rotation averaging task in turn. Experiments on both 1DSfM and KITTI odometry datasets demonstrate the effectiveness of our proposed IRAv3 on large-scale rotation averaging problems and its advantages over its previous works (IRA and IRA++) and other state of the arts. Xiang Gao 0009, Hainan Cui, Zexiao Xie, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | 3D Semantic Segmentation of Aerial Photogrammetry Models Based on Orthographic ProjectionabstractSemantic segmentation of 3D scenes is one of the most important tasks in the field of computer vision and has attracted much attention. In this paper, we propose a novel framework for 3D semantic segmentation of aerial photogrammetry models, which uses orthographic projection to improve efficiency while still ensuring high precision, and can also be applied to multiple types of models (i.e., textured mesh or colored point cloud). In our pipeline, we first obtain RGB images and elevation images from the 3D scene through orthographic projection, then use the image semantic segmentation network to segment these images to obtain pixel-wise semantic predictions, and finally back-project the segmentation results to the 3D model for fusion. Specifically, for the image semantic segmentation model, we design a cross-modality feature aggregation module and a context guidance module based on category features, which assist the network in learning more discriminative features between different objects. For the 2D-3D semantic fusion, we combine the segmentation results of the 2D images with the geometric consistency of the 3D models for joint optimization, which further improves the accuracy of the 3D semantic segmentation. Extensive experiments on two large-scale urban scenes demonstrate the efficiency and feasibility of our algorithm and surpass the current mainstream 3D deep learning methods. Mengqi Rong, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | MCSfM: Multi-Camera-Based Incremental Structure-From-MotionabstractFully perceiving the surrounding world is a vital capability for autonomous robots. To achieve this goal, a multi-camera system is usually equipped on the data collecting platform and the structure from motion (SfM) technology is used for scene reconstruction. However, although incremental SfM achieves high-precision modeling, it is inefficient and prone to scene drift in large-scale reconstruction tasks. In this paper, we propose a tailored incremental SfM framework for multi-camera systems, where the internal relative poses between cameras can not only be calibrated automatically but also serve as an additional constraint to improve the system robustness. Previous multi-camera based modeling work has mainly focused on stereo setups or multi-camera systems with known calibration information, but we allow arbitrary configurations and only require images as input. First, one camera is selected as the reference camera, and the other cameras in the multi-camera system are denoted as non-reference cameras. Based on the pose relationship between the reference and non-reference camera, the non-reference camera pose can be derived from the reference camera pose and internal relative poses. Then, a two-stage multi-camera based camera registration module is proposed, where the internal relative poses are computed first by local motion averaging, and then the rigid units are registered incrementally. Finally, a multi-camera based bundle adjustment is put forth to iteratively refine the reference camera and the internal relative poses. Experiments demonstrate that our system achieves higher accuracy and robustness on benchmark data compared to the state-of-the-art SfM and SLAM (simultaneous localization and mapping) methods. Hainan Cui, Xiang Gao 0009, Shuhan Shen |
IEEE Trans. Image Process. | 3 |
| 2023 | Efficient 3D Scene Semantic Segmentation via Active Learning on Rendered 2D ImagesabstractInspired by Active Learning and 2D-3D semantic fusion, we proposed a novel framework for 3D scene semantic segmentation based on rendered 2D images, which could efficiently achieve semantic segmentation of any large-scale 3D scene with only a few 2D image annotations. In our framework, we first render perspective images at certain positions in the 3D scene. Then we continuously fine-tune a pre-trained network for image semantic segmentation and project all dense predictions to the 3D model for fusion. In each iteration, we evaluate the 3D semantic model and re-render images in several representative areas where the 3D segmentation is not stable and send them to the network for training after annotation. Through this iterative process of rendering-segmentation-fusion, it can effectively generate difficult-to-segment image samples in the scene, while avoiding complex 3D annotations, so as to achieve label-efficient 3D scene segmentation. Experiments on three large-scale indoor and outdoor 3D datasets demonstrate the effectiveness of the proposed method compared with other state-of-the-art. Mengqi Rong, Hainan Cui, Shuhan Shen |
IEEE Trans. Image Process. | 3 |
| 2023 | Vehicle-Borne Multi-Sensor Temporal-Spatial Pose Globalization via Cross-Domain Data AssociationabstractLarge-scale urban scene 3D mapping has urgent demands and wide applications in many areas, where sensor pose globalization remains its fundamental problem and critical step. As the street-view images and vehicle-borne Light Detection And Ranging (LiDAR) points contain complementary advantages in urban scene 3D mapping, it is desirable to make the most of both to facilitate this task. Most existing methods make strong assumptions of strict synchronization, and even further, exact calibration between the vehicle-borne cameras and LiDARs, which are hard to guarantee in practice. To deal with this, we propose a novel pipeline for vehicle-borne camera and LiDAR temporal and spatial pose globalization with the guidance of Global Navigation Satellite System/Inertial Measurement Unit (GNSS/IMU), where both of the assumptions on strict synchronization and exact calibration are loosened. Specifically, the global poses of both cameras and LiDARs are first initialized by leveraging GNSS/IMU signals and multi-sensor pre-calibrations, and then refined by a global optimization scheme. To perform the global pose optimization, image-based, LiDAR-based, and cross-domain data association and constraint construction are conducted. Among them, the cross-domain ones, which are achieved by LiDAR point projection, image feature back-projection, and spatial point association, provide key clues for associating these two kinds of data with significant differences. Comprehensive experiments on both of a self-collected and the KITTI Odometry datasets demonstrate the effectiveness of our proposed method on multi-sensor pose globalization for large-scale urban scene 3D mapping. Xiang Gao 0009, Dongdong Tao, Yuqian Liu, Zexiao Xie, Shuhan Shen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | MMA: Multi-Camera Based Global Motion AveragingabstractIn order to fully perceive the surrounding environment, many intelligent robots and self-driving cars are equipped with a multi-camera system. Based on this system, the structure-from-motion (SfM) technology is used to realize scene reconstruction, but the fixed relative poses between cameras in the multi-camera system are usually not considered. This paper presents a tailor-made multi-camera based motion averaging system, where the fixed relative poses are utilized to improve the accuracy and robustness of SfM. Our approach starts by dividing the images into reference images and non-reference images, and edges in view-graph are divided into four categories accordingly. Then, a multi-camera based rotating averaging problem is formulated and solved in two stages, where an iterative re-weighted least squares scheme is used to deal with outliers. Finally, a multi-camera based translation averaging problem is formulated and a l1-norm based optimization scheme is proposed to compute the relative translations of multi-camera system and reference camera positions simultaneously. Experiments demonstrate that our algorithm achieves superior accuracy and robustness on various data sets compared to the state-of-the-art methods. Hainan Cui, Shuhan Shen |
AAAI | 2 |
| 2022 | Multi-Camera-LiDAR Auto-Calibration by Joint Structure-from-MotionabstractMultiple sensors, especially cameras and LiDARs, are widely used in autonomous vehicles. In order to fuse data from different sensors accurately, precise calibrations are required, including camera intrinsic parameters, and relative poses between multiple cameras and LiDARs. However, most existing camera-LiDAR calibration methods need to place manually designed calibration objects in multiple locations and multiple times, which are time-consuming and labor-intensive, and are not suitable for frequent use. To address that, in this paper we proposed a novel calibration pipeline that can automatically calibrate multiple cameras and multiple LiDARs in a Structure-from-Motion (SfM) process. In our pipeline, we first perform a global SfM on all images with the help of rough LiDAR data to get the initial poses of all sensors. Then, feature points on lines and planes are extracted from both SfM point cloud and LiDARs. With these features, a global Bundle Adjustment is performed to minimize the point reprojection errors, point-to-line errors, and point-to-plane errors together. During this minimization process, camera intrinsic parameters, camera and LiDAR poses, and SfM point cloud are refined jointly. The proposed method uses the characteristics of natural scenes, does not require manually designed calibration objects, and incorporates all calibration parameters into a unified optimization framework. Experiments on autonomous vehicles with different sensor configurations demonstrate the effectiveness and robustness of the proposed method. Diantao Tu, Baoyu Wang, Hainan Cui, Yuqian Liu, Shuhan Shen |
IROS | 5 |
| 2022 | IRA++: Distributed Incremental Rotation AveragingabstractBy observing that the recently presented Incremental Rotation Averaging (IRA) suffers from drifting and efficiency problems in large-scale situations, it is upgraded in this work to possess stronger scalability in both accuracy and efficiency based on the thought of divide and conquer. This upgraded version is termed as IRA++. Specifically, the original Epipolar-geometry Graph (EG) is clustered into several sub-graphs and inner-rotation averaging is distributedly performed in each of them with IRA at first. Then, the relative rotation between each pair of inner-sub-EG coordinate systems is distributedly estimated by a voting-based single rotation averaging method. Subsequently, IRA-based inter-rotation averaging is performed to obtain the absolute rotation of each inner-sub-EG coordinate system. And finally, the absolute rotations of all the cameras in the original EG are globally aligned and optimized to get the final rotation averaging result. Comprehensive evaluations on the 1DSfM, Campus, and San Francisco datasets demonstrate the advantages of our proposed IRA++ over IRA and several other state-of-the-art rotation averaging methods in both efficiency and accuracy, especially the accuracy in noise-polluted and efficiency in large-scale situations. Xiang Gao 0009, Lingjie Zhu, Hainan Cui, Zexiao Xie, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Incremental Translation AveragingabstractTranslation averaging is known to be more difficult than rotation averaging due to scale ambiguity, estimation sensitivity, and solution uncertainty. Existing approaches have exposed their limitations in terms of accuracy, robustness, simplicity, or efficiency. To tackle this tough problem, a simple yet effective translation averaging pipeline, termed as Incremental Translation Averaging (ITA), is proposed in this paper. It combines the advantages of high accuracy and robustness in incremental parameter estimation pipeline and the advantages of high simplicity and efficiency in global motion averaging approach. Unlike the traditional translation averaging methods which estimate all the absolute camera locations simultaneously and suffer from inaccuracy in parameter estimation and incompleteness in scene reconstruction, our ITA computes them novelly in an incremental way with higher accuracy and robustness. Thanks to the introduction of incremental parameter estimation thought into the translation averaging pipeline, 1) our ITA is robust to measurement outliers and accurate in parameter estimation; and 2) our ITA is simple and efficient because of its less dependency on complicated optimization, carefully-designed preprocessing, or additional information. Comprehensive evaluations on the 1DSfM dataset demonstrate the effectiveness of our ITA and its advantages over several state-of-the-art translation averaging approaches. Xiang Gao 0009, Lingjie Zhu, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Active Learning Based 3D Semantic Labeling From Images and Videosabstract3D semantic segmentation is one of the most fundamental problems for 3D scene understanding and has attracted much attention in the field of computer vision. In this paper, we propose an active learning based 3D semantic labeling method for large-scale 3D mesh model generated from images or videos. Taking as input a 3D mesh model reconstructed from the image based 3D modeling system, coupled with the calibrated images, our method outputs a fine 3D semantic mesh model in which each facet is assigned a semantic label. There are three major steps in our framework: 2D semantic segmentation, 2D-3D semantic fusion, and batch image selection. A limited annotation image set is first used to fine-tune a pre-trained semantic segmentation network for obtaining the pixel-wise semantic probability maps. Then all these maps are back-projected into 3D space and fused on the 3D mesh model using Markov Random Field optimization, thus yield a preliminary 3D semantic mesh model and a heat model showing each facet’s confidence. This 3D semantic model is used as a reliable supervisor to select the parts that are not well segmented for manual annotation to boost the performance of the 2D semantic segmentation network, as well as the 3D mesh labeling, in the next iteration. This Training-Fusion-Selection process continues until the label assignment of the 3D mesh model becomes steady. By this means, we significantly reduce the amount for annotation but not the labeling quality of 3D semantic models. Extensive experiments demonstrate the effectiveness and generalization ability of our method on a wide variety of datasets. Mengqi Rong, Hainan Cui, Zhanyi Hu, Hanqing Jiang, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | VidSfM: Robust and Accurate Structure-From-Motion for Monocular VideosabstractWith the popularization of smartphones, larger collection of videos with high quality is available, which makes the scale of scene reconstruction increase dramatically. However, high-resolution video produces more match outliers, and high frame rate video brings more redundant images. To solve these problems, a tailor-made framework is proposed to realize an accurate and robust structure-from-motion based on monocular videos. The key ideas include two points: one is to use the spatial and temporal continuity of video sequences to improve the accuracy and robustness of reconstruction; the other is to use the redundancy of video sequences to improve the efficiency and scalability of system. Our technical contributions include an adaptive way to identify accurate loop matching pairs, a cluster-based camera registration algorithm, a local rotation averaging scheme to verify the pose estimate and a local images extension strategy to reboot the incremental reconstruction. In addition, our system can integrate data from different video sequences, allowing multiple videos to be simultaneously reconstructed. Extensive experiments on both indoor and outdoor monocular videos demonstrate that our method outperforms the state-of-the-art approaches in robustness, accuracy and scalability. Hainan Cui, Diantao Tu, Fulin Tang, Pengfei Xu 0013, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Image Process. | 6 |
| 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road MappingabstractCompared to widely used LiDAR-based mapping in autonomous driving field, image-based mapping method has the advantages of low cost, high resolution, and no need for complex calibration. However, the image-based 3D mapping depends heavily on the texture richness and always leaves holes and outliers in low-textured areas, such as the road surface. To this end, this paper proposed a novel semantically guided Multi-View Stereo method for dense 3D road mapping, which integrates semantic information into PatchMatch-based MVS pipeline and uses image semantic segmentation as soft constraints in neighbor views selection, depth-map initialization, depth propagation, and depth-map completion. Experimental results on public and our own datasets show that, with the help of semantics, the proposed method achieves superior completeness with comparable accuracy for 3D road mapping compared to state-of-the-art MVS methods. Mingzhe Lv, Diantao Tu, Xincheng Tang, Yuqian Liu, Shuhan Shen |
ICRA | 5 |
| 2021 | Recalling Direct 2D-3D Matches for Large-Scale Visual LocalizationabstractEstimating the 6-DoF camera pose of an image with respect to a 3D scene model, known as visual localization, is a fundamental problem in many computer vision and robotics tasks. Among various visual localization methods, the direct 2D-3D matching method has become the preferred method for many practical applications due to its computational efficiency. When using direct 2D-3D matching methods in large-scale scenes, a vocabulary tree can be used to accelerate the matching process, which will also induce the quantization artifacts leading to reduce the inlier ratio and decrease the localization accuracy. To this end, in this paper two simple and effective mechanisms, called visibility-based recalling and space-based recalling, are proposed to recover lost matches caused by the quantization artifacts, thus can largely improve the localization accuracy and success rate without increasing too much computational time. Experimental results on long-term visual localization benchmarks demonstrate the effectiveness of our method compared with state-of-the-arts. Chuting Wang, Yuqian Liu, Shuhan Shen |
IROS | 4 |
| 2021 | Incremental Rotation Averaging
Xiang Gao 0009, Lingjie Zhu, Zexiao Xie, Hongmin Liu 0001, Shuhan Shen |
Int. J. Comput. Vis. | 5 |
| 2021 | View-graph construction framework for robust and efficient structure-from-motion
Hainan Cui, Tianxin Shi, Pengfei Xu 0013, Yiping Meng, Shuhan Shen |
Pattern Recognit. | 6 |
| 2021 | Urban Scene LOD Vectorized Modeling From Photogrammetry MeshesabstractUrban scene modeling is a challenging task for the photogrammetry and computer vision community due to its large scale, structural complexity, and topological delicacy. This paper presents an efficient multistep modeling framework for large-scale urban scenes from aerial images. It takes aerial images and a textured 3D mesh model generated by an image-based modeling system as the input and outputs compact polygon models with semantics at different levels of detail (LODs). Based on the key observation that urban buildings usually have piecewise planar rooftops and vertical walls, we propose a segment-based modeling method, which consists of three major stages: scene segmentation, roof contour extraction, and building modeling. By combining the deep neural network predictions with geometric constraints of the 3D mesh, the scene is first segmented into three classes. Then, for each building mesh, the 2D line segments are detected and used to slice the ground into polygon cells, followed by assigning each cell a roof plane via a MRF optimization. Finally, the LOD model is obtained by extruding cells to their corresponding planes. Compared with direct modeling in 3D space, we transform the mesh into a uniform 2D image grid representation and most of the modeling work is performed in 2D space, which has the advantages of low computational complexity and high robustness. In addition, our method doesn't require any global prior, such as the Manhattan or Atlanta world assumption, making it flexible to model scenes with different characteristics and complexity. Experiments on both single buildings and large-scale urban scenes demonstrate that by combining 2D photometric with 3D geometric information, the proposed algorithm is robust and efficient in urban scene LOD vectorized modeling compared with the state-of-the-art approaches. Jiali Han, Lingjie Zhu, Xiang Gao 0009, Zhanyi Hu, Liyang Zhou, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Image Process. | 7 |
| 2020 | 3D Semantic Labeling of Photogrammetry Meshes Based on Active LearningabstractAs different urban scenes are similar but still not completely consistent, coupled with the complexity of labeling directly in 3D, high-level understanding of 3D scenes has always been a tricky problem. In this paper, we propose a procedural approach for 3D semantic expression of urban scenes based on active learning. We first start with a small labeled image set to fine-tune a semantic segmentation network and then project its probability map onto a 3D mesh model for fusion, finally outputs a 3D semantic mesh model in which each facet has a semantic label and a heat model showing each facet's confidence. Our key observation is that our algorithm is iterative, in each iteration, we use the output semantic model as a supervision to select several valuable images for annotation to co-participate in the fine-tuning for overall improvement. In this way, we reduce the workload of labeling but not the quality of 3D semantic model. Using urban areas from two different cities, we show the potential of our method and demonstrate its effectiveness. Mengqi Rong, Shuhan Shen, Zhanyi Hu |
ICPR | 2 |
| 2020 | Graph-based parallel large scale structure from motion
Shuhan Shen, Yisong Chen |
Pattern Recognit. | 2 |
| 2020 | Depth-map completion for large indoor scene reconstruction
Hongmin Liu 0001, Xincheng Tang, Shuhan Shen |
Pattern Recognit. | 3 |
| 2020 | Complete Scene Reconstruction by Merging Images and Laser ScansabstractImage based modeling and laser scanning are two commonly used approaches in large-scale architectural scene reconstruction nowadays. In order to generate a complete scene reconstruction, an effective way is to completely cover the scene using ground and aerial images, supplemented by laser scanning on certain regions with low textures and complicated structures. Thus, the key issue is to accurately calibrate cameras and register laser scans in a unified framework. To this end, we proposed a three-step pipeline for complete scene reconstruction by merging images and laser scans. First, images are captured around the architecture in a multiview and multiscale way and are feed into a structure-from-motion (SfM) pipeline to generate SfM points. Then, based on the SfM result, the laser scanning locations are automatically planned by considering textural richness, structural complexity of the scene and spatial layout of the laser scans. Finally, the images and laser scans are accurately merged in a coarse-to-fine manner. Experimental evaluations on two ancient Chinese architecture datasets demonstrate the effectiveness of our proposed complete scene reconstruction pipeline. Xiang Gao 0009, Shuhan Shen, Lingjie Zhu, Tianxin Shi, Zhiheng Wang 0001, Zhanyi Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Distributed Surface Reconstruction from Point Cloud for City-Scale ScenesabstractImage based 3D modeling is an effective way to reconstruct large-scale scenes, especially city-level scenarios. In the image based modeling pipeline, obtaining a watertight mesh model from a noisy multiple view stereo point cloud is a key step to ensure the model quality. However, state-of-the-art method [34] relies on the global Delaunay-based optimization formed by all points and cameras, and will encounter scale problem when dealing with large scenes. To circumvent this limitation, this paper proposes a distributed surface reconstruction approach which could handle city-scale scenes with limited memory and time consumption. Firstly, the whole scene is adaptively divided into several chunks with overlapping boundaries, and each chunk can satisfy the memory limit. Then, the Delaunay-based optimization is performed to extract meshes for each chunk in parallel. Finally, the local meshes are merged together by resolving local inconsistencies in the overlapping areas. We test the proposed method on three city-scale scenes with billions of points and tens of thousands of images, and demonstrate its scalability and completeness compared with the state-of-the-art methods. Jiali Han, Shuhan Shen |
3DV | 2 |
| 2019 | Visual Localization Using Sparse Semantic 3D MapabstractAccurate and robust visual localization under a wide range of viewing condition variations including season and illumination changes, as well as weather and day-night variations, is the key component for many computer vision and robotics applications. Under these conditions, most traditional methods would fail to locate the camera. In this paper we present a visual localization algorithm that combines structure-based method and image-based method with semantic information. Given semantic information about the query and database images, the retrieved images are scored according to the semantic consistency of the 3D model and the query image. Then the semantic matching score is used as weight for RANSAC's sampling and the pose is solved by a standard PnP solver. Experiments on the challenging long-term visual localization benchmark dataset demonstrate that our method has significant improvement compared with the state-of-the-arts. Tianxin Shi, Shuhan Shen, Xiang Gao 0009, Lingjie Zhu |
ICIP | 2 |
| 2019 | Active Semantic Labeling of Street View Point CloudsabstractSemantic 3D models have shown their importance in many fields such as autonomous driving. However, it remains a tough task to assign semantic labels to various scenes. In this paper, we propose an Active Learning based method for semantic labeling of street view point clouds with a small amount of annotated data samples. The proposed method takes a point cloud and registrated images as the input, and yields a point cloud with semantic labels. We iteratively fine-tunes a network with the ever-enlarging training set to exploit the semantic information of the scene, and fuse the semantic labels in 3D space. To deal with the imbalanced data in street view scenes, a label biased criterion for query selection is proposed to help select images to efficiently improve the performance of the network and the quality of the semantic model. Experimental result shows that the proposed method demands limited human labor and works well in assigning semantic labels to the imbalanced scenes like street view scenes. Shuhan Shen, Zhanyi Hu |
ICME | 2 |
| 2019 | Ground and aerial meta-data integration for localization and reconstruction: A review
Xiang Gao 0009, Shuhan Shen, Zhanyi Hu, Zhiheng Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | Multi-source data-based 3D digital preservation of largescale ancient chinese architecture: A case reportabstractThe 3D digitalization and documentation of ancient Chinese architecture is challenging because of architectural complexity and structural delicacy. To generate complete and detailed models of this architecture, it is better to acquire, process, and fuse multi-source data instead of single-source data. In this paper, we describe our work on 3D digital preservation of ancient Chinese architecture based on multisource data. We first briefly introduce two surveyed ancient Chinese temples, Foguang Temple and Nanchan Temple. Then, we report the data acquisition equipment we used and the multi-source data we acquired. Finally, we provide an overview of several applications we conducted based on the acquired data, including ground and aerial image fusion, image and LiDAR (light detection and ranging) data fusion, and architectural scene surface reconstruction and semantic modeling. We believe that it is necessary to involve multi-source data for the 3D digital preservation of ancient Chinese architecture, and that the work in this paper will serve as a heuristic guideline for the related research communities. Xiang Gao 0009, Hainan Cui, Lingjie Zhu, Tianxin Shi, Shuhan Shen |
Virtual Real. Intell. Hardw. | 5 |
| 2018 | Progressive Large-Scale Structure-from-Motion with Orthogonal MSTsabstractPairwise image matching plays a vital role in Structure-from-Motion (SfM). Though the image-retrieval method accelerates the matching process, the number of neighbors is usually hard to determine. Insufficient feature matches could break the completeness of reconstructed scene, while redundant pairs may bring in many erroneous ones. In this paper, we propose a progressive SfM method to tackle the completeness, robustness and efficiency problems in a united framework, where two loops are contained. The outer loop is a feature matching loop, where the orthogonal MSTs (maximum spanning trees) of the image similarity graph is iteratively selected to perform the image matching. The inner loop is an incremental camera calibration loop, where the initial camera poses in each iteration are inherited from those calibrated in the last one. By progressively performing the image matching and calibration, we find both loops converge fast and a large number of redundant pairs are excluded. Experiments demonstrate the superior performance of our method in terms of both efficiency and robustness on various image datasets, and our method also has a large potential to tackle the ambiguity problems in SfM. Hainan Cui, Shuhan Shen, Wei Gao 0014, Zhiheng Wang 0001 |
3DV | 2 |
| 2018 | Fine-Level Semantic Labeling of Large-Scale 3D Model by Active LearningabstractSemantic labeling of 3D models has been a challenging task in recent years. Due to the various categories and shapes of 3D objects in different scenes, it is hard to develop a versatile method suitable for most scenes. In this paper, we propose an Active Learning based method to tackle the problem. The proposed method takes a 3D mesh model generated from images using SfM and MVS, as well as the calibrated images, as the input, and outputs a semantic mesh model in which each facet takes a fine-level semantic label. Starting with a small annotated image set, we progressively fine-tune a Convolutional Neural Network (CNN) with the ever-enlarging annotated image set for image semantic segmentation. In each iteration, by back-projecting the pixel labels to the 3D model and fusing them in 3D space, a semantic 3D model is generated. The semantic 3D model functions as a supervisor to select a batch of worthy images for annotation to boost the performance of the CNN in next iteration. This process iterates until the label assignment of the 3D model becomes steady. By making full use of the 3D geometric information, the proposed method could significantly reduce the annotation cost without losing the labeling quality of 3D models. Experimental results of fine-level labeling on two large-scale ancient Chinese architectures demonstrate the effectiveness of the proposed method. Shuhan Shen, Zhanyi Hu |
3DV | 2 |
| 2018 | Large Scale Urban Scene Modeling from MVS Meshes
Lingjie Zhu, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ECCV (11) | 2 |
| 2018 | Voting-based Incremental Structure-from-MotionabstractIncremental Structure-from-Motion (SfM) technique is the most prevalent way for image-based reconstruction, but its robustness is highly relying on each camera registration, where a false calibration could make everything following fail. In this paper, we propose a voting-based incremental SfM approach to improve upon the camera registration process. First, the degree of closeness between cameras is used as the vote to determine which cameras are going to register. Then, for each camera, two methods are simultaneously used to estimate the camera pose, and the number of inliers is used as the vote to determine which pose is more accurate. Finally, by estimating the priori global camera rotations from the view-graph, the camera poses that are consistent with the priori camera rotations are considered as getting double votes and preferentially kept. After all these prioritized cameras are calibrated, the other cameras are then incrementally registered. Compared to the state-of-the-art incremental SfM approaches, extensive experiments demonstrate that our system performs similarly or better in terms of reconstruction efficiency, while achieves a better robustness and accuracy. Especially for the ambiguous datasets, our system has a better potential to reconstruct them. Hainan Cui, Shuhan Shen, Wei Gao 0014 |
ICPR | 2 |
| 2018 | Accurate and efficient ground-to-aerial model alignment
Xiang Gao 0009, Lihua Hu, Hainan Cui, Shuhan Shen, Zhanyi Hu |
Pattern Recognit. | 4 |
| 2017 | Batched Incremental Structure-from-MotionabstractThe incremental Structure-from-Motion (SfM) technique has advanced in both robustness and accuracy, but the efficiency and scalability remain its key challenges. In this paper, we propose a novel batched incremental SfM technique to tackle these problems in a unified framework, where two iteration loops are contained. The inner loop is a tracks triangulation loop, where a novel tracks selection method is proposed to find a compact subset of tracks for the bundle adjustment (BA). The outer loop is a camera registration loop, where a batch of cameras are simultaneously added to alleviate the drifting risk and reduce the running times of BA. By the tracks selection and batched camera registration, we find these two iteration loops converge fast. Extensive experiments demonstrate that our new SfM system performs similarly or better than many of the state-of-the-art SfM systems in terms of camera calibration accuracy, while is more efficient, robust and scalable for large-scale scene reconstruction. Hainan Cui, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
3DV | 2 |
| 2017 | Variational Building Modeling from Urban MVS MeshesabstractIn this paper, we introduce a method for building LOD (levels of detail) modeling from urban multi-view stereo (MVS) meshes. Using city MVS meshes as input, our algorithm proceeds in three main steps: segmentation, contour extraction and modeling. With the prior knowledge and span constraint, we first segment the scene with an adapted variational measure to discover the underlying structures. The next contour extraction step projects the vertical structures onto the ground as line segments and extract the facade contours from them with a Markov random field. In the last modeling step, the contours are used to label the roof sections out and extruded to generate models of LODs with semantics. Experiments on complex and noisy urban meshes show that our approach could generate compact and accurate building models when compared with stateof- art methods. Lingjie Zhu, Shuhan Shen, Lihua Hu, Zhanyi Hu |
3DV | 2 |
| 2017 | HSfM: Hybrid Structure-from-MotionabstractStructure-from-Motion (SfM) methods can be broadly categorized as incremental or global according to their ways to estimate initial camera poses. While incremental system has advanced in robustness and accuracy, the efficiency remains its key challenge. To solve this problem, global reconstruction system simultaneously estimates all camera poses from the epipolar geometry graph, but it is usually sensitive to outliers. In this work, we propose a new hybrid SfM method to tackle the issues of efficiency, accuracy and robustness in a unified framework. More specifically, we propose an adaptive community-based rotation averaging method first to estimate camera rotations in a global manner. Then, based on these estimated camera rotations, camera centers are computed in an incremental way. Extensive experiments show that our hybrid method performs similarly or better than many of the state-of-the-art global SfM approaches, in terms of computational efficiency, while achieves similar reconstruction accuracy and robustness with two other state-of-the-art incremental SfM approaches. Hainan Cui, Xiang Gao 0009, Shuhan Shen, Zhanyi Hu |
CVPR | 3 |
| 2017 | CSFM: Community-based structure from motionabstractStructure-from-Motion approaches could be broadly divided into two classes: incremental and global. While incremental manner is robust to outliers, it suffers from error accumulation and heavy computation load. The global manner has the advantage of simultaneously estimating all camera poses, but it is usually sensitive to epipolar geometry outliers. In this paper, we propose an adaptive community-based SfM (CSfM) method which takes both robustness and efficiency into consideration. First, the epipolar geometry graph is partitioned into separate communities. Then, the reconstruction problem is solved for each community in parallel. Finally, the reconstruction results are merged by a novel global similarity averaging method, which solves three convex L1 optimization problems. Experimental results show that our method performs better than many of the state-of-the-art global SfM approaches in terms of computational efficiency, while achieves similar or better reconstruction accuracy and robustness than many of the state-of-the-art incremental SfM approaches. Hainan Cui, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ICIP | 2 |
| 2017 | Accurate mesh-based alignment for ground and aerial multi-view stereo modelsabstractWe propose a method for accurate alignment of ground and aerial multi-view stereo (MVS) models. We achieve this goal by reconstructing the surface meshes from MVS point clouds generated by aerial and ground images respectively, and then iteratively removing the gap between them. The key issue is how to establish reliable correspondences between two meshes. To address this issue, we introduce a new set called the skeleton facet set (SFS) to represent the locally smooth part on the mesh, and then compute the transformation matrix by comparing the depths of the facets in SFS between aerial and ground models. Experimental results show that the proposed method is able to yield accurate alignment results and is robust to noise as well. Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ICIP | 2 |
| 2017 | Global fusion of generalized camera model for efficient large-scale structure from motion
Hainan Cui, Shuhan Shen, Zhanyi Hu |
Sci. China Inf. Sci. | 2 |
| 2017 | Tracks selection for robust, efficient and scalable large-scale structure from motion
Hainan Cui, Shuhan Shen, Zhanyi Hu |
Pattern Recognit. | 2 |
| 2017 | Dynamic Graph Cuts in ParallelabstractThis paper aims at bridging the two important trends in efficient graph cuts in the literature, the one is to decompose a graph into several smaller subgraphs to take the advantage of parallel computation, the other is to reuse the solution of the max-flow problem on a residual graph to boost the efficiency on another similar graph. Our proposed parallel dynamic graph cuts algorithm takes the advantages of both, and is extremely efficient for certain dynamically changing MRF models in computer vision. The performance of our proposed algorithm is validated on two typical dynamic graph cuts problems: the foreground-background segmentation in video, where similar graph cuts problems need to be solved in sequential and GrabCut, where graph cuts are used iteratively. Miao Yu 0005, Shuhan Shen, Zhanyi Hu |
IEEE Trans. Image Process. | 2 |
| 2016 | Robust global translation averaging with feature tracksabstractHow to average translations is the single most difficult task in global structure-from-motion (SfM) to fully tap its potentials in terms of reconstruction efficiency and accuracy since usually only noisy translation directions can be factored out from essential matrices due to the inevitable matching outliers. To tackle this problem, this work proposes a two-step strategy. Firstly, a “2-point method” is introduced to refine the epipolar geometry by which a more accurate track set is generated. Then, translation lengths are computed by solving a convex L1 optimization according to the adjacent triangles induced by the selected tracks and translations. Extensive experiments show that our method performs similarly or better than the state-of-art SfM approaches in terms of the reconstruction accuracy, completeness and efficiency. Hainan Cui, Shuhan Shen, Zhanyi Hu |
ICPR | 2 |
| 2016 | Automatic building extraction from oblique aerial imagesabstractIn this paper we propose an automatic urban building extraction method for oblique aerial images. Five steps are included in this method: point cloud generation, grid partition, feature extraction, building detection and building reconstruction. Taking advantages of recent progress in large-scale Structure from Motion (SfM) and Multiple View Stereo (MVS), dense point cloud is generated first. Then, we project the point cloud into a regularly spaced grid in XY plan, and convert the building extraction problem into an image segmentation problem. By combining the strength of the geometric attribute and spectral attribute, three complementary features are extracted and a MRF based graph model along with an energy function is created. Points belonging to buildings are recognized by minimizing this function, and prismatic 3D building models are reconstructed accordingly. Shuhan Shen, Zhanyi Hu |
ICPR | 2 |
| 2016 | Dynamic Parallel and Distributed Graph CutsabstractGraph cuts are widely used in computer vision. To speed up the optimization process and improve the scalability for large graphs, Strandmark and Kahl introduced a splitting method to split a graph into multiple subgraphs for parallel computation in both shared and distributed memory models. However, this parallel algorithm (the parallel BK-algorithm) does not have a polynomial bound on the number of iterations and is found to be non-convergent in some cases due to the possible multiple optimal solutions of its sub-problems. To remedy this non-convergence problem, in this paper, we first introduce a merging method capable of merging any number of those adjacent sub-graphs that can hardly reach agreement on their overlapping regions in the parallel BK-algorithm. Based on the pseudo-boolean representations of graph cuts, our merging method is shown to be effectively reused all the computed flows in these sub-graphs. Through both splitting and merging, we further propose a dynamic parallel and distributed graph cuts algorithm with guaranteed convergence to the globally optimal solutions within a predefined number of iterations. In essence, this paper provides a general framework to allow more sophisticated splitting and merging strategies to be employed to further boost performance. Our dynamic parallel algorithm is validated with extensive experimental results. Miao Yu 0005, Shuhan Shen, Zhanyi Hu |
IEEE Trans. Image Process. | 2 |
| 2015 | Deep neural network based image annotation
Songhao Zhu, Chengjian Sun, Shuhan Shen |
Pattern Recognit. Lett. | 4 |
| 2015 | Efficient Large-Scale Structure From Motion by Fusing Auxiliary Imaging InformationabstractOne of the potentially effective means for large-scale 3D scene reconstruction is to reconstruct the scene in a global manner, rather than incrementally, by fully exploiting available auxiliary information on the imaging condition, such as camera location by Global Positioning System (GPS), orientation by inertial measurement unit (or compass), focal length from EXIF, and so on. However, such auxiliary information, though informative and valuable, is usually too noisy to be directly usable. In this paper, we present an approach by taking advantage of such noisy auxiliary information to improve structure from motion solving. More specifically, we introduce two effective iterative global optimization algorithms initiated with such noisy auxiliary information. One is a robust rotation averaging algorithm to deal with contaminated epipolar graph, the other is a robust scene reconstruction algorithm to deal with noisy GPS data for camera centers initialization. We found that by exclusively focusing on the estimated inliers at the current iteration, the optimization process initialized by such noisy auxiliary information could converge well and efficiently. Our proposed method is evaluated on real images captured by unmanned aerial vehicle, StreetView car, and conventional digital cameras. Extensive experimental results show that our method performs similarly or better than many of the state-of-art reconstruction approaches, in terms of reconstruction accuracy and completeness, but is more efficient and scalable for large-scale image data sets. Hainan Cui, Shuhan Shen, Wei Gao 0014, Zhanyi Hu |
IEEE Trans. Image Process. | 2 |
| 2014 | Fusion of Auxiliary Imaging Information for Robust, Scalable and Fast 3D Reconstruction
Hainan Cui, Shuhan Shen, Wei Gao 0014, Zhanyi Hu |
ACCV (1) | 2 |
| 2014 | How to Select Good Neighboring Images in Depth-Map Merging Based 3D ModelingabstractDepth-map merging based 3D modeling is an effective approach for reconstructing large-scale scenes from multiple images. In addition to generate high quality depth maps at each image, how to select suitable neighboring images for each image is also an important step in the reconstruction pipeline, unfortunately to which little attention has been paid in the literature until now. This paper is intended to tackle this issue for large scale scene reconstruction where many unordered images are captured and used with substantial varying scale and view-angle changes. We formulate the neighboring image selection as a combinatorial optimization problem and use the quantum-inspired evolutionary algorithm to seek its optimal solution. Experimental results on the ground truth data set show that our approach can significantly improve the quality of the depth-maps as well as final 3D reconstruction results with high computational efficiency. Shuhan Shen, Zhanyi Hu |
IEEE Trans. Image Process. | 1 |
| 2013 | Image annotation using high order statistics in non-Euclidean spaces
Songhao Zhu, Juanjuan Hu, Baoyun Wang, Shuhan Shen |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Accurate Multiple View 3D Reconstruction Using Patch-Based Stereo for Large-Scale ScenesabstractIn this paper, we propose a depth-map merging based multiple view stereo method for large-scale scenes which takes both accuracy and efficiency into account. In the proposed method, an efficient patch-based stereo matching process is used to generate depth-map at each image with acceptable errors, followed by a depth-map refinement process to enforce consistency over neighboring views. Compared to state-of-the-art methods, the proposed method can reconstruct quite accurate and dense point clouds with high computational efficiency. Besides, the proposed method could be easily parallelized at image level, i.e., each depth-map is computed individually, which makes it suitable for large-scale scene reconstruction with high resolution images. The accuracy and efficiency of the proposed method are evaluated quantitatively on benchmark data and qualitatively on large data sets. Shuhan Shen |
IEEE Trans. Image Process. | 1 |
| 2012 | Depth-map merging for Multi-View Stereo with high resolution images
Shuhan Shen |
ICPR | 1 |
| 2011 | A fast approach to deformable surface 3D tracking
Chenhao Wang 0003, Shuhan Shen, Yuncai Liu |
Pattern Recognit. | 2 |
| 2010 | Monocular 3D tracking of deformable surfaces using sequential second order cone programming
Shuhan Shen, Yuncai Liu, Wu-Sheng Lu |
Pattern Recognit. | 1 |
| 2010 | An Optimization Based Framework for Human Pose EstimationabstractIn computer vision community, human pose estimation and nonrigid shape recovery have evolved into different subfields. The state-of-the-art optimization techniques have been applied to the problem of deformable surface reconstruction successfully and recent methods in this area have focused on designing formulations that are easier to solve. In general, these techniques lay their success on the assumption that sufficient 2-D-3-D correspondences can be detected. By contrast, confronted with the similar ambiguity problem, many techniques for human pose estimation adopt stochastic searching or discriminative predictions, which allow for more generative image cues. However, the global optimization cannot be guaranteed via the stochastic methods; and discriminative techniques usually suffer from inaccuracy. In this letter, we absorb ideas from both domains and propose a unified approach for articulated human pose estimation. Specifically, we optimize the human pose to account for the discriminative pose prediction, bone length preservation in parallel with the point-topoint image observation. Moreover, the L2norm minimization is solved iteratively as a linear system with high computational efficiency. Junchi Yan, Shuhan Shen, Yin Li 0003, Yuncai Liu |
IEEE Signal Process. Lett. | 2 |
| 2010 | Convex Optimization for Nonrigid Stereo ReconstructionabstractWe present a method for recovering 3-D nonrigid structure from an image pair taken with a stereo rig. More specifically, we dedicate to recover shapes of nearly inextensible deformable surfaces. In our approach, we represent the surface as a 3-D triangulated mesh and formulate the reconstruction problem as an optimization problem consisting of data terms and shape terms. The data terms are model to image keypoint correspondences which can be formulated as second-order cone programming (SOCP) constraints using L(infinity) norm. The shape terms are designed to retaining original lengths of mesh edges which are typically nonconvex constraints. We will show that this optimization problem can be turned into a sequence of SOCP feasibility problems in which the nonconvex constraints are approximated as a set of convex constraints. Thanks to the efficient SOCP solver, the reconstruction problem can then be solved reliably and efficiently. As opposed to previous methods, ours neither involves smoothness constraints nor need an initial estimation, which enables us to recover shapes of surfaces with smooth, sharp and other complex deformations from a single image pair. The robustness and accuracy of our approach are evaluated quantitatively on synthetic data and qualitatively on real data. Shuhan Shen, Wenjuan Ma, Wenhuan Shi, Yuncai Liu |
IEEE Trans. Image Process. | 1 |
| 2010 | Monocular 3-D Tracking of Inextensible Deformable Surfaces Under L2 -NormabstractWe present a method for recovering the 3-D shape of an inextensible deformable surface from a monocular image sequence. State-of-the-art methods on this problem , utilize L(infinity)-norm of reprojection residual vectors and formulate the tracking problem as a Second-Order Cone Programming (SOCP) problem. Instead of using L(infinity) which is sensitive to outliers, we use L(2)-norm of reprojection errors. Generally, using L(2) leads a nonconvex optimization problem which is difficult to minimize. Instead of solving the nonconvex problem directly, we design an iterative L(2)-norm approximation process to approximate the nonconvex objective function, in which only a linear system needs to be solved at each iteration. Furthermore, we introduce a shape regularization term into this iterative process in order to keep the inextensibility of the recovered mesh. Compared with previous methods, ours performs more robust to image noises, outliers and large interframe motions with high computational efficiency. The robustness and accuracy of our approach are evaluated quantitatively on synthetic data and qualitatively on real data. Shuhan Shen, Wenhuan Shi, Yuncai Liu |
IEEE Trans. Image Process. | 1 |
| 2009 | Monocular Template-Based Tracking of Inextensible Deformable Surfaces under L2-Norm
Shuhan Shen, Wenhuan Shi, Yuncai Liu |
ACCV (2) | 1 |
| 2008 | Probability Evolutionary Algorithm based human motion tracking using voxel dataabstractA novel evolutionary algorithm called Probability Evolutionary Algorithm (PEA), and a method based on PEA for visual tracking of human body using voxel data are presented. PEA is inspired by the Quantum computation and the Quantum-inspired Evolutionary Algorithm, and it has a good balance between exploration and exploitation with very fast computation speed. The individual in PEA is encoded by the probabilistic compound bit, defined as the smallest unit of information, for the probabilistic representation. The observation step is used in PEA to obtain the observed states of the individual, and the update operator is used to evolve the individual. In the PEA based human tracking framework, tracking is considered to be a function optimization problem, so the aim is to optimize the matching function between the model and the image observation. Since the matching function is a very complex function in high-dimensional space, PEA is used to optimize it. Experiments on 3D human motion tracking using voxel data demonstrate the effectiveness, significance and computation efficiency of the proposed human tracking method. Shuhan Shen, Haolong Deng, Yuncai Liu |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | Deformable surface stereo tracking-by-detection using Second Order Cone ProgrammingabstractWe present a method for tracking deformable surfaces in 3D using a stereo rig. Different from traditional recursive tracking approaches that provide a strong prior on the pose for each new frame, the proposed method tracks deformable surfaces by detecting them in individual frames. In our method, the model of the surface is represented by a triangulated mesh. The constraints for model to image keypoint correspondences, together with the constraints that preserve the lengths of mesh edges, are formulated as Second Order Cone Programming (SOCP) constraints, leading this tracking-by-detection method to be an SOCP problem that can be effectively solved. Experiments on a piece of deformed paper demonstrate the capability of the proposed tracking-by-detection method. Shuhan Shen, Yinqiang Zheng, Yuncai Liu |
ICPR | 1 |
| 2008 | Efficient multiple faces tracking based on Relevance Vector Machine and Boosting learning
Shuhan Shen, Yuncai Liu |
J. Vis. Commun. Image Represent. | 1 |
| 2008 | Model based human motion tracking using probability evolutionary algorithm
Shuhan Shen, Minglei Tong, Haolong Deng, Yuncai Liu, Kaoru Wakabayashi, Hideki Koike |
Pattern Recognit. Lett. | 1 |