VLDB 2026 Research / reviewers in the wild / expert
Lingjie Zhu
dblp:220/7763
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with TransformerabstractIn this paper, we present BWFormer, a novel Transformerbased model for building wireframe reconstruction from airborne LiDAR point cloud. The problem is solved in a ground-up manner here by detecting the building corners in 2D, lifting and connecting them in 3D space afterwards with additional data augmentation. Due to the 2.5D characteristic of the airborne LiDAR point cloud, we simplify the problem by projecting the points on the ground plane to produce a 2D height map. With the height map, a heat map is first generated with pixel-wise corner likelihood to predict the possible 2D corners. Then, 3D corners are predicted by a Transformer-based network with extra height embedding initialization. This 2D-to-3D corner detection strategy reduces the search space significantly. To recover the topological connections among the corners, edges are finally predicted from the height map with the proposed edge attention mechanism, which extracts holistic features and preserves local details simultaneously. In addition, due to the limited datasets in the field and the irregularity of the point clouds, a conditional latent diffusion model for LiDAR scanning simulation is utilized for data augmentation. BW-Former surpasses other state-of-the-art methods, especially in reconstruction completeness. Our code is available at: https : //github.com/3dv-casia/BWformer/. Lingjie Zhu, Hanqiao Ye, Shangfeng Huang, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen |
CVPR | 2 |
| 2025 | BPN: Building Pointer Network for Satellite Imagery Building Contour ExtractionabstractExtracting structured building contours from satellite imagery plays an important role in many geospatial tasks. However, it still remains a challenge due to the high cost of manual labeling, and models trained on simple polygons show poor generalization on buildings with more complex shapes. To deal with this, we propose a novel neural network called building pointer network (BPN) in this letter, which builds upon a recurrent neural network (RNN) architecture that integrates visual and geometric signals with an input-focused attention mechanism, making it more general for various shape complexity. Given an RGB satellite image, the model first uses a convolutional neural network (CNN) to obtain the set of key points for each building. Then, the coordinates of the key points and their image features are fused and fed into the RNN which ultimately predicts the index of the building corners sequentially. Results show that our method has good generalization ability for building data with complex shapes, provided that a dataset with relatively simple shapes is used as the training set. Lingjie Zhu, Zexiao Xie, Xiang Gao 0009, Shuhan Shen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | PolyRoom: Room-Aware Transformer for Floorplan Reconstruction
Lingjie Zhu, Hanqiao Ye, Xiang Gao 0009, Xianwei Zheng, Shuhan Shen |
ECCV (50) | 2 |
| 2022 | GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD DrawingsabstractSpotting graphical symbols from the computer-aided design (CAD) drawings is essential to many industrial applications. Different from raster images, CAD drawings are vector graphics consisting of geometric primitives such as segments, arcs, and circles. By treating each CAD drawing as a graph, we propose a novel graph attention network GAT-CADNet to solve the panoptic symbol spotting problem: vertex features derived from the GAT branch are mapped to semantic labels, while their attention scores are cascaded and mapped to instance prediction. Our key contributions are three-fold: 1) the instance symbol spotting task is formulated as a subgraph detection problem and solved by predicting the adjacency matrix; 2) a relative spatial encoding (RSE) module explicitly encodes the relative positional and geometric relation among vertices to enhance the vertex attention; 3) a cascaded edge encoding (CEE) module extracts vertex attentions from multiple stages of GAT and treats them as edge encoding to predict the adjacency matrix. The proposed GAT-CADNet is intuitive yet effective and manages to solve the panoptic symbol spotting problem in one consolidated network. Extensive experiments and ablation studies on the public benchmark show that our graph-based approach surpasses existing state-of-the-art methods by a large margin. Zhaohua Zheng, Jianfang Li 0001, Lingjie Zhu, Honghua Li, Frank Petzold, Ping Tan 0002 |
CVPR | 3 |
| 2022 | IRA++: Distributed Incremental Rotation AveragingabstractBy observing that the recently presented Incremental Rotation Averaging (IRA) suffers from drifting and efficiency problems in large-scale situations, it is upgraded in this work to possess stronger scalability in both accuracy and efficiency based on the thought of divide and conquer. This upgraded version is termed as IRA++. Specifically, the original Epipolar-geometry Graph (EG) is clustered into several sub-graphs and inner-rotation averaging is distributedly performed in each of them with IRA at first. Then, the relative rotation between each pair of inner-sub-EG coordinate systems is distributedly estimated by a voting-based single rotation averaging method. Subsequently, IRA-based inter-rotation averaging is performed to obtain the absolute rotation of each inner-sub-EG coordinate system. And finally, the absolute rotations of all the cameras in the original EG are globally aligned and optimized to get the final rotation averaging result. Comprehensive evaluations on the 1DSfM, Campus, and San Francisco datasets demonstrate the advantages of our proposed IRA++ over IRA and several other state-of-the-art rotation averaging methods in both efficiency and accuracy, especially the accuracy in noise-polluted and efficiency in large-scale situations. Xiang Gao 0009, Lingjie Zhu, Hainan Cui, Zexiao Xie, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Incremental Translation AveragingabstractTranslation averaging is known to be more difficult than rotation averaging due to scale ambiguity, estimation sensitivity, and solution uncertainty. Existing approaches have exposed their limitations in terms of accuracy, robustness, simplicity, or efficiency. To tackle this tough problem, a simple yet effective translation averaging pipeline, termed as Incremental Translation Averaging (ITA), is proposed in this paper. It combines the advantages of high accuracy and robustness in incremental parameter estimation pipeline and the advantages of high simplicity and efficiency in global motion averaging approach. Unlike the traditional translation averaging methods which estimate all the absolute camera locations simultaneously and suffer from inaccuracy in parameter estimation and incompleteness in scene reconstruction, our ITA computes them novelly in an incremental way with higher accuracy and robustness. Thanks to the introduction of incremental parameter estimation thought into the translation averaging pipeline, 1) our ITA is robust to measurement outliers and accurate in parameter estimation; and 2) our ITA is simple and efficient because of its less dependency on complicated optimization, carefully-designed preprocessing, or additional information. Comprehensive evaluations on the 1DSfM dataset demonstrate the effectiveness of our ITA and its advantages over several state-of-the-art translation averaging approaches. Xiang Gao 0009, Lingjie Zhu, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol SpottingabstractAccess to large and diverse computer-aided design (CAD) drawings is critical for developing symbol spotting algorithms. In this paper, we present FloorPlan-CAD, a large-scale real-world CAD drawing dataset containing over 10,000 floor plans, ranging from residential to commercial buildings. CAD drawings in the dataset are all represented as vector graphics, which enable us to provide line-grained annotations of 30 object categories. Equipped by such annotations, we introduce the task of panoptic symbol spotting, which requires to spot not only instances of countable things, but also the semantic of uncountable stuff. Aiming to solve this task, we propose a novel method by combining Graph Convolutional Networks (GCNs) with Convolutional Neural Networks (CNNs), which captures both non-Euclidean and Euclidean features and can be trained end-to-end. The proposed CNN-GCN method achieved state-of-the-art (SOTA) performance on the task of semantic symbol spotting, and help us build a baseline network for the panoptic symbol spotting task. Our contributions are three-fold: 1) to the best of our knowledge, the presented CAD drawing dataset is the first of its kind; 2) the panoptic symbol spotting task considers the spotting of both thing instances and stuff semantic as one recognition problem; and 3) we presented a baseline solution to the panoptic symbol spotting task based on a novel CNN-GCN method, which achieved SOTA performance on semantic symbol spotting. We believe that these contributions will boost research in related areas. The dataset and code is publicly available at https://floorplancad.github.io/. Zhiwen Fan, Lingjie Zhu, Honghua Li, Xiaohao Chen, Siyu Zhu 0001, Ping Tan 0002 |
ICCV | 2 |
| 2021 | Incremental Rotation Averaging
Xiang Gao 0009, Lingjie Zhu, Zexiao Xie, Hongmin Liu 0001, Shuhan Shen |
Int. J. Comput. Vis. | 2 |
| 2021 | Urban Scene LOD Vectorized Modeling From Photogrammetry MeshesabstractUrban scene modeling is a challenging task for the photogrammetry and computer vision community due to its large scale, structural complexity, and topological delicacy. This paper presents an efficient multistep modeling framework for large-scale urban scenes from aerial images. It takes aerial images and a textured 3D mesh model generated by an image-based modeling system as the input and outputs compact polygon models with semantics at different levels of detail (LODs). Based on the key observation that urban buildings usually have piecewise planar rooftops and vertical walls, we propose a segment-based modeling method, which consists of three major stages: scene segmentation, roof contour extraction, and building modeling. By combining the deep neural network predictions with geometric constraints of the 3D mesh, the scene is first segmented into three classes. Then, for each building mesh, the 2D line segments are detected and used to slice the ground into polygon cells, followed by assigning each cell a roof plane via a MRF optimization. Finally, the LOD model is obtained by extruding cells to their corresponding planes. Compared with direct modeling in 3D space, we transform the mesh into a uniform 2D image grid representation and most of the modeling work is performed in 2D space, which has the advantages of low computational complexity and high robustness. In addition, our method doesn't require any global prior, such as the Manhattan or Atlanta world assumption, making it flexible to model scenes with different characteristics and complexity. Experiments on both single buildings and large-scale urban scenes demonstrate that by combining 2D photometric with 3D geometric information, the proposed algorithm is robust and efficient in urban scene LOD vectorized modeling compared with the state-of-the-art approaches. Jiali Han, Lingjie Zhu, Xiang Gao 0009, Zhanyi Hu, Liyang Zhou, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Image Process. | 2 |
| 2020 | Complete Scene Reconstruction by Merging Images and Laser ScansabstractImage based modeling and laser scanning are two commonly used approaches in large-scale architectural scene reconstruction nowadays. In order to generate a complete scene reconstruction, an effective way is to completely cover the scene using ground and aerial images, supplemented by laser scanning on certain regions with low textures and complicated structures. Thus, the key issue is to accurately calibrate cameras and register laser scans in a unified framework. To this end, we proposed a three-step pipeline for complete scene reconstruction by merging images and laser scans. First, images are captured around the architecture in a multiview and multiscale way and are feed into a structure-from-motion (SfM) pipeline to generate SfM points. Then, based on the SfM result, the laser scanning locations are automatically planned by considering textural richness, structural complexity of the scene and spatial layout of the laser scans. Finally, the images and laser scans are accurately merged in a coarse-to-fine manner. Experimental evaluations on two ancient Chinese architecture datasets demonstrate the effectiveness of our proposed complete scene reconstruction pipeline. Xiang Gao 0009, Shuhan Shen, Lingjie Zhu, Tianxin Shi, Zhiheng Wang 0001, Zhanyi Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Visual Localization Using Sparse Semantic 3D MapabstractAccurate and robust visual localization under a wide range of viewing condition variations including season and illumination changes, as well as weather and day-night variations, is the key component for many computer vision and robotics applications. Under these conditions, most traditional methods would fail to locate the camera. In this paper we present a visual localization algorithm that combines structure-based method and image-based method with semantic information. Given semantic information about the query and database images, the retrieved images are scored according to the semantic consistency of the 3D model and the query image. Then the semantic matching score is used as weight for RANSAC's sampling and the pose is solved by a standard PnP solver. Experiments on the challenging long-term visual localization benchmark dataset demonstrate that our method has significant improvement compared with the state-of-the-arts. Tianxin Shi, Shuhan Shen, Xiang Gao 0009, Lingjie Zhu |
ICIP | 4 |
| 2019 | Multi-source data-based 3D digital preservation of largescale ancient chinese architecture: A case reportabstractThe 3D digitalization and documentation of ancient Chinese architecture is challenging because of architectural complexity and structural delicacy. To generate complete and detailed models of this architecture, it is better to acquire, process, and fuse multi-source data instead of single-source data. In this paper, we describe our work on 3D digital preservation of ancient Chinese architecture based on multisource data. We first briefly introduce two surveyed ancient Chinese temples, Foguang Temple and Nanchan Temple. Then, we report the data acquisition equipment we used and the multi-source data we acquired. Finally, we provide an overview of several applications we conducted based on the acquired data, including ground and aerial image fusion, image and LiDAR (light detection and ranging) data fusion, and architectural scene surface reconstruction and semantic modeling. We believe that it is necessary to involve multi-source data for the 3D digital preservation of ancient Chinese architecture, and that the work in this paper will serve as a heuristic guideline for the related research communities. Xiang Gao 0009, Hainan Cui, Lingjie Zhu, Tianxin Shi, Shuhan Shen |
Virtual Real. Intell. Hardw. | 3 |
| 2018 | Large Scale Urban Scene Modeling from MVS Meshes
Lingjie Zhu, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ECCV (11) | 1 |
| 2017 | Variational Building Modeling from Urban MVS MeshesabstractIn this paper, we introduce a method for building LOD (levels of detail) modeling from urban multi-view stereo (MVS) meshes. Using city MVS meshes as input, our algorithm proceeds in three main steps: segmentation, contour extraction and modeling. With the prior knowledge and span constraint, we first segment the scene with an adapted variational measure to discover the underlying structures. The next contour extraction step projects the vertical structures onto the ground as line segments and extract the facade contours from them with a Markov random field. In the last modeling step, the contours are used to label the roof sections out and extruded to generate models of LODs with semantics. Experiments on complex and noisy urban meshes show that our approach could generate compact and accurate building models when compared with stateof- art methods. Lingjie Zhu, Shuhan Shen, Lihua Hu, Zhanyi Hu |
3DV | 1 |