EDBT 2026 Demo / reviewers in the wild / expert
Liangliang Nan
dblp:17/1760
· DBLP profile ↗
40ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-5629-9975ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuSEditor: From Multi-View Images to Text-Guided Neural Surface EditsabstractImplicit surface representations are valued for their compactness and continuity, but they pose significant challenges for editing. Despite recent advancements, existing methods often fail to preserve identity and maintain geometric consistency during editing. To address these challenges, we present NeuSEditor, a novel method for text-guided editing of neural implicit surfaces derived from multi-view images. NeuSEditor introduces an identity-preserving architecture that efficiently separates scenes into foreground and background, enabling precise modifications without altering the scene-specific elements. Our geometry-aware distillation loss significantly enhances rendering and geometric quality. Our method simplifies the editing workflow by eliminating the need for continuous dataset updates and source prompting. NeuSEditor outperforms recent state-of-the-art methods, delivering superior quantitative and qualitative results. For visual results, visit: neuseditor.github.io. Nail Ibrahimli, Julian F. P. Kooij, Liangliang Nan |
3DV | 3 |
| 2026 | MorphCut: an efficient convex decomposition method of 3D building models for urban morphological analyticsabstractUrban morphological analytics on buildings is informative for sustainable development. 3D building massing features, such as courtyards and setbacks, reflect spatial organizations and circulations, while influence daylight access, ventilation, and shading. However, existing 3D GIS methods usually overlook such 3D massing features, further obscure morphological analytics and environmental assessment. This article proposes MorphCut, an efficient convex decomposition method that segments 3D shapes into mass-aligned parts. MorphCut leverages key morphological properties—planarity, regularity, and Gestalt laws—after a topological preprocessing step to enable mass-aware decomposition. Experiments on representative samples, ranging from small houses to complex skyscrapers, showed that MorphCut outperformed four baseline methods in (i) balancing convexity and compactness, (ii) aligning decomposed parts with building masses, and (iii) preserving geometric fidelity (average deviation: 0.25 m). An urban-scale validation on datasets from Delft and Hong Kong, comprising over 30,000 buildings across 18.3 km², demonstrated MorphCut’s robustness, scalability, and generalizability. MorphCut successfully decomposed 98% of buildings in low-rise regions (+78% over the second-best method) and 93% in high-rise areas (+2%), completing processing in 13 hours (3 hours faster). These results position MorphCut as a foundational 3D GIS tool for large-scale, mass-aware morphological analysis, with implications for digital twins, sustainable planning, and environmental modeling. Fan Xue, Liangliang Nan, Longyong Wu, Jantien E. Stoter, Anthony Gar-On Yeh |
Int. J. Geogr. Inf. Sci. | 3 |
| 2026 | CrossTracker: Robust Multi-Modal 3D Multi-Object Tracking via Cross CorrectionabstractInaccurate detections remain a critical bottleneck in 3D multi-object tracking (MOT). Recent detection fusion-based methods incorporate camera detections as supplementary to reduce false detections and compensate for missing ones in LiDAR. However, their unidirectional camera-LiDAR correction lacks a feedback mechanism, precluding iterative mutual refinement between modalities for more robust LiDAR-based tracking. Inspired by the coarse-to-fine strategy in two-stage object detection, we introduceCrossTracker, a novel two-stage framework for online multi-modal 3D MOT. CrossTracker first constructs coarse camera and LiDAR trajectories independently, then performs trajectory fusion using both current and historical frames, without requiring future data. This ensures more robust mutual refinement between modalities. Specifically, CrossTracker comprises three core modules: i) the multi-modal modeling (M3) module, which fuses data from images, point clouds, and even planar geometry derived from images to establish a robust tracking constraint; ii) the coarse trajectory generation (C-TG) module, which independently generates coarse trajectories for both modalities using the M3constraint; and iii) the trajectory fusion (TF) module, which applies mutual refinement between coarse LiDAR and camera trajectories through cross correction to ensure robust LiDAR trajectories. Extensive experiments show that CrossTracker outperforms 19 state-of-the-art methods, highlighting its effectiveness in leveraging the synergistic strengths of camera and LiDAR sensors for robust multi-modal 3D MOT. The code is available at https://github.com/lipeng-gu/CrossTracker. Lipeng Gu, Xuefeng Yan 0001, Weiming Wang 0002, Honghua Chen, Dingkun Zhu, Liangliang Nan, Mingqiang Wei |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | CosCAD: Cross-Modal CAD Model Retrieval and Pose Alignment from a Single Image
Zhikun Wen, Honghua Chen, Zhe Zhu, Zeyong Wei, Liangliang Nan, Mingqiang Wei |
CVM (2) | 5 |
| 2025 | Parametric Point Cloud Completion for Polygonal Surface ReconstructionabstractExisting polygonal surface reconstruction methods heavily depend on input completeness and struggle with incomplete point clouds. We argue that while current point cloud completion techniques may recover missing points, they are not optimized for polygonal surface reconstruction, where the parametric representation of underlying surfaces remains overlooked. To address this gap, we introduce parametric completion, a novel paradigm for point cloud completion, which recovers parametric primitives instead of individual points to convey high-level geometric structures. Our presented approach, PaCo, enables high-quality polygonal surface reconstruction by leveraging plane proxies that encapsulate both plane parameters and inlier points, proving particularly effective in challenging scenarios with highly incomplete data. Comprehensive evaluations of our approach on the ABC dataset establish its effectiveness with superior performance and set a new standard for polygonal surface reconstruction from incomplete data. Project page: https://parametric-completion.github.io. Zhaiyu Chen, Liangliang Nan, Xiao Xiang Zhu 0001 |
CVPR | 3 |
| 2025 | SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban MeshesabstractSemantic segmentation in urban scene analysis has mainly focused on images or point clouds, while textured meshes—offering richer spatial representation—remain underexplored. This paper introduces SUM Parts, the first large-scale dataset for urban textured meshes with part-level semantic labels, covering about 2.5 km2with 21 classes. The dataset was created using our own annotation tool, which supports both face- and texture-based annotations with efficient interactive selection. We also provide a comprehensive evaluation of 3D semantic segmentation and interactive annotation methods on this dataset. Our project page is available at https://tudelft3d.github.io/SUMParts/. Weixiao Gao, Liangliang Nan, Hugo Ledoux |
CVPR | 2 |
| 2025 | VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene FlowabstractScene flow estimation aims to recover per-point motion from two adjacent LiDAR scans. However, in real-world applications such as autonomous driving, points rarely move independently of others, especially for nearby points belonging to the same object, which often share the same motion. Incorporating this locally rigid motion constraint has been a key challenge in self-supervised scene flow estimation, which is often addressed by post-processing or appending extra regularization. While these approaches are able to improve the rigidity of predicted flows, they lack an architectural inductive bias for local rigidity within the model structure, leading to suboptimal learning efficiency and inferior performance. In contrast, we enforce local rigidity with a lightweight add-on module in neural network design, enabling end-to-end learning. We design a discretized voting space that accommodates all possible translations and then identify the one shared by nearby points by differentiable voting. Additionally, to ensure computational efficiency, we operate on pillars rather than points and learn representative features for voting per pillar. We plug the Voting Module into popular model designs and evaluate its benefit on Argoverse 2 and Waymo datasets. We outperform baseline works with only marginal compute overhead. Code is available at https://github.com/tudelft-iv/VoteFlow. Yancong Lin, Liangliang Nan, Julian F. P. Kooij, Holger Caesar |
CVPR | 3 |
| 2025 | MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPabstractImage-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. The differential perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing the projective constraints on 2D-3D correspondences. However, differential PnP is highly sensitive to noise and outliers in the predicted correspondences. This issue hinders the effectiveness of correspondence learning. Inspired by the robustness of blind PnP against noise and outliers in correspondences, we propose an approximated blind PnP based correspondence learning approach. To mitigate the high computational cost of blind PnP, we simplify blind PnP to an amenable task of minimizing Chamfer distance between learned 2D and 3D keypoints, called MinCD-PnP. To effectively solve MinCD-PnP, we design a lightweight multi-task learning module, named as MinCD-Net, which can be easily integrated into the existing I2P registration architectures. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected datasets demonstrate that MinCD-Net outperforms state-of-the-art methods and achieves a higher inlier ratio (IR) and registration recall (RR) in both cross-scene and cross-dataset settings. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Liangliang Nan |
ICCV | 7 |
| 2025 | Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology RelationshipabstractImage-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Jie Ma 0003, Liangliang Nan |
IJCAI | 7 |
| 2025 | Enhance Image-to-Point-Cloud Registration with Beltrami Flow
Pei An, You Yang 0002, Jiaqi Yang 0002, Muyao Peng, Qiong Liu 0001, Liangliang Nan |
Int. J. Comput. Vis. | 6 |
| 2025 | AdLeaf: Quantitative Leaf Reconstruction From TLS Point CloudsabstractQuantitatively reconstructing the 3D structure of individual leaves within tree canopies is critical for understanding forest function and environmental responses to climate change. While quantitative structure models (QSMs) using terrestrial laser scanning (TLS) effectively capture woody structures, they lack the capability to accurately reconstruct non-woody leaf components. This study proposes AdLeaf (Accurate and Detailed Leaf), a novel approach for fine-scale reconstruction of individual leaves using TLS point clouds. AdLeaf combines wood-leaf separation, individual leaf segmentation, detection and repair of incomplete leaves, explicit reconstruction, and parameter extraction. It automates semantic segmentation at the tree scale to separate woody and leafy components. Instance segmentation is refined through similarity graphs. Incomplete leaves are detected and repaired using shape concavity analysis and symmetry-based mirroring. AdLeaf enables direct measurement of leaf attributes, including count, area, inclination, volume, and azimuth. Validation using field scans, synthetic data, and both in-situ and destructive measurements shows high accuracy: leaf counting errors ranged from 0.58% to 8.23% for trees with 201-4,000 leaves. Reconstructed leaf geometries had mean and standard deviations below 0.83 cm and 0.70 cm, respectively. Leaf area measurements (10–180 cm2) achieved a coefficient of determination (R²) of 0.95, bias of -0.20 cm², and root mean square error of 5.63 cm2. Incomplete leaf detection errors were below 28%, with the repaired area relative RMSE reduced by 9.4%. By addressing QSM limitations, AdLeaf enables explicit 3D leaf reconstructions that support detailed analysis of canopy light interception, spatial heterogeneity, and photosynthesis. It provides a robust framework for linking leaf structure to function at the tree level, advancing forest structure and radiative transfer research. Guangpeng Fan, Liangliang Xu, Jiani Guo, Ruoyoulan Wang, Hao Lu 0004, Jinhu Wang, Di Wang 0006, Feixiang Chen, Liangliang Nan |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2025 | PointCG: Self-Supervised Point Cloud Learning via Joint Completion and GenerationabstractThe core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points' representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance. Yun Liu 0002, Peng Li 0064, Xuefeng Yan 0001, Liangliang Nan, Bing Wang 0013, Honghua Chen, Lina Gong, Wei Zhao 0039, Mingqiang Wei |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | MuVieCAST: Multi-View Consistent Artistic Style TransferabstractWe introduce MuVieCAST, a modular multi-view consistent style transfer network architecture that enables consistent style transfer between multiple viewpoints of the same scene. This network architecture supports both sparse and dense views, making it versatile enough to handle a wide range of multi-view image datasets. The approach consists of three modules that perform specific tasks related to style transfer, namely content preservation, image transformation, and multi-view consistency enforcement. We extensively evaluate our approach across multiple application domains including depth-map-based point cloud fusion, mesh reconstruction, and novel-view synthesis. Our experiments reveal that the proposed framework achieves an exceptional generation of stylized images, exhibiting consistent outcomes across perspectives. A user study focusing on novel-view synthesis further confirms these results, with approximately $68 \%$ of cases participants expressing a preference for our generated outputs compared to the recent state-of-the-art method. Our modular framework is extensible and can easily be integrated with various backbone architectures, making it a flexible solution for multi-view style transfer. More results are demonstrated on our project page: muviecast.github.io Nail Ibrahimli, Julian F. P. Kooij, Liangliang Nan |
3DV | 3 |
| 2024 | On the Estimation of Image-Matching Uncertainty in Visual Place RecognitionabstractIn Visual Place Recognition (VPR) the pose of a query image is estimated by comparing the image to a map of reference images with known reference poses. As is typical for image retrieval problems, a feature extractor maps the query and reference images to a feature space, where a nearest neighbor search is then performed. However, till recently little attention has been given to quantifying the confidence that a retrieved reference image is a correct match. Highly certain but incorrect retrieval can lead to catastrophic failure of VPR-based localization pipelines. This work compares for the first time the main approaches for estimating the image-matching uncertainty, including the traditional retrieval-based uncertainty estimation, more recent data-driven aleatoric uncertainty estimation, and the compute-intensive geometric verification. We further formulate a simple baseline method, “SUE”, which unlike the other methods considers the freely-available poses of the reference images in the map. Our experiments reveal that a simple L2-distance between the query and reference descriptors is already a better estimate of image-matching uncertainty than current data-driven approaches. SUE outperforms the other efficient uncertainty estimation methods, and its uncertainty estimates complement the computationally expensive geometric verification approach. Future works for uncertainty estimation in VPR should consider the baselines discussed in this work. Mubariz Zaffar, Liangliang Nan, Julian F. P. Kooij |
CVPR | 2 |
| 2024 | UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor ModalitiesabstractMulti-sensor object detection is an active research topic in automated driving, but the robustness of such detection models against missing sensor input (modality missing), e.g., due to a sudden sensor failure, is a critical problem which remains under-studied. In this work, we propose UniBEV, an end-to-end multi-modal 3D object detection framework designed for robustness against missing modalities: UniBEV can operate on LiDAR plus camera input, but also on LiDAR-only or camera-only input without retraining. To facilitate its detector head to handle different input combinations, UniBEV aims to create well-aligned Bird’s Eye View (BEV) feature maps from each available modality. Unlike prior BEV-based multi-modal detection methods, all sensor modalities follow a uniform approach to resample features from the original sensor coordinate systems to the BEV features. We furthermore investigate the robustness of various fusion strategies w.r.t. missing modalities: the commonly used feature concatenation, but also channel-wise averaging, and a generalization to weighted averaging termed Channel Normalized Weights. To validate its effectiveness, we compare UniBEV to state-of-the-art BEVFusion and MetaBEV on nuScenes over all sensor input combinations. In this setting, UniBEV achieves better performance than these baselines for all input combinations. An ablation study shows the robustness benefits of fusing by weighted averaging over regular concatenation, and of sharing queries between the BEV encoders of each modality. Our code is available at https://github.com/tudelft-iv/UniBEV. Holger Caesar, Liangliang Nan, Julian F. P. Kooij |
IV | 3 |
| 2024 | PointeNet: A lightweight framework for effective and efficient point cloud analysis
Lipeng Gu, Xuefeng Yan 0001, Liangliang Nan, Dingkun Zhu, Honghua Chen, Weiming Wang 0002, Mingqiang Wei |
Comput. Aided Geom. Des. | 3 |
| 2024 | Fast Building Instance Proxy Reconstruction for Large Urban ScenesabstractDigitalization of large-scale urban scenes (in particular buildings) has been a long-standing open problem, which attributes to the challenges in data acquisition, such as incomplete scene coverage, lack of semantics, low efficiency, and low reliability in path planning. In this paper, we address these challenges in urban building reconstruction from aerial images, and we propose an effective workflow and a few novel algorithms for efficient 3D building instance proxy reconstruction for large urban scenes. Specifically, we propose a novel learning-based approach to instance segmentation of urban buildings from aerial images followed by a voting-based algorithm to fuse the multi-view instance information to a sparse point cloud (reconstructed using a standard Structure from Motion pipeline). Our method enables effective instance segmentation of the building instances from the point cloud. We also introduce a layer-based surface reconstruction method dedicated to the 3D reconstruction of building proxies from extremely sparse point clouds. Extensive experiments on both synthetic and real-world aerial images of large urban scenes have demonstrated the effectiveness of our approach. The generated scene proxy models can already provide a promising 3D surface representation of the buildings in large urban scenes, and when applied to aerial path planning, the instance-enhanced building proxy models can significantly improve data completeness and accuracy, yielding highly detailed 3D building models. Jianwei Guo 0003, Haobo Qin, Yinchang Zhou, Xin Chen 0124, Liangliang Nan, Hui Huang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | PathNet: Path-Selective Point Cloud DenoisingabstractCurrent point cloud denoising (PCD) models optimize single networks, trying to make their parameters adaptive to each point in a large pool of point clouds. Such a denoising network paradigm neglects that different points are often corrupted by different levels of noise and they may convey different geometric structures. Thus, the intricacy of both noise and geometry poses side effects including remnant noise, wrongly-smoothed edges, and distorted shape after denoising. We propose PathNet, a path-selective PCD paradigm based on reinforcement learning (RL). Unlike existing efforts, PathNet enables dynamic selection of the most appropriate denoising path for each point, best moving it onto its underlying surface. We have two more contributions besides the proposed framework of path-selective PCD for the first time. First, to leverage geometry expertise and benefit from training data, we propose a noise- and geometry-aware reward function to train the routing agent in RL. Second, the routing agent and the denoising network are trained jointly to avoid under- and over-smoothing. Extensive experiments show promising improvements of PathNet over its competitors, in terms of the effectiveness for removing different levels of noise and preserving multi-scale surface geometries. Furthermore, PathNet generalizes itself more smoothly to real scans than cutting-edge models. Zeyong Wei, Honghua Chen, Liangliang Nan, Jun Wang 0039, Harry Qin, Mingqiang Wei |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | SimLOG: Simultaneous Local-Global Feature Learning for 3D Object Detection in Indoor Point CloudsabstractThe acquisition of both local and global features from irregular point clouds is crucial for 3D object detection (3DOD). Current mainstream 3D detectors neglect significant local features during pooling operations or disregard many global features of the overall scene context. This paper proposes new techniques for simultaneously learning local-global features of scene point clouds to enhance 3DOD. Specifically, we propose an efficient 3DOD network in indoor point clouds, named SimLOG, which utilizes simultaneous local-global feature learning. SimLOG has two main contributions: a Dynamic Points Interaction (DPI) module to recover local features lost during pooling, and a Global Context Aggregation(GCA) module to aggregate multi-scale features from various layers of the encoder to improve scene context awareness. Unlike traditional local-global feature learning methods, our DPI and GCA modules are integrated into a single feature learning module, making it easily detachable and able to be incorporated into existing 3DOD networks to enhance their performance. SimLOG demonstrates superior performance over twenty competitors in terms of detection accuracy and robustness on both the SUN RGB-D and ScanNet V2 datasets. Specifically, SimLOG boosts the baseline VoteNet by 8.1% of [email protected] on ScanNet V2 and by 3.9% of [email protected] on SUN RGB-D. Code is publicly available athttps://github.com/chenbaian-cs/SimLOG. Mingqiang Wei, Baian Chen, Liangliang Nan, Haoran Xie 0001, Lipeng Gu, Dening Lu, Fu Lee Wang, Qing Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | CSDN: Cross-Modal Shape-Transfer Dual-Refinement Network for Point Cloud CompletionabstractHow will you repair a physical object with some missings? You may imagine its original shape from previously captured images, recover its overall (global) but coarse shape first, and then refine its local details. We are motivated to imitate the physical repair procedure to address point cloud completion. To this end, we propose a cross-modal shape-transfer dual-refinement network (termed CSDN), a coarse-to-fine paradigm with images of full-cycle participation, for quality point cloud completion. CSDN mainly consists of "shape fusion" and "dual-refinement" modules to tackle the cross-modal challenge. The first module transfers the intrinsic shape characteristics from single images to guide the geometry generation of the missing regions of point clouds, in which we propose IPAdaIN to embed the global features of both the image and the partial point cloud into completion. The second module refines the coarse output by adjusting the positions of the generated points, where the local refinement unit exploits the geometric relation between the novel and the input points by graph convolution, and the global constraint unit utilizes the input image to fine-tune the generated offset. Different from most existing approaches, CSDN not only explores the complementary information from images but also effectively exploits cross-modal data in the whole coarse-to-fine completion procedure. Experimental results indicate that CSDN performs favorably against twelve competitors on the cross-modal benchmark. Zhe Zhu, Liangliang Nan, Haoran Xie 0001, Honghua Chen, Jun Wang 0039, Mingqiang Wei, Harry Qin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Symmetrization of 2D Polygonal Shapes Using Mixed-Integer ProgrammingabstractSymmetry widely exists in nature and man-made shapes, but it is unavoidably distorted during the process of growth, design, digitalization, and reconstruction steps. To enhance symmetry, traditional methods follow the detect-then-symmetrize paradigm, which is sensitive to noise in the detection phase, resulting in ambiguities for the subsequent symmetrization step. In this work, we propose a novel optimization-based framework that jointly detects and optimizes symmetry for 2D shapes represented as polygons. Our method can detect and optimize symmetry using a single objective function. Specifically, we formulate symmetry detection and optimization as a mixed-integer program. Our method first generates a set of candidate symmetric edge pairs, which are then encoded as binary variables in our optimization. The geometry of the shape is expressed as continuous variables, which are then optimized together with the binary variables. The symmetry of the shape is enforced by the designed hard constraints. After the optimization, both the optimal symmetric edge correspondences and the geometry are obtained. Our method simultaneously detects all the symmetric primitive pairs and enhances the symmetry of a model while minimally altering its geometry. We have tested our method on a variety of shapes from designs and vectorizations, and the results have demonstrated its effectiveness. Jantien E. Stoter, Liangliang Nan |
Comput. Aided Des. | 3 |
| 2023 | CoPR: Toward Accurate Visual Localization With Continuous Place-Descriptor RegressionabstractVisual place recognition (VPR) is an image-based localization method that estimates the camera location of a query image by retrieving the most similar reference image from a map of geo-tagged reference images. In this work, we look into two fundamental bottlenecks for its localization accuracy: 1) reference map sparseness and 2) viewpoint invariance. First, the reference images for VPR are only available at sparse poses in a map, which enforces an upper bound on the maximum achievable localization accuracy through VPR. We, therefore, propose Continuous Place-descriptor Regression (CoPR) to densify the map and improve localization accuracy. We study various interpolation and extrapolation models to regress additional VPR feature descriptors from only the existing references. Second, we compare different feature encoders and show that CoPR presents value for all of them. We evaluate our models on three existing public datasets and report on average around 30% improvement in VPR-based localization accuracy using CoPR, on top of the 15% increase by using a viewpoint-variant loss for the feature encoder. The complementary relation between CoPR and relative pose estimation is also discussed. Mubariz Zaffar, Liangliang Nan, Julian F. P. Kooij |
IEEE Trans. Robotics | 2 |
| 2022 | Push-the-Boundary: Boundary-aware Feature Propagation for Semantic Segmentation of 3D Point CloudsabstractFeedforward fully convolutional neural networks currently dominate in semantic segmentation of 3D point clouds. Despite their great success, they suffer from the loss of local information at low-level layers, posing significant challenges to accurate scene segmentation and precise object boundary delineation. Prior works either address this issue by post-processing or jointly learn object boundaries to implicitly improve feature encoding of the networks. These approaches often require additional modules which are difficult to integrate into the original architecture. To improve the segmentation near object boundaries, we propose a boundary-aware feature propagation mechanism. This mechanism is achieved by exploiting a multitask learning framework that aims to explicitly guide the boundaries to their original locations. With one shared encoder, our network outputs (i) boundary localization, (ii) prediction of directions pointing to the object's interior, and (iii) semantic segmentation, in three parallel streams. The predicted boundaries and directions are fused to propagate the learned features to refine the segmentation. We conduct extensive experiments on the S3DIS and SensatUrban datasets against various baseline methods, demonstrating that our proposed approach yields consistent improvements by reducing boundary errors. Our code is available at https://github.com/shenglandu/PushBoundary. Shenglan Du, Nail Ibrahimli, Jantien E. Stoter, Julian F. P. Kooij, Liangliang Nan |
3DV | 5 |
| 2022 | 3-D Instance Segmentation of MVS BuildingsabstractWe present a novel 3D instance segmentation framework for Multi-View Stereo (MVS) buildings in urban scenes. Unlike existing works focusing on semantic segmentation of urban scenes, the emphasis of this work lies in detecting and segmenting 3D building instances even if they are attached and embedded in a large and imprecise 3D surface model. Multi-view RGB images are first enhanced to RGBH images by adding a heightmap and are segmented to obtain all roof instances using a fine-tuned 2D instance segmentation neural network. Instance masks from different multi-view images are then clustered into global masks. Our mask clustering accounts for spatial occlusion and overlapping, which can eliminate segmentation ambiguities among multi-view images. Based on these global masks, 3D roof instances are segmented out by mask back-projections and extended to the entire building instances through a Markov random field optimization. A new dataset that contains instance-level annotation for both 3D urban scenes (roofs and buildings) and drone images (roofs) is provided. To the best of our knowledge, it is the first outdoor dataset dedicated for 3D instance segmentation with much more annotations of attached 3D buildings than existing datasets1. Quantitative evaluations and ablation studies have shown the effectiveness of all major steps and the advantages of our multi-view framework over the orthophoto-based method. Jiazhou Chen 0002, Yanghui Xu, Shufang Lu, Ronghua Liang, Liangliang Nan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | HRBF-Fusion: Accurate 3D Reconstruction from RGB-D Data Using On-the-fly ImplicitsabstractReconstruction of high-fidelity 3D objects or scenes is a fundamental research problem. Recent advances in RGB-D fusion have demonstrated the potential of producing 3D models from consumer-level RGB-D cameras. However, due to the discrete nature and limited resolution of their surface representations (e.g., point or voxel based), existing approaches suffer from the accumulation of errors in camera tracking and distortion in the reconstruction, which leads to an unsatisfactory 3D reconstruction. In this article, we present a method using on-the-fly implicits of Hermite Radial Basis Functions (HRBFs) as a continuous surface representation for camera tracking in an existing RGB-D fusion framework. Furthermore, curvature estimation and confidence evaluation are coherently derived from the inherent surface properties of the on-the-fly HRBF implicits, which are devoted to a data fusion with better quality. We argue that our continuous but on-the-fly surface representation can effectively mitigate the impact of noise with its robustness and constrain the reconstruction with inherent surface smoothness when being compared with discrete representations. Experimental results on various real-world and synthetic datasets demonstrate that our HRBF-fusion outperforms the state-of-the-art approaches in terms of tracking robustness and reconstruction accuracy. Yabin Xu, Liangliang Nan, Laishui Zhou, Jun Wang 0039, Charlie C. L. Wang |
ACM Trans. Graph. | 2 |
| 2021 | ComNet: Combinational Neural Network for Object Detection in UAV-Borne Thermal ImagesabstractWe propose a deep learning-based method for object detection in UAV-borne thermal images that have the capability of observing scenes in both day and night. Compared with visible images, thermal images have lower requirements for illumination conditions, but they typically have blurred edges and low contrast. Using a boundary-aware salient object detection network, we extract the saliency maps of the thermal images to improve the distinguishability. Thermal images are augmented with the corresponding saliency maps through channel replacement and pixel-level weighted fusion methods. Considering the limited computing power of UAV platforms, a lightweight combinational neural network ComNet is used as the core object detection method. The YOLOv3 model trained on the original images is used as a benchmark and compared with the proposed method. In the experiments, we analyze the detection performances of the ComNet models with different image fusion schemes. The experimental results show that the average precisions (APs) for pedestrian and vehicle detection have been improved by 2%~5% compared with the benchmark without saliency map fusion and MobileNetv2. The detection speed is increased by over 50%, while the model size is reduced by 58%. The results demonstrate that the proposed method provides a compromise model, which has application potential in UAV-borne detection tasks. Minglei Li 0003, Xingke Zhao, Jiasong Li, Liangliang Nan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | PLADE: A Plane-Based Descriptor for Point Cloud Registration With Small OverlapabstractTraditional point cloud registration methods require large overlap between scans, which imposes strict constraints on data acquisition. To facilitate registration, users have to carefully position scanners to ensure sufficient overlap. In this article, we propose to use high-level structural information (i.e., plane/line features and their interrelationship) for registration, which is capable of registering point clouds with small overlap, allowing more freedom in data acquisition. We design a novel plane-/line-based descriptor dedicated to establishing structure-level correspondences between point clouds. Based on this descriptor, we propose a simple but effective registration algorithm. We also provide a data set of real-world scenes containing a larger number of scans with a wide range of overlap. Experiments and comparisons with state-of-the-art methods on various data sets reveal that our method is superior to existing techniques. Though the proposed algorithm outperforms state-of-the-art methods on the most challenging data set, the point cloud registration problem is still far from being solved, leaving significant room for improvement and future work. Liangliang Nan, Renbo Xia, Peter Wonka |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Near support-free multi-directional 3D printing via global-optimal decomposition
Yisong Gao, Lifang Wu, Dong-Ming Yan 0001, Liangliang Nan |
Graph. Model. | 4 |
| 2017 | PolyFit: Polygonal Surface Reconstruction from Point CloudsabstractWe propose a novel framework for reconstructing lightweight polygonal surfaces from point clouds. Unlike traditional methods that focus on either extracting good geometric primitives or obtaining proper arrangements of primitives, the emphasis of this work lies in intersecting the primitives (planes only) and seeking for an appropriate combination of them to obtain a manifold polygonal surface model without boundary. We show that reconstruction from point clouds can be cast as a binary labeling problem. Our method is based on a hypothesizing and selection strategy. We first generate a reasonably large set of face candidates by intersecting the extracted planar primitives. Then an optimal subset of the candidate faces is selected through optimization. Our optimization is based on a binary linear programming formulation under hard constraints that enforce the final polygonal surface model to be manifold and watertight. Experiments on point clouds from various sources demonstrate that our method can generate lightweight polygonal surface models of arbitrary piecewise planar objects. Besides, our method is capable of recovering sharp features and is robust to noise, outliers, and missing data. Liangliang Nan, Peter Wonka |
ICCV | 1 |
| 2016 | Large Scale Asset Extraction for Urban Images
Lama Affara, Liangliang Nan, Bernard Ghanem, Peter Wonka |
ECCV (3) | 2 |
| 2016 | Manhattan-World Urban Reconstruction from Point Clouds
Minglei Li 0003, Peter Wonka, Liangliang Nan |
ECCV (4) | 3 |
| 2016 | Reconstructing building mass models from UAV images
Minglei Li 0003, Liangliang Nan, Neil Smith, Peter Wonka |
Comput. Graph. | 2 |
| 2016 | Symmetrization of facade layouts
Haiyong Jiang, Dong-Ming Yan 0001, Weiming Dong, Fuzhang Wu, Liangliang Nan, Xiaopeng Zhang 0001 |
Graph. Model. | 5 |
| 2016 | Block assembly for global registration of building scansabstractWe propose a framework for global registration of building scans. The first contribution of our work is to detect and use portals (e.g., doors and windows) to improve the local registration between two scans. Our second contribution is an optimization based on a linear integer programming formulation. We abstract each scan as a block and model the blocks registration as an optimization problem that aims at maximizing the overall matching score of the entire scene. We propose an efficient solution to this optimization problem by iteratively detecting and adding local constraints. We demonstrate the effectiveness of the proposed method on buildings of various styles and that our approach is superior to the current state of the art. Feilong Yan, Liangliang Nan, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2016 | Automatic Constraint Detection for 2D Layout RegularizationabstractIn this paper, we address the problem of constraint detection for layout regularization. The layout we consider is a set of two-dimensional elements where each element is represented by its bounding box. Layout regularization is important in digitizing plans or images, such as floor plans and facade images, and in the improvement of user-created contents, such as architectural drawings and slide layouts. To regularize a layout, we aim to improve the input by detecting and subsequently enforcing alignment, size, and distance constraints between layout elements. Similar to previous work, we formulate layout regularization as a quadratic programming problem. In addition, we propose a novel optimization algorithm that automatically detects constraints. We evaluate the proposed framework using a variety of input layouts from different applications. Our results demonstrate that our method has superior performance to the state of the art. Haiyong Jiang, Liangliang Nan, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Template Assembly for Detailed Urban ReconstructionabstractAbstract We propose a new framework to reconstruct building details by automatically assembling 3D templates on coarse textured building models. In a preprocessing step, we generate an initial coarse model to approximate a point cloud computed using Structure from Motion and Multi View Stereo, and we model a set of 3D templates of facade details. Next, we optimize the initial coarse model to enforce consistency between geometry and appearance (texture images). Then, building details are reconstructed by assembling templates on the textured faces of the coarse model. The 3D templates are automatically chosen and located by our optimization‐based template assembly algorithm that balances image matching and structural regularity. In the results, we demonstrate how our framework can enrich the details of coarse models using various data sets. Liangliang Nan, Caigui Jiang, Bernard Ghanem, Peter Wonka |
Comput. Graph. Forum | 1 |
| 2014 | 2D-D Lifting for Shape ReconstructionabstractAbstract We present an algorithm for shape reconstruction from incomplete 3D scans by fusing together two acquisition modes: 2D photographs and 3D scans. The two modes exhibit complementary characteristics: scans have depth information, but are often sparse and incomplete; photographs, on the other hand, are dense and have high resolution, but lack important depth information. In this work we fuse the two modes, taking advantage of their complementary information, to enhance 3D shape reconstruction from an incomplete scan with a 2D photograph. We compute geometrical and topological shape properties in 2D photographs and use them to reconstruct a shape from an incomplete 3D scan in a principled manner. Our key observation is that shape properties such as boundaries, smooth patches and local connectivity, can be inferred with high confidence from 2D photographs. Thus, we register the 3D scan with the 2D photograph and use scanned points as 3D depth cues for lifting 2D shape structures into 3D. Our contribution is an algorithm which significantly regularizes and enhances the problem of 3D reconstruction from partial scans by lifting 2D shape structures into 3D. We evaluate our algorithm on various shapes which are loosely scanned and photographed from different views, and compare them with state‐of‐the‐art reconstruction methods. Liangliang Nan, Andrei Sharf, Baoquan Chen |
Comput. Graph. Forum | 1 |
| 2012 | A search-classify approach for cluttered indoor scene understandingabstractWe present an algorithm for recognition and reconstruction of scanned 3D indoor scenes. 3D indoor reconstruction is particularly challenging due to object interferences, occlusions and overlapping which yield incomplete yet very complex scene arrangements. Since it is hard to assemble scanned segments into complete models, traditional methods for object recognition and reconstruction would be inefficient. We present a search-classify approach which interleaves segmentation and classification in an iterative manner. Using a robust classifier we traverse the scene and gradually propagate classification information. We reinforce classification by a template fitting step which yields a scene reconstruction. We deform-to-fit templates to classified objects to resolve classification ambiguities. The resulting reconstruction is an approximation which captures the general scene arrangement. Our results demonstrate successful classification and reconstruction of cluttered indoor scenes, captured in just few minutes. Liangliang Nan, Ke Xie 0001, Andrei Sharf |
ACM Trans. Graph. | 1 |
| 2011 | Conjoining Gestalt rules for abstraction of architectural drawingsabstractWe present a method for structural summarization and abstraction of complex spatial arrangements found in architectural drawings. The method is based on the well-known Gestalt rules, which summarize how forms, patterns, and semantics are perceived by humans from bits and pieces of geometric information. Although defining a computational model for each rule alone has been extensively studied, modeling a conjoint of Gestalt rules remains a challenge. In this work, we develop a computational framework which models Gestalt rules and more importantly, their complex interactions. We apply conjoining rules to line drawings, to detect groups of objects and repetitions that conform to Gestalt principles. We summarize and abstract such groups in ways that maintain structural semantics by displaying only a reduced number of repeated elements, or by replacing them with simpler shapes. We show an application of our method to line drawings of architectural models of various styles, and the potential of extending the technique to other computer-generated illustrations, and three-dimensional models. Liangliang Nan, Andrei Sharf, Ke Xie 0001, Tien-Tsin Wong, Oliver Deussen, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 1 |
| 2010 | SmartBoxes for interactive urban reconstructionabstractWe introduce an interactive tool which enables a user to quickly assemble an architectural model directly over a 3D point cloud acquired from large-scale scanning of an urban scene. The user loosely defines and manipulates simple building blocks, which we call SmartBoxes, over the point samples. These boxes quickly snap to their proper locations to conform to common architectural structures. The key idea is that the building blocks are smart in the sense that their locations and sizes are automatically adjusted on-the-fly to fit well to the point data, while at the same time respecting contextual relations with nearby similar blocks. SmartBoxes are assembled through a discrete optimization to balance between two snapping forces defined respectively by a data-fitting term and a contextual term, which together assist the user in reconstructing the architectural model from a sparse and noisy point cloud. We show that a combination of the user's interactive guidance and high-level knowledge about the semantics of the underlying model, together with the snapping forces, allows the reconstruction of structures which are partially or even completely missing from the input. Liangliang Nan, Andrei Sharf, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 1 |