EDBT 2026 Demo / reviewers in the wild / expert
Lun Luo
dblp:189/3267
· DBLP profile ↗
17ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-0531-9171ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dust to Tower: Prior-Driven Coarse-to-Fine Photo-Realistic Scene Reconstruction From Sparse Uncalibrated ImagesabstractPhoto-realistic scene reconstruction from sparse-view, uncalibrated images is highly required in practice. Although some successes have been made, existing methods are either Sparse-View but require accurate camera parameters (i.e., intrinsic and extrinsic), or SfM-free but need densely captured images. This paper proposes Dust to Tower (D2T), a novel coarse-to-fine framework to address the coupled difficulty. The key idea is to explicitly narrow down the solution space and then introduce reliable supervision at novel viewpoints without resorting to expensive diffusion-based view synthesis. To do this, we first introduce a Coarse Construction Module (CCM) which exploits a fast Multi-View Stereo model to initialize a 3D Gaussian Splatting (3DGS) and recover initial camera poses. To refine the 3D model at novel viewpoints, we introduce Confidence-Aware Depth Alignment (CADA), which aligns a monocular inverse-depth prior to the reliable regions of the coarse depth using DUSt3R confidence, producing sharp and scale-consistent depth maps for accurate warping. We further propose Warped Image-Guided Inpainting (WIGI), which converts the accurate warped views into multi-view-consistent pseudo supervision via elaborate warping and inpainting process. Experiments on three benchmark datasets show that D2T achieves superior novel view synthesis quality and pose accuracy over ten representative baselines, while keeping high efficiency. Yongcai Wang, Zhaoxin Fan, Shuo Wang 0015, Deying Li 0001, Lun Luo, Minhang Wang, Hongyuan Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | InsCMPR: Efficient Cross-Modal Place Recognition via Instance-Aware Hybrid Mamba-TransformerabstractPlace recognition is an important technique for autonomous mobile robotic applications. While single-modal sensor-based approaches have shown satisfactory performance, cross-modal place recognition remains underexplored due to the challenge of bridging the cross-modal heterogeneity gap. In this work, we introduce an instance-aware cross-modal place recognition approach, named InsCMPR. We design a novel instance-aware modality alignment module, which aligns multi-modal data at both pixel-level and instance-level by leveraging a pre-trained vision foundation model SAM. Then a novel dual-branch hybrid Mamba-Transformer network is proposed to efficiently enhance the distinctiveness of the produced descriptors by integrating global features with local instance features. Experimental results on the KITTI, NCLT, and HAOMO datasets show that our proposed methods achieve state-of-the-art performance while operating in real time. We will open source the implementation of our method at: https://github.com/nubot-nudt/InsCMPR. Shuaifeng Jiao, Zhuoqun Su, Lun Luo, Hongshan Yu, Zongtan Zhou, Huimin Lu 0002, Xieyuanli Chen |
ICRA | 3 |
| 2025 | Image-Goal Navigation Using Refined Feature Guidance and Scene Graph EnhancementabstractIn this paper, we introduce a novel image-goal navigation approach, named RFSG. Our focus lies in leveraging the fine-grained connections between goals, observations, and the environment within limited image data, all the while keeping the navigation architecture simple and lightweight. To this end, we propose the spatial-channel attention mechanism, enabling the network to learn the importance of multi-dimensional features to fuse the goal and observation features. In addition, a self-distillation mechanism is incorporated to further enhance the feature representation capabilities. Given that the navigation task needs surrounding environmental information for more efficient navigation, we propose an image scene graph to establish feature associations at both the image and object levels, effectively encoding the surrounding scene information. Cross-scene performance validation was conducted on the Gibson and HM3D datasets, and the proposed method achieved state-of-the-art results among mainstream methods, with a speed of up to 53.5 frames per second on an RTX3080. This contributes to the realization of end-to-end image-goal navigation in real-world scenarios. The implementation and model of our method have been released at: https://github.com/nubot-nudt/RFSG. Zhicheng Feng, Xieyuanli Chen, Chenghao Shi, Lun Luo, Zhichao Chen 0002, Huimin Lu 0002 |
IROS | 4 |
| 2025 | BEVPlace++: Fast, Robust, and Lightweight LiDAR Global Localization for Autonomous Ground VehiclesabstractThis article introduces BEVPlace++, a novel, fast, and robust LiDAR global localization method for unmanned ground vehicles. It uses lightweight convolutional neural networks (CNNs) on Bird's Eye View (BEV) image-like representations of LiDAR data to achieve accurate global localization through place recognition, followed by 3-DoF pose estimation. Our detailed analyses reveal an interesting fact that CNNs are inherently effective at extracting distinctive features from LiDAR BEV images. Remarkably, keypoints of two BEV images with large translations can be effectively matched using CNN-extracted features. Building on this insight, we design a Rotation Equivariant Module (REM) to obtain distinctive features while enhancing robustness to rotational changes. A Rotation Equivariant and Invariant Network (REIN) is then developed by cascading REM and a descriptor generator, NetVLAD, to sequentially generate rotation equivariant local features and rotation invariant global descriptors. The global descriptors are used first to achieve robust place recognition, and then local features are used for accurate pose estimation. Experimental results on seven public datasets and our UGV platform demonstrate that BEVPlace++, even when trained on a small dataset (3000 frames of KITTI) only with place labels, generalizes well to unseen environments, performs consistently across different days and years, and adapts to various types of LiDAR scanners. BEVPlace++ achieves state-of-theart performance in multiple tasks, including place recognition, loop closure detection, and global localization. Additionally, BEVPlace++ is lightweight, runs in real-time, and does not require accurate pose supervision, making it highly convenient for deployment. The source codes are publicly available at https://github.com/zjuluolun/BEVPlace2. Lun Luo, Si-Yuan Cao, Jintao Xu 0001, Rui Ai 0001, Zhu Yu 0001, Xieyuanli Chen |
IEEE Trans. Robotics | 1 |
| 2024 | SCPNet: Unsupervised Cross-Modal Homography Estimation via Intra-modal Self-supervised Learning
Runmin Zhang, Si-Yuan Cao, Lun Luo, Beinan Yu, Shujie Chen 0001, Junwei Li 0009 |
ECCV (23) | 4 |
| 2024 | ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place RecognitionabstractPlace recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a challenging problem. Current cross-modal methods transform images into 3D points using depth estimation for modality conversion, which are usually computationally intensive and need expensive labeled data for depth supervision. In this work, we introduce a fast and lightweight framework to encode images and point clouds into place-distinctive descriptors. We propose an effective Field of View (FoV) transformation module to convert point clouds into an analogous modality as images. This module eliminates the necessity for depth estimation and helps subsequent modules achieve real-time performance. We further design a non-negative factorization-based encoder to extract mutually consistent semantic features between point clouds and images. This encoder yields more distinctive global descriptors for retrieval. Experimental results on the KITTI dataset show that our proposed methods achieve state-of-the-art performance while running in real time. Additional evaluation on the HAOMO dataset covering a 17 km trajectory further shows the practical generalization capabilities. We have released the implementation of our methods as open source at: https://github.com/haomo-ai/ModaLink.git. Weidong Xie, Lun Luo, Nanfei Ye, Shaoyi Du, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Xieyuanli Chen |
IROS | 2 |
| 2024 | PRISM: PRogressive dependency maxImization for Scale-invariant image MatchingabstractImage matching aims at identifying corresponding points between a pair of images. Currently, detector-free methods have shown impressive performance in challenging scenarios, thanks to their capability of generating dense matches and global receptive field. However, performing feature interaction and proposing matches across the entire image is unnecessary, because not all image regions contribute to the matching process. Interacting and matching in unmatchable areas can introduce errors, reducing matching accuracy and efficiency. Meanwhile, the scale discrepancy issue still troubles existing methods. To address above issues, we propose PRogressive dependency maxImization for Scale-invariant image Matching (PRISM), which jointly prunes irrelevant patch features and tackles the scale discrepancy. To do this, we firstly present a Multi-scale Pruning Module (MPM) to adaptively prune irrelevant features by maximizing the dependency between the two feature sets. Moreover, we design the Scale-Aware Dynamic Pruning Attention (SADPA) to aggregate information from different scales via a hierarchical design. Our method's superior matching performance and generalization capability are confirmed by leading accuracy across various evaluation benchmarks and downstream tasks. The code is publicly available at https://github.com/Master-cai/PRISM. Yongcai Wang, Lun Luo, Minhang Wang, Deying Li 0001, Jintao Xu 0001, Weihao Gu, Rui Ai 0001 |
ACM Multimedia | 3 |
| 2024 | Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionabstractVision-based Semantic Scene Completion (SSC) has gained much attention due to its widespread applications in various 3D perception tasks. Existing sparse-to-dense approaches typically employ shared context-independent queries across various input images, which fails to capture distinctions among them as the focal regions of different inputs vary and may result in undirected feature aggregation of cross-attention. Additionally, the absence of depth information may lead to points projected onto the image plane sharing the same 2D position or similar sampling points in the feature map, resulting in depth ambiguity. In this paper, we present a novel context and geometry aware voxel transformer. It utilizes a context aware query generator to initialize context-dependent queries tailored to individual input images, effectively capturing their unique characteristics and aggregating information within the region of interest. Furthermore, it extend deformable cross-attention from 2D to 3D pixel space, enabling the differentiation of points with similar image coordinates based on their depth coordinates. Building upon this module, we introduce a neural network named CGFormer to achieve semantic scene completion. Simultaneously, CGFormer leverages multiple 3D representations (i.e., voxel and TPV) to boost the semantic and geometric representation abilities of the transformed 3D volume from both local and global perspectives. Experimental results demonstrate that CGFormer achieves state-of-the-art performance on the SemanticKITTI and SSCBench-KITTI-360 benchmarks, attaining a mIoU of 16.87 and 20.05, as well as an IoU of 45.99 and 48.07, respectively. Remarkably, CGFormer even outperforms approaches employing temporal images as inputs or much larger image backbone networks. Zhu Yu 0001, Runmin Zhang, Jiacheng Ying, Junchen Yu, Xiaohai Hu, Lun Luo, Si-Yuan Cao |
NeurIPS | 6 |
| 2024 | A Clustering Election Game-Based and Two-Level Management Protocol for Wireless Sensor NetworksabstractEnergy balance consumption is an important research topic in the field of wireless sensor networks (WSNs). Clustering protocols are widely used to reduce WSN energy consumption. However, this many-to-one cluster head (CH) model can also easily lead to energy exhaustion of some critical nodes near the base station results in the decline of the network quality and lifetime. This article proposes a clustering election game-based and two-level management clustering protocol (CEGT) for WSNs; the network is divided into layers of equal size, and the layer heads (LHs) are selected within the layers based on the location and energy information. Each layer is further divided into clusters of equal size, and the nodes within the clusters undergo a two-stage screening followed by local clustering and multiple rounds of election games to dynamically select the staged CHs. This protocol effectively prevents low-energy nodes from participating in the CH competitions, and selects high-quality nodes through election games which are based on customized qualification values, and ensures reasonable energy utilization as well as efficient network load balancing. Finally, our simulations show that our CEGT protocol outperforms other clustering protocols in terms of the network lifetime, the number of surviving nodes, and the average number of clusters, and the throughput, which also indicate that the CEGT protocol is more suitable for large-scale networks. Lingru Cai, Ruisong Huang, Zhangjie Li, Lun Luo, Zhi Xiong 0001, Yindong Chen |
IEEE Internet Things J. | 4 |
| 2024 | Game-Based Dynamic Clustering Routing Strategy for Mobile Wireless Sensor NetworksabstractIn mobile wireless sensor networks, the mobility of nodes imposes heightened energy management demands on the network. This study proposes a game-based dynamic clustering routing (GDCR) protocol for mobile wireless sensor networks. The GDCR protocol employs multidimensional clustering and constructs a mixed strategy game model among heterogeneous nodes. This model considers various factors for cluster head selection, including residual energy, movement, distance to the sink, and node spacing, ensuring comprehensive decision-making. In addition, the protocol incorporates detached node management to ensure network security and stability. Simulation experiments demonstrate that the proposed protocol reduces the average network energy consumption when nodes have unrestricted movements. Furthermore, it effectively curtails fluctuations in energy consumption as the network scales up, thereby prolonging network lifetime and enhancing stability. Lingru Cai, Lun Luo, Zhangjie Li, Zhi Xiong 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerabstractWe propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the attention-focusing mechanism in FocusFormer aggregates the intra-inter correspondence in a global→nonlocal→local manner. Thanks to the above strategies, RHWF ranks top in accuracy on a variety of datasets, including the challenging cross-resolution and cross-modal ones. Meanwhile, benefiting from the recurrent framework, RHWF achieves parameter efficiency despite the transformer architecture. Compared to previous state-of-the-art approaches LocalTrans and IHN, RHWF reduces the mean average corner error (MACE) by about 70% and 38.1% on the MSCOCO dataset, while saving the parameter costs by 86.5% and 24.6%. Similar to the previous works, RHWF can also be arranged in 1-scale for efficiency and 2-scale for accuracy, with the 1-scale RHWF already outperforming most of the previous methods. Source code is available at https://github.com/imdump178/RHWF. Si-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu, Zehua Sheng, Junwei Li 0009 |
CVPR | 3 |
| 2023 | BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesabstractPlace recognition is a key module for long-term SLAM systems. Current LiDAR-based place recognition methods usually use representations of point clouds such as unordered points or range images. These methods achieve high recall rates of retrieval, but their performance may degrade in the case of view variation or scene changes. In this work, we explore the potential of a different representation in place recognition, i.e. bird’s eye view (BEV) images. We validate that, in scenes of slight viewpoint changes, a simple NetVLAD network trained on BEV images achieves comparable performance to the state-of-the-art place recognition methods. For robustness to view variations, we propose a rotation-invariant network called BEVPlace. We use group convolution to extract rotation-equivariant local features from the images and NetVLAD for global feature aggregation. In addition, we observe that the distance between BEV features is correlated with the geometry distance of point clouds. Based on the observation, we develop a method to estimate the position of the query cloud, extending the usage of place recognition. The experiments conducted on large-scale public datasets show that our method 1) achieves state-of-the-art performance in terms of recall rates, 2) is robust to view changes, 3) shows strong generalization ability, and 4) can estimate the positions of query point clouds. Source codes are publicly available at https://github.com/zjuluolun/BEVPlace. Lun Luo, Shuhang Zheng, Yongzhi Fan, Beinan Yu, Si-Yuan Cao, Junwei Li 0009 |
ICCV | 1 |
| 2023 | Aggregating Feature Point Cloud for Depth CompletionabstractGuided depth completion aims to recover dense depth maps by propagating depth information from the given pixels to the remaining ones under the guidance of RGB images. However, most of the existing methods achieve this using a large number of iterative refinements or stacking repetitive blocks. Due to the limited receptive field of conventional convolution, the generalizability with respect to different sparsity levels of input depth maps is impeded. To tackle these problems, we propose a feature point cloud aggregation framework to directly propagate 3D depth information between the given points and the missing ones. We extract 2D feature map from images and transform the sparse depth map to point cloud to extract sparse 3D features. By regarding the extracted features as two sets of feature point clouds, the depth information for a target location can be reconstructed by aggregating adjacent sparse 3D features from the known points using cross attention. Based on this, we design a neural network, called as PointDC, to complete the entire depth information reconstruction process. Experimental results show that, our PointDC achieves superior or competitive results on the KITTI benchmark and NYUv2 dataset. In addition, the proposed PointDC demonstrates its higher generalizability to different sparsity levels of the input depth maps and cross-dataset evaluation. Zhu Yu 0001, Zehua Sheng, Lun Luo, Si-Yuan Cao, Huaqi Zhang |
ICCV | 4 |
| 2023 | I2P-Rec: Recognizing Images on Large-Scale Point Cloud Maps Through Bird's Eye View ProjectionsabstractPlace recognition is an important technique for autonomous cars to achieve full autonomy since it can provide an initial guess to online localization algorithms. Although current methods based on images or point clouds have achieved satisfactory performance, localizing the images on a large-scale point cloud map remains a fairly unexplored problem. This cross-modal matching task is challenging due to the difficulty in extracting consistent descriptors from images and point clouds. In this paper, we propose the I2P-Rec method to solve the problem by transforming the cross-modal data into the same modality. Specifically, we leverage on the recent success of depth estimation networks to recover point clouds from images. We then project the point clouds into Bird's Eye View (BEV) images. Using the BEV image as an intermediate representation, we extract global features with a Convolutional Neural Network followed by a NetVLAD layer to perform matching. The experimental results evaluated on the KITTI dataset show that, with only a small set of training data, I2P-Rec achieves recall rates at Top-l % over 80% and 90%, when localizing monocular and stereo images on point cloud maps, respectively. We further evaluate I2P-Rec on a 1 km trajectory dataset collected by an autonomous logistics car and show that I2P- Rec can generalize well to previously unseen environments. Shuhang Zheng, Zhu Yu 0001, Beinan Yu, Si-Yuan Cao, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Lun Luo |
IROS | 10 |
| 2022 | FreSCo: Frequency-Domain Scan Context for LiDAR-based Place Recognition with Translation and Rotation InvarianceabstractPlace recognition plays a crucial role in relocalization and loop closure detection tasks for robots and vehicles. This paper seeks a well-defined global descriptor for LiDAR-based place recognition. Compared to local descriptors, global descriptors show remarkable performance in urban road scenes but are usually viewpoint-dependent. To this end, we propose a simple yet robust global descriptor dubbed FreSCo that decomposes the viewpoint difference during revisit and achieves both translation and rotation invariance by leveraging Fourier Transform and circular shift technique. Besides, a fast two-stage pose estimation method is proposed to estimate the relative pose after place retrieval by utilizing the compact 2D point clouds extracted from the original data. Experiments show that FreSCo exhibited superior performance than contemporaneous methods on sequences of different scenes from multiple datasets. Code will be publicly available at https://github.com/soytony/FreSCo. Yongzhi Fan, Xin Du 0005, Lun Luo, Jizhong Shen |
ICARCV | 3 |
| 2017 | Joint Enhancing Filtering for Road Network ExtractionabstractIn this paper, we propose a task-oriented enhancing technique for extracting road networks from satellite images. By exploiting an approximate estimation of the potential road edges for guidance, we developed a joint enhancing filtering framework to generate a version of the input image that facilitates road network extraction. First, an adaptive smoothing scheme is designed to suppress the interference of noise or heavy textures, such as residential areas or terrain boundaries. By combining this scheme with the proposed novel anisotropic shock filter, the edges of the potential road regions can be kept sharp and clear. Through abundant experimental comparisons with state-of-the-art filtering techniques and quantitative evaluations using data from various satellite sensors, the performance of the proposed approach is comprehensively evaluated. The experimental results demonstrate that our system can address heavy high contrast textures and provide a meaningful improvement in the feature detection for road extraction. Cheng Wang 0003, Lun Luo, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | Road network extraction via deep learning and line integral convolutionabstractIn this paper, we propose a learning-based road network extraction scheme from high resolution satellite. First, the convolutional neural network (CNN), which is able to capture large context of local structures, are applied to predict the probability of a pixel belonging to road regions, and assign labels to each pixel to describe whether it is road. Then, a line integral convolution based algorithm is developed to smooth the rough map to connect small gaps. Finally, by combining with some common image processing operators, road centerlines are able to be acquired. Attribute to the learning capacity of CNN, and the line integral convolution based connection scheme, the proposed road extraction method is able to provide high quality results comparing to current state-of-art road extraction methods. Peikang Li, Cheng Wang 0003, Jonathan Li 0001, Ming Cheng 0002, Lun Luo |
IGARSS | 6 |