EDBT 2026 Demo / reviewers in the wild / expert
Li Yan 0003
dblp:71/7028-3
· DBLP profile ↗
23ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-5507-810XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HFCH: Hybrid frontier guided fast UAV autonomous exploration for complete and high-quality mapping in unknown environment
Yuquan Zhou, Li Yan 0003, Yaxi Han, Longze Zhu, Hong Xie 0002 |
Adv. Eng. Informatics | 2 |
| 2026 | HALE: a hierarchical autonomous exploration framework for UAVs with limited field-of-view in large-span environments
Yuquan Zhou, Li Yan 0003, Longze Zhu, Yaxi Han, Yukang Liu, Hong Xie 0002 |
Expert Syst. Appl. | 2 |
| 2026 | HELF-SLAM: A hybrid-enhanced learning-based feature for distortion-resilient monocular SLAM
Longze Zhu, Li Yan 0003, Hong Xie 0002, Xiaoteng Yang, Jiang Song, Linxia Ji, Aoran Li |
Neurocomputing | 2 |
| 2026 | MAHF-LIO: Motion-Aware Hierarchical Fusion for Robust LiDAR-Inertial OdometryabstractAccurate and robust localization is critical for Internet of Things (IoT)-enabled autonomous vehicles and intelligent mobile robots operating in complex and dynamic environments. However, existing learning-based LiDAR-inertial odometry (LIO) methods often underutilize the complementarity between LiDAR and IMU measurements, as they typically rely on either simple weighted fusion strategies or computationally intensive attention-based interaction mechanisms, without explicitly modeling hierarchical cross-modal interactions. To address this limitation, we propose MAHF-LIO, a motion-aware hierarchical fusion framework that combines local motion-conditioned modulation with global lightweight feature mixing for robust and accurate LIO. Specifically, a Motion-Conditioned Gated Local Modulation (MCGLM) module uses IMU-derived temporal motion cues to adaptively recalibrate local LiDAR geometric features, improving feature reliability under varying motion conditions. In addition, a Lightweight Global MLP-Mixer Fusion (LGMMF) module performs efficient cross-modal feature integration to capture long-range dependencies between LiDAR and IMU representations with modest computational cost. Extensive experiments on both public and self-collected datasets show that MAHF-LIO delivers accurate and robust localization across diverse environments. On the KITTI benchmark, MAHF-LIO reduces the average relative translational and rotational errors by 7.1% and 29.8%, respectively, compared with Adaptive-LIO. It also demonstrates greater robustness than learning-based baselines under challenging scenarios, including rapid motion and perceptual degradation. Ablation studies further verify the effectiveness of the proposed hierarchical fusion design. Xin Yang 0039, Li Yan 0003, Hong Xie 0002, Xiaohu Lin, Longze Zhu, Aiqiang Ma |
IEEE Internet Things J. | 2 |
| 2025 | VRGNet: A Relative Geometric-Driven Network for Point Cloud Registration with Virtual Correspondences
Li Yan 0003, Changjun Chen, Binbing Wang, Yuquan Zhou, Yiming Yang 0002 |
PRICAI (5) | 2 |
| 2025 | SED-SLAM: Enhancing Monocular SLAM Under Image Distortions via Spatially Equalized Deep Feature
Longze Zhu, Li Yan 0003, Hong Xie 0002, Xiaoteng Yang, Aoran Li |
PRICAI (5) | 2 |
| 2025 | MaCon: A Generic Self-Supervised Framework for Unsupervised Multimodal Change DetectionabstractChange detection(CD) is important for Earth observation, emergency response and time-series understanding. Recently, data availability in various modalities has increased rapidly, and multimodal change detection (MCD) is gaining prominence. Given the scarcity of datasets and labels for MCD, unsupervised approaches are more practical for MCD. However, previous methods typically either merely reduce the gap between multimodal data through transformation or feed the original multimodal data directly into the discriminant network for difference extraction. The former faces challenges in extracting precise difference features. The latter contains the pronounced intrinsic distinction between the original multimodal data; direct extraction and comparison of features usually introduce significant noise, thereby compromising the quality of the resultant difference image. In this article, we proposed the MaCon framework to synergistically distill the common and discrepancy representations. The MaCon framework unifies mask reconstruction (MR) and contrastive learning (CL) self-supervised paradigms, where the MR serves the purpose of transformation while CL focuses on discrimination. Moreover, we presented an optimal sampling strategy in the CL architecture, enabling the CL subnetwork to extract more distinguishable discrepancy representations. Furthermore, we developed an effective silent attention mechanism that not only enhances contrast in output representations but stabilizes the training. Experimental results on both multimodal and monomodal datasets demonstrate that the MaCon framework effectively distills the intrinsic common representations between varied modalities and manifests state-of-the-art performance across both multimodal and monomodal CD. Such findings imply that the MaCon possesses the potential to serve as a unified framework in the CD and relevant fields. Source code will be publicly available once the article is accepted. Jian Wang 0138, Li Yan 0003, Jianbing Yang, Hong Xie 0002, Qiangqiang Yuan, Pengcheng Wei, Zhao Gao, Ce Zhang 0005, Peter M. Atkinson |
IEEE Trans. Image Process. | 2 |
| 2024 | Unsupervised Multimodal Change Detection by Distilling Common and Discrepant RepresentationsabstractChange detection (CD) has become increasingly important in remote sensing and Earth observation. Currently, the data in various modalities has rapidly increased, and multimodal change detection is gaining prominence and holds substantial potential for applications demanding high temporal frequency or rapid response. In this research, we proposed a novel CDR-Net architecture for unsupervised multimodal change detection. The CDR-Net fuses the merits of the mask reconstruction and contrastive learning self-supervised paradigm. Within this architecture, the mask reconstruction subnetwcork pays more attention to low-level details, distilling common representations between multimodal remote sensing images to make them comparable, while the CL subnetwork emphasizes high-level semantics, extracting discrepant representations to facilitate the change detection task. Experimental results demonstrated that the CDR-Net achieved outstanding performance. This implies that the CDR-Net is of great value for resource investigation, emergency response and time-series understanding. Jian Wang 0138, Li Yan 0003, Hong Xie 0002, Tingyuan Zhou, Wenxu Shi, Peter M. Atkinson |
IGARSS | 2 |
| 2024 | ESR-DMNet: Enhanced Super-Resolution-Based Dual-Path Metric Change Detection Network for Remote Sensing Images With Different ResolutionsabstractRemote sensing change detection has always been one of the research hot issues in remote sensing. Current research focuses on studying deep learning change detection methods for remote sensing images with the same resolution. With the prevalence of multi-resolution remote sensing images, how to effectively utilize remote sensing images with different resolutions for change detection is a key issue. To solve this problem, this paper proposes an enhanced super-resolution-based dual-path metric change detection network (ESR-DMNet) to realize high-accuracy and high-efficiency end-to-end change detection of remote sensing images with different resolutions. ESR-DMNet provides a new enhanced super-resolution module for the change detection of remote sensing images with different resolutions, which can perceptively reconstruct low-resolution images into more realistic high-resolution images. ESR-DMNet proposes an effective and efficient dual-path metric change detection network, which processes shallow spatial details information and deep semantic information separately to achieve high-accuracy and high-efficiency change detection. Compared with nine state-of-the-art methods, our method shows good performance at three resolutions on three datasets, SYSU, CDD and CLCD, confirming its potential for change detection tasks in remote sensing images with different resolutions. Xi Li 0018, Li Yan 0003, Yi Zhang 0132, Huaien Zeng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Micro-Structures Graph-Based Point Cloud Registration for Balancing Efficiency and AccuracyabstractPoint cloud registration (PCR) is a fundamental and significant issue in photogrammetry and remote sensing, aiming to seek the optimal rigid transformation between sets of points. Achieving efficient and precise PCR poses a considerable challenge. We propose a novel micro-structures graph-based global PCR method. The overall method is comprised of two stages. 1) Coarse registration (CR): We develop a graph incorporating micro-structures, employing an efficient graph-based hierarchical strategy to remove outliers for obtaining the maximal consensus set. We propose a robust GNC-Welsch estimator for optimization derived from a robust estimator to the outlier process in the Lie algebra space, achieving fast and robust alignment. 2) Fine registration (FR): To refine local alignment further, we use the octree approach to adaptive search plane features in the micro-structures. By minimizing the distance from the point-to-plane, we can obtain a more precise local alignment, and the process will also be addressed effectively by being treated as a planar adjustment (PA) algorithm combined with Anderson accelerated (PA-AA) optimization. After extensive experiments on real data, our proposed method performs well on the 3DMatch and ETH datasets compared to the most advanced methods, achieving higher accuracy metrics and reducing the time cost by at least one-third. Rongling Zhang, Li Yan 0003, Pengcheng Wei, Hong Xie 0002, Pinzhuo Wang, Binbing Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | TSTD:A Cross-modal Two Stages Network with New Trans-decoder for Point Cloud Semantic Segmentation
Zhao Gao, Li Yan 0003, Hong Xie 0002, Pengcheng Wei, Jian Wang 0138 |
PRCV (8) | 2 |
| 2023 | A Voxel-Based Multiview Point Cloud Refinement Method via Factor Graph Optimization
Li Yan 0003, Hong Xie 0002, Pengcheng Wei, Jicheng Dai, Zhao Gao, Rongling Zhang |
PRCV (2) | 2 |
| 2023 | A New Outlier Removal Strategy Based on Reliability of Correspondence Graph for Fast Point Cloud RegistrationabstractRegistration is a basic yet crucial task in point cloud processing. In correspondence-based point cloud registration, matching correspondences by point feature techniques may lead to an extremely high outlier (false correspondence) ratio. Current outlier removal methods still suffer from low efficiency, accuracy, and recall rate. We use an intuitive method to describe the 6-DOF (degree of freedom) curtailment process in point cloud registration and propose an outlier removal strategy based on the reliability of the correspondence graph. The method constructs the corresponding graph according to the given correspondences and designs the concept of the reliability degree of the graph node for optimal candidate selection and the reliability degree of the graph edge to obtain the global maximum consensus set. The presented method achieves fast and accurate outliers removal along with gradual aligning parameters estimation. Extensive experiments on simulations and challenging real-world datasets demonstrate that the proposed method can still perform effective point cloud registration even the correspondence outlier ratio is over 99%, and the efficiency is better than the state-of-the-art. Code is available at https://github.com/WPC-WHU/GROR. Li Yan 0003, Pengcheng Wei, Hong Xie 0002, Jicheng Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | UBMDP: Urban Building Mesh Decoupling and PolygonizationabstractWith the development of photogrammetry, digital city, and metaverse, the 3-D representation of urban buildings has attracted more and more attention. As the main form of the 3-D urban building model, the triangular mesh model has deficiencies such as high complexity, high-data volume, and low-structural information, which seriously restrict its application in spatial analysis and urban planning. This article proposes a hybrid modeling strategy geared toward the mesh model generated from oblique images to obtain building models that are compact, manifold, watertight, and have certain structural and semantic information. First of all, when the planar region topology graph has been established, a topology decoupling strategy is designed to obtain a set of relatively independent topology subgraphs which form a hierarchical structure. After that, to improve model quality, topology optimization of parallel planes has also been studied systematically. Then, we adopt a divide-and-conquer strategy to perform data-driven and model-driven building modeling for the primary and ancillary structures. Finally, a component-level simple polygon model combination is generated. Experiments prove that the proposed method has excellent visual authenticity, structural completeness advantages, and decent LoD3 ability. As a mesh simplification method, the data is compressed to 0.11%–0.75% in a Hausdorff metric around 0.3 m, which further proves that this method is state-of-the-art. Li Yan 0003, Yao Li 0030, Jicheng Dai, Hong Xie 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | SDMNet: A Deep-Supervised Dual Discriminative Metric Network for Change Detection in High-Resolution Remote Sensing ImagesabstractHigh-resolution remote sensing image change detection (CD) is one of the main methods to analyze land surface changes. How to effectively distinguish interesting changes and pseudochanges in high-resolution remote sensing images and form accurate and robust CD results is crucial. To deal with these problems, in this letter, we propose a deep-supervised dual discriminative metric network (SDMNet) trained end-to-end for CD in high-resolution bitemporal remote sensing images. A discriminative decoder network is designed in SDMNet to aggregate global and multiscale contextual information, utilizing high stage features to guide the selection of low stage features stage-by-stage to obtain more consistent and robust features. A discriminative implicit metric module is designed in SDMNet to measure the distance between features to detect changes and utilize batch-balanced contrastive loss (BCL) to enlarge the distance difference between unchanged pairs and changed pairs, while alleviate the problem of sample imbalance, and multiple change graph losses are introduced in the intermediate layer of the network for deep supervision. Quantitative evaluation of our method on the CD datasets SYSU, DSIFN, and CLCD demonstrates that the proposed method can provide superior performance than the other state-of-the-art methods. Xi Li 0018, Li Yan 0003, Yi Zhang 0132, Nan Mo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A DSM-Based Co-Occurrence Matrix for Semantic ClassificationabstractTraditional 2D textures cannot reflect objects’ real textures in 3D world, since they only consider spectral distribution in a 2D region which is a projection of 3D objects at a certain angle of view. The existing researches of 3D textures can only process volumetric data (VD) like multi/hyper-spectral images which are not real 3D geometric data. In this letter, we proposed a digital surface model (DSM)-based co-occurrence matrix (DSMB-CM) which extended 2D co-occurrence matrix (2D-CM) to 3D space for multispectral images with DSM. DSMB-CM is the first 3D feature in remote sensing areas considering the spectral distribution over 3D surface to represent real textures of objects in 3D space. Besides, a dimension reduction method was proposed to avoid curse of dimensionality. Experiments compared classification accuracies of different feature combinations of two data sets from ISPRS Benchmark of Semantic Labeling Contest. The results proved that DSMB-CM had better performance than traditional textures in identification of all categories. Li Yan 0003, Hong Xie 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Deriving Mining-Induced 3-D Deformations at Any Moment and Assessing Building Damage by Integrating Single InSAR Interferogram and Gompertz Probability Integral Model (SII-GPIM)abstractIt is necessary to timely and accurately estimate the surface deformations in mining areas, especially the three-dimensional (3D) deformations during surface movement. At present, nearly all mining-induced 3D deformations retrieved by interferometric synthetic aperture radar (InSAR) pertain to the SAR imaging interval. Research on progressive 3D deformations during surface movement is limited, and the existing approaches are unsatisfactory in practical engineering. Aiming at these challenges, we proposed a novel method for deriving mining-induced 3D surface deformations at any moment by integrating single InSAR interferogram (SII), the Gompertz time function, and the probability integral model (PIM), named the SII-GPIM method. We established an inversion approach for GPIM parameters and derived the mining-induced 3D surface deformations at any moment. Subsequently, we conducted experiments considering two ALOS PALSAR images in the Huaibei mining area. The accuracy of the proposed method was evaluated in subsidence, tilt, curvature, horizontal displacement, and horizontal strain. Compared with existing methods, the SII-GPIM method is state of the art. Additionally, we assessed the building damage, performance of parameter inversion, and method generality. The results demonstrated that the proposed method can accurately determine the mining-induced 3D surface deformations and deformation level at any moment under different geological mining conditions. Moreover, accurate GPIM parameters can be acquired with only two SAR images and traditional measurement is nearly not required. Consequently, the SII-GPIM method owns great value for improving economic efficiency, assessing building damage, and restoring the ecological environment in the mining area. Jian Wang 0138, Li Yan 0003, Keming Yang, Wei Tang 0008, Hong Xie 0002, Shuyi Yao, Zhihua Xu, Jianbing Yang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Class-Specific Dictionary Based Semi-Supervised Domain Adaptation for Land-Cover Classification of Aerial ImagesabstractSupervised scene classification of aerial images plays an important role in land-cover classification. However, it is difficult and time-consuming to annotate the required number of samples. Moreover, the conventional classifiers cannot produce satisfactory results without sufficient labelled data. Semi-supervised domain adaptation methods can overcome this problem to some extent by transferring previously labelled data. But the feature distribution bias caused by different sensors, seasons or locations may lead to a lower performance. In order to reduce the feature distribution bias and keep the discriminative ability, we propose a novel class-specific dictionary-based semi-supervised domain adaptation (CDSDA) framework when newly labelled data are unavailable. The CDSDA first learns a discriminative class-specific dictionary in the source domain. Then the target dictionary are obtained by designing an objective function in an iterative process in order to reduce the feature distribution bias and keep the discriminative ability. The target and source features are both mapped to the new feature space by the learnt target dictionary for classification. The CDSDA method was tested on a large aerial image where two benchmark datasets serve as the training dataset. The experiments on a larger aerial image demonstrate that the CDSDA method performs better than some previous domain adaptation methods in the case of no target labelled data. Li Yan 0003, Ruixi Zhu, Yi Liu 0028, Nan Mo |
IGARSS | 1 |
| 2019 | Cross-Domain Distance Metric Learning Framework With Limited Target Samples for Scene Classification of Aerial ImagesabstractIn this paper, we concentrate on the problem of cross-domain aerial scene classification. The primary assumption of the proposed cross-domain distance metric learning (CDDML) framework is that training data are adequate in the source domain but limited in the target domain. One major problem of cross-domain scene classification caused by different dates, sensor positions, lighting conditions, and sensor types is data distribution bias. To solve this problem, the CDDML framework first replaces the existing color space with the proposed hybrid color features derived from all candidate color components to decrease the spectral shift between domains. Then, hybrid color features and bag of convolution features (BOCFs) are put into a discriminating DML (DDML) method to reduce the data distribution bias in the feature space. Finally, the image-to-subcategory distance measure is proposed to decrease the effect of intraclass variability on the nearest neighbor classifier by fusing hybrid color features and BOCF in the distance space. The experiments on three aerial target images or data sets confirm that the CDDML framework can obtain better results than most of the previous methods in the case of inadequate samples. Experimental results also demonstrate that DDML, hybrid color features, and the image-to-subcategory distance measure can increase the classification performance. Li Yan 0003, Ruixi Zhu, Nan Mo, Yi Liu 0028 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Scene Capture and Selected Codebook-Based Refined Fuzzy Classification of Large High-Resolution ImagesabstractScene classification has been successfully applied to the semantic interpretation of large high-resolution images (HRIs). The bag-of-words (BOW) model has been proven to be effective but inadequate for HRIs because of the complex arrangement of the ground objects and the multiple types of land cover. How to define the scenes in HRIs is still a problem for scene classification. The previous methods involve selecting the scenes manually or with a fixed spatial distribution, leading to scenes with a mixture of objects from different categories. In this paper, to address these issues, a scene capture method using adjacent segmented images and a support vector machine classifier is proposed to generate scenes dominated by one category. The codebook in BOW is obtained from clustering features extracted from all the categories, which may lose the discrimination in some vocabularies. Thus, more discriminative visual vocabularies are selected by the introduced mutual information and the proposed intraclass variability balance in each category, to decrease the redundancy of the codebook. In addition, a refined fuzzy classification strategy is presented to avoid misclassification in similar categories. The experimental results obtained with three different types of HRI data sets confirm that the proposed method obtains classification results better than those obtained by most of the previous methods in all the large HRIs, demonstrating that the selection of representative vocabularies, the refined fuzzy classification, and the scene capture strategy are all effective in improving the performance of scene classification. Li Yan 0003, Ruixi Zhu, Yi Liu 0028, Nan Mo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | OSSIM: An Object-Based Multiview Stereo Algorithm Using SSIM Index Matching CostabstractMultiview stereo (MVS) is a crucial process in image-based automatic 3-D reconstruction and mapping applications. In a dense matching process, the matching cost is generally computed between image pairs, making the efficiency low due to the large number of stereo pairs. This paper presents a novel object-based MVS algorithm using structural similarity (SSIM) index matching cost in a coarse-to-fine workflow. As far as we know, this is the first time SSIM index is introduced to calculate the matching cost of MVS applications. In contrast to classical stereo methods, the proposed object-based structural similarity (OSSIM) method computes only a depth map for each image. Thus, the efficiency can be greatly improved when the overlap between images is large. To obtain an optimized depth map, the winner-take-all and semi-global matching strategies are implemented. Moreover, an object-based multiview consistency checking strategy is also proposed to eliminate wrong matches and perform pixelwise view selection. The proposed method was successfully applied on a close-range Fountain-P11 data set provided by EPFL and aerial data sets of Vaihingen and Zürich by the ISPRS. Experimental results demonstrate that the proposed method can deliver matches at high completeness and accuracy. For the Vaihingen data set, the correctness and completeness rate were 71.12% and 95.99% with an RMSE of 2.8 GSD. For the Foutain-P11 data set, the proposed method outperformed the other existing methods with the ratio of pixels less than 2 cm. Extensive comparison using Zürich data set shows that it can derive results comparable to the state-of-the-art software (PhotoScan, Pix4d, and Smart3D) in urban buildings areas. Liang Fei, Li Yan 0003, Changhai Chen, Zhiyun Ye, Jiantong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | An Adaptive Weighted Tensor Completion Method for the Recovery of Remote Sensing Images With Missing DataabstractMissing information, such as dead pixel values and cloud effects, is very common image quality degradation problems in remote sensing. Missing information can reduce the accuracy of the subsequent image processing, in applications such as classification, unmixing, and target detection, and even the quantitative retrieval process. The main aim of this paper is to study an adaptive weighted tensor completion (AWTC) method for the recovery of remote sensing images with missing data. Our idea is to collectively make use of the spatial, spectral, and temporal information to build a new weighted tensor low-rank regularization model for recovering the missing data. In the model, the weights are determined adaptively by considering the contribution of the spatial, spectral, and temporal information in each dimension. Experimental results based on both simulated and real data sets are presented to verify that the proposed method can recover missing data, and its performance is found to be better than the other tested methods. In the simulated experiments, the peak signal-to-noise ratio is improved by more than 3 dB, compared with the original tensor completion model. In the real data experiments, the proposed AWTC model can better recover the dead line problem in Aqua Moderate Resolution Imaging Spectroradiometer band 6 and the scan-line corrector-off problem in enhanced thematic mapper plus images, with the smallest spectral distortion. Michael Kwok-Po Ng, Qiangqiang Yuan, Li Yan 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Remote sensing image super-resolution via regional spatially adaptive total variation modelabstractTotal variation has been used as a popular and effective image prior model in the regularization-based image processing fields. However, as the total variation model favors a piecewise constant solution, the processing result under high noise intensity in the flat regions of the image is often poor, and some “pseudo-edges” are produced. In this paper, we develop a regional spatially adaptive total variation (RSATV) model. Firstly, the spatial information is extracted based on each pixel, and then two filtering processes are respectively added to suppress the effect of “pseudo-edges”. After that, the spatial information weight is constructed and classified with k-means clustering, and the regularization strength in each region is controlled by the clustering center value. The experimental results, on both simulated and real datasets, show that the proposed approach can effectively reduce the “pseudo-edges” of the total variation regularization in the flat regions, and maintain the partial smoothness of the highresolution image. More importantly, compared with the traditional pixel-based spatial information adaptive approach, the proposed region-based spatial information adaptive total variation model can better avoid the effect of noise on the spatial information extraction, and maintains robustness with changes in the noise intensity in the super-resolution process. Qiangqiang Yuan, Li Yan 0003, Jiancheng Li, Liangpei Zhang 0001 |
IGARSS | 2 |