VLDB 2026 Research / reviewers in the wild / expert
Hong Xie 0002
dblp:39/3657-2
· DBLP profile ↗
16ranked-venue papers
0as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HFCH: Hybrid frontier guided fast UAV autonomous exploration for complete and high-quality mapping in unknown environment
Yuquan Zhou, Li Yan 0003, Yaxi Han, Longze Zhu, Hong Xie 0002 |
Adv. Eng. Informatics | 6 |
| 2026 | HALE: a hierarchical autonomous exploration framework for UAVs with limited field-of-view in large-span environments
Yuquan Zhou, Li Yan 0003, Longze Zhu, Yaxi Han, Yukang Liu, Hong Xie 0002 |
Expert Syst. Appl. | 7 |
| 2026 | HELF-SLAM: A hybrid-enhanced learning-based feature for distortion-resilient monocular SLAM
Longze Zhu, Li Yan 0003, Hong Xie 0002, Xiaoteng Yang, Jiang Song, Linxia Ji, Aoran Li |
Neurocomputing | 3 |
| 2026 | MAHF-LIO: Motion-Aware Hierarchical Fusion for Robust LiDAR-Inertial OdometryabstractAccurate and robust localization is critical for Internet of Things (IoT)-enabled autonomous vehicles and intelligent mobile robots operating in complex and dynamic environments. However, existing learning-based LiDAR-inertial odometry (LIO) methods often underutilize the complementarity between LiDAR and IMU measurements, as they typically rely on either simple weighted fusion strategies or computationally intensive attention-based interaction mechanisms, without explicitly modeling hierarchical cross-modal interactions. To address this limitation, we propose MAHF-LIO, a motion-aware hierarchical fusion framework that combines local motion-conditioned modulation with global lightweight feature mixing for robust and accurate LIO. Specifically, a Motion-Conditioned Gated Local Modulation (MCGLM) module uses IMU-derived temporal motion cues to adaptively recalibrate local LiDAR geometric features, improving feature reliability under varying motion conditions. In addition, a Lightweight Global MLP-Mixer Fusion (LGMMF) module performs efficient cross-modal feature integration to capture long-range dependencies between LiDAR and IMU representations with modest computational cost. Extensive experiments on both public and self-collected datasets show that MAHF-LIO delivers accurate and robust localization across diverse environments. On the KITTI benchmark, MAHF-LIO reduces the average relative translational and rotational errors by 7.1% and 29.8%, respectively, compared with Adaptive-LIO. It also demonstrates greater robustness than learning-based baselines under challenging scenarios, including rapid motion and perceptual degradation. Ablation studies further verify the effectiveness of the proposed hierarchical fusion design. Xin Yang 0039, Li Yan 0003, Hong Xie 0002, Xiaohu Lin, Longze Zhu, Aiqiang Ma |
IEEE Internet Things J. | 3 |
| 2025 | Casual3DHDR: High Dynamic Range 3D Gaussian Splatting from Casually Captured VideosabstractPhoto-realistic novel view synthesis from multi-view images, such as neural radiance field (NeRF) and 3D Gaussian Splatting (3DGS), has gained significant attention for its superior performance. However, most existing methods rely on low dynamic range (LDR) images, limiting their ability to capture detailed scenes in high-contrast environments. While some prior works address high dynamic range (HDR) scene reconstruction, they typically require multi-view sharp images with varying exposure times captured at fixed camera positions-a process that is time-consuming and impractical. To make data acquisition more flexible, we propose Casual3DHDR, a robust one-stage method that reconstructs 3D HDR scenes from casually-captured auto-exposure (AE) videos, even under severe motion blur and unknown, varying exposure times. Our approach integrates a continuous camera trajectory into a unified physical imaging model, jointly optimizing exposure times, camera poses, and the camera response function (CRF). Extensive experiments on synthetic and real-world datasets demonstrate that Casual3DHDR outperforms existing methods in robustness and rendering quality. Shucheng Gong, Lingzhe Zhao, Wenpu Li, Hong Xie 0002, Shiyu Zhao 0002, Peidong Liu 0001 |
ACM Multimedia | 4 |
| 2025 | SED-SLAM: Enhancing Monocular SLAM Under Image Distortions via Spatially Equalized Deep Feature
Longze Zhu, Li Yan 0003, Hong Xie 0002, Xiaoteng Yang, Aoran Li |
PRICAI (5) | 3 |
| 2025 | Hyperspectral Video Tracking With Spectral-Spatial Fusion and Memory EnhancementabstractHyperspectral video (HSV) provides rich spectral-spatial-temporal information, enabling the capture of complex object dynamics beyond the limitations of conventional single- and multi-modal tracking. However, current HSV tracking methods face challenges such as data scarcity, band gaps, spectral fragmentation, temporal underutilization, and high computational load, which constrain performance. In this article, we present SpectralTrack, a novel HSV tracking framework with spectral-spatial fusion and memory enhancement. SpectralTrack incorporates an explicit visual prompting module to mitigate band gaps and spectral fragmentation. We further introduce an extraction-matching-interaction module, which leverages a template-bridging search adapter and a multi-layer perceptron adapter within a multi-modal Transformer architecture for efficient cross-modal feature extraction-matching-interaction. Additionally, a memory perception module enhances state reasoning by injecting temporal prompts to refine spectral and spatial cues. SpectralTrack follows parameter-efficient fine-tuning and feature-level fusion to alleviate data scarcity and reduce computational overhead. We instantiate two variants, SpectralTrack and SpectralTrack+, across nine HSV tracking datasets, demonstrating superior effectiveness over extensive trackers. Implementations and results will be available at https://github.com/YZCU/SpectralTrack. Yuzeng Chen, Qiangqiang Yuan, Hong Xie 0002, Yi Xiao 0003, Renxiang Guan, Xinwang Liu 0002, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | MaCon: A Generic Self-Supervised Framework for Unsupervised Multimodal Change DetectionabstractChange detection(CD) is important for Earth observation, emergency response and time-series understanding. Recently, data availability in various modalities has increased rapidly, and multimodal change detection (MCD) is gaining prominence. Given the scarcity of datasets and labels for MCD, unsupervised approaches are more practical for MCD. However, previous methods typically either merely reduce the gap between multimodal data through transformation or feed the original multimodal data directly into the discriminant network for difference extraction. The former faces challenges in extracting precise difference features. The latter contains the pronounced intrinsic distinction between the original multimodal data; direct extraction and comparison of features usually introduce significant noise, thereby compromising the quality of the resultant difference image. In this article, we proposed the MaCon framework to synergistically distill the common and discrepancy representations. The MaCon framework unifies mask reconstruction (MR) and contrastive learning (CL) self-supervised paradigms, where the MR serves the purpose of transformation while CL focuses on discrimination. Moreover, we presented an optimal sampling strategy in the CL architecture, enabling the CL subnetwork to extract more distinguishable discrepancy representations. Furthermore, we developed an effective silent attention mechanism that not only enhances contrast in output representations but stabilizes the training. Experimental results on both multimodal and monomodal datasets demonstrate that the MaCon framework effectively distills the intrinsic common representations between varied modalities and manifests state-of-the-art performance across both multimodal and monomodal CD. Such findings imply that the MaCon possesses the potential to serve as a unified framework in the CD and relevant fields. Source code will be publicly available once the article is accepted. Jian Wang 0138, Li Yan 0003, Jianbing Yang, Hong Xie 0002, Qiangqiang Yuan, Pengcheng Wei, Zhao Gao, Ce Zhang 0005, Peter M. Atkinson |
IEEE Trans. Image Process. | 4 |
| 2024 | Unsupervised Multimodal Change Detection by Distilling Common and Discrepant RepresentationsabstractChange detection (CD) has become increasingly important in remote sensing and Earth observation. Currently, the data in various modalities has rapidly increased, and multimodal change detection is gaining prominence and holds substantial potential for applications demanding high temporal frequency or rapid response. In this research, we proposed a novel CDR-Net architecture for unsupervised multimodal change detection. The CDR-Net fuses the merits of the mask reconstruction and contrastive learning self-supervised paradigm. Within this architecture, the mask reconstruction subnetwcork pays more attention to low-level details, distilling common representations between multimodal remote sensing images to make them comparable, while the CL subnetwork emphasizes high-level semantics, extracting discrepant representations to facilitate the change detection task. Experimental results demonstrated that the CDR-Net achieved outstanding performance. This implies that the CDR-Net is of great value for resource investigation, emergency response and time-series understanding. Jian Wang 0138, Li Yan 0003, Hong Xie 0002, Tingyuan Zhou, Wenxu Shi, Peter M. Atkinson |
IGARSS | 3 |
| 2024 | Micro-Structures Graph-Based Point Cloud Registration for Balancing Efficiency and AccuracyabstractPoint cloud registration (PCR) is a fundamental and significant issue in photogrammetry and remote sensing, aiming to seek the optimal rigid transformation between sets of points. Achieving efficient and precise PCR poses a considerable challenge. We propose a novel micro-structures graph-based global PCR method. The overall method is comprised of two stages. 1) Coarse registration (CR): We develop a graph incorporating micro-structures, employing an efficient graph-based hierarchical strategy to remove outliers for obtaining the maximal consensus set. We propose a robust GNC-Welsch estimator for optimization derived from a robust estimator to the outlier process in the Lie algebra space, achieving fast and robust alignment. 2) Fine registration (FR): To refine local alignment further, we use the octree approach to adaptive search plane features in the micro-structures. By minimizing the distance from the point-to-plane, we can obtain a more precise local alignment, and the process will also be addressed effectively by being treated as a planar adjustment (PA) algorithm combined with Anderson accelerated (PA-AA) optimization. After extensive experiments on real data, our proposed method performs well on the 3DMatch and ETH datasets compared to the most advanced methods, achieving higher accuracy metrics and reducing the time cost by at least one-third. Rongling Zhang, Li Yan 0003, Pengcheng Wei, Hong Xie 0002, Pinzhuo Wang, Binbing Wang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | TSTD:A Cross-modal Two Stages Network with New Trans-decoder for Point Cloud Semantic Segmentation
Zhao Gao, Li Yan 0003, Hong Xie 0002, Pengcheng Wei, Jian Wang 0138 |
PRCV (8) | 3 |
| 2023 | A Voxel-Based Multiview Point Cloud Refinement Method via Factor Graph Optimization
Li Yan 0003, Hong Xie 0002, Pengcheng Wei, Jicheng Dai, Zhao Gao, Rongling Zhang |
PRCV (2) | 3 |
| 2023 | A New Outlier Removal Strategy Based on Reliability of Correspondence Graph for Fast Point Cloud RegistrationabstractRegistration is a basic yet crucial task in point cloud processing. In correspondence-based point cloud registration, matching correspondences by point feature techniques may lead to an extremely high outlier (false correspondence) ratio. Current outlier removal methods still suffer from low efficiency, accuracy, and recall rate. We use an intuitive method to describe the 6-DOF (degree of freedom) curtailment process in point cloud registration and propose an outlier removal strategy based on the reliability of the correspondence graph. The method constructs the corresponding graph according to the given correspondences and designs the concept of the reliability degree of the graph node for optimal candidate selection and the reliability degree of the graph edge to obtain the global maximum consensus set. The presented method achieves fast and accurate outliers removal along with gradual aligning parameters estimation. Extensive experiments on simulations and challenging real-world datasets demonstrate that the proposed method can still perform effective point cloud registration even the correspondence outlier ratio is over 99%, and the efficiency is better than the state-of-the-art. Code is available at https://github.com/WPC-WHU/GROR. Li Yan 0003, Pengcheng Wei, Hong Xie 0002, Jicheng Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | UBMDP: Urban Building Mesh Decoupling and PolygonizationabstractWith the development of photogrammetry, digital city, and metaverse, the 3-D representation of urban buildings has attracted more and more attention. As the main form of the 3-D urban building model, the triangular mesh model has deficiencies such as high complexity, high-data volume, and low-structural information, which seriously restrict its application in spatial analysis and urban planning. This article proposes a hybrid modeling strategy geared toward the mesh model generated from oblique images to obtain building models that are compact, manifold, watertight, and have certain structural and semantic information. First of all, when the planar region topology graph has been established, a topology decoupling strategy is designed to obtain a set of relatively independent topology subgraphs which form a hierarchical structure. After that, to improve model quality, topology optimization of parallel planes has also been studied systematically. Then, we adopt a divide-and-conquer strategy to perform data-driven and model-driven building modeling for the primary and ancillary structures. Finally, a component-level simple polygon model combination is generated. Experiments prove that the proposed method has excellent visual authenticity, structural completeness advantages, and decent LoD3 ability. As a mesh simplification method, the data is compressed to 0.11%–0.75% in a Hausdorff metric around 0.3 m, which further proves that this method is state-of-the-art. Li Yan 0003, Yao Li 0030, Jicheng Dai, Hong Xie 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A DSM-Based Co-Occurrence Matrix for Semantic ClassificationabstractTraditional 2D textures cannot reflect objects’ real textures in 3D world, since they only consider spectral distribution in a 2D region which is a projection of 3D objects at a certain angle of view. The existing researches of 3D textures can only process volumetric data (VD) like multi/hyper-spectral images which are not real 3D geometric data. In this letter, we proposed a digital surface model (DSM)-based co-occurrence matrix (DSMB-CM) which extended 2D co-occurrence matrix (2D-CM) to 3D space for multispectral images with DSM. DSMB-CM is the first 3D feature in remote sensing areas considering the spectral distribution over 3D surface to represent real textures of objects in 3D space. Besides, a dimension reduction method was proposed to avoid curse of dimensionality. Experiments compared classification accuracies of different feature combinations of two data sets from ISPRS Benchmark of Semantic Labeling Contest. The results proved that DSMB-CM had better performance than traditional textures in identification of all categories. Li Yan 0003, Hong Xie 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Deriving Mining-Induced 3-D Deformations at Any Moment and Assessing Building Damage by Integrating Single InSAR Interferogram and Gompertz Probability Integral Model (SII-GPIM)abstractIt is necessary to timely and accurately estimate the surface deformations in mining areas, especially the three-dimensional (3D) deformations during surface movement. At present, nearly all mining-induced 3D deformations retrieved by interferometric synthetic aperture radar (InSAR) pertain to the SAR imaging interval. Research on progressive 3D deformations during surface movement is limited, and the existing approaches are unsatisfactory in practical engineering. Aiming at these challenges, we proposed a novel method for deriving mining-induced 3D surface deformations at any moment by integrating single InSAR interferogram (SII), the Gompertz time function, and the probability integral model (PIM), named the SII-GPIM method. We established an inversion approach for GPIM parameters and derived the mining-induced 3D surface deformations at any moment. Subsequently, we conducted experiments considering two ALOS PALSAR images in the Huaibei mining area. The accuracy of the proposed method was evaluated in subsidence, tilt, curvature, horizontal displacement, and horizontal strain. Compared with existing methods, the SII-GPIM method is state of the art. Additionally, we assessed the building damage, performance of parameter inversion, and method generality. The results demonstrated that the proposed method can accurately determine the mining-induced 3D surface deformations and deformation level at any moment under different geological mining conditions. Moreover, accurate GPIM parameters can be acquired with only two SAR images and traditional measurement is nearly not required. Consequently, the SII-GPIM method owns great value for improving economic efficiency, assessing building damage, and restoring the ecological environment in the mining area. Jian Wang 0138, Li Yan 0003, Keming Yang, Wei Tang 0008, Hong Xie 0002, Shuyi Yao, Zhihua Xu, Jianbing Yang |
IEEE Trans. Geosci. Remote. Sens. | 5 |