Liangzhi Li 0002

dblp:169/4123-2 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Digital Buildings Analysis: 3-D Modeling, GIS Integration, and Visual Descriptions Using Gaussian Splatting, ChatGPT/Deepseek, and Google Maps Platform
abstract
We propose a Digital Building Analysis (DBA), a digital system for building-scale cloud-based data integration and data analytics. By connecting to cloud mapping platforms such as Google Map Platforms APIs, by leveraging state-of-the-art multi-agent Large Language Models data analysis using ChatGPT(4o) and Deepseek-V3/R1, and by using our Gaussian Splatting-based mesh extraction pipeline, our framework can retrieve a building’s 3D model, visual descriptions, and achieve cloud-based mapping integration with large language model-based data analytics using a building’s address, postal code, or geographic coordinates, and be easily extended to perform data analysis on other cloud-based data streams.
Kyle Gao, Dening Lu, Liangzhi Li 0002, Hongjie He 0003, Linlin Xu, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 RCE-Net: A Novel Multiscale Attention Network for Large-Scale Aerial 3-D Reconstruction
abstract
Recent studies demonstrate deep learning’s effectiveness for multi-view stereo (MVS) applications. However, most existing approaches focus primarily on close-range objects, with limited dedicated solutions for large-scale urban reconstruction. We present RCE-Net, a novel network specifically designed for large-area aerial imagery. Our method enhances multi-scale feature perception through an efficient attention module, addressing traditional convolutional neural networks’ limitations in complex scene feature learning. The proposed architecture incorporates both channel and spatial attention mechanisms during cost volume regularization. This dual attention approach significantly improves the network’s ability to discriminate between local and global features. Experimental results show our method achieves excellent performance across varying resolutions of the WHU dataset.
Jiaming Mo, Liangzhi Li 0002
IEEE Geosci. Remote. Sens. Lett.3
2025 SCECA-Net: A Deep Learning-Based Model for Precipitation Nowcasting
abstract
In the context of precipitation nowcasting of severe convective weather, radar echo extrapolation is a commonly employed method. However, existing methods still face numerous challenges, such as inaccurate echo boundary predictions, redundant feature extraction, and prolonged inference time, which reduce efficiency. This article proposes an innovative spatial-channel enhanced convolutional attention network (SCECA-Net) model aimed at improving feature extraction and enhancing prediction accuracy. SCECA-Net adopts a convolutional neural network (CNN) architecture and incorporates SCECA modules [spatial and channel reconstruction convolution (SCConv) and efficient channel attention (ECA)], effectively reducing spatial and channel redundancies while increasing attention to critical echo regions and enhancing the extraction of temporal sequence features. Additionally, continuous convolutions in the Dense Layer further mitigate the risk of overfitting and reduce interference between features. The experimental results demonstrate that the proposed model exhibits outstanding performance in both efficiency and accuracy.
Liangzhi Li 0002
IEEE Trans. Geosci. Remote. Sens.1
2024 A Spatiotemporal Fusion Transformer Model for Chlorophyll-a Concentrations Prediction Over Large Areas With Satellite Time Series Data
abstract
Predicting Chlorophyll-a (Chla) is essential to support the marine environment changes and marine ecosystem health, and provide early warning of algae blooms. The development of learning-based methods has facilitated Chla prediction research. Still, most of the current methods can only predict short-term Chla changes in small areas, which is limited by the ability of the model to exploit spatiotemporal dependencies. Thus, this article proposes a spatiotemporal fusion transformer prediction model (STF_Transformer) to predict relatively long-term Chla changes (15 days ahead). This model utilizes temporal and spatial transformer modules to extract the temporal and spatial correlations of the input spatiotemporal sequences, which are then fused to predict the 15-day Chla. The experimental results show that the proposed model has the optimal performance compared to the existing methods [e.g., convolutional neural network (CNN), long- and short-term memory (LSTM), and convolutional LSTM (ConvLSTM)], with root mean squared error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) less than 0.61 mg/m3, 0.315 mg/m3, and 22.5%, respectively, for 15 days Chla prediction in a large area (including Bohai, Yellow, and East China Sea). In addition, the temporal and spatial prediction results of the proposed model show that the predicted Chla has consistent temporal and spatial patterns with the observed Chla. This study indicates that the proposed STF_Transformer model can provide a highly accurate prediction of Chla over a large area in the relatively long term (15 days), providing data and technical support for marine ecosystem-related applications.
Gaoxiang Zhou, Liangzhi Li 0002
IEEE Trans. Geosci. Remote. Sens.3
2023 Multimodal Image Fusion Framework for End-to-End Remote Sensing Image Registration
abstract
We formulate the registration as a function that maps the input reference and sensed images to eight displacement parameters between prescribed matching points, as opposed to the usual techniques (feature extraction–description–matching–geometric restrictions). The projection transformation matrix (PTM) is then computed in the neural network and used to warp the sensed image, uniting all matching tasks under one framework. In this article, we offer a multimodal image fusion network with self-attention to merge the feature representation of the reference and sensed images. The integration information is then utilized to regress the prescribed points’ displacement parameters to get PTM between the reference and sensed images. Finally, PTM is supplied into the spatial transformation network (STN), which warps the sensed image to the same coordinates as the reference image, achieving end-to-end matching. In addition, a dual-supervised loss function is proposed to optimize the network from both the prescribed point displacement and the overall pixel matching perspectives. The effectiveness of our method is validated by qualitative and quantitative experimental results on multimodal remote sensing image matching tasks. The code is available at:https://github.com/liliangzhi110/E2EIR.
Liangzhi Li 0002, Mingtao Ding, Hongye Cao
IEEE Trans. Geosci. Remote. Sens.1
2023 SAR-Optical Image Matching With Semantic Position Probability Distribution
abstract
We propose a deep learning framework of Semantic Position Probability Distribution for SAR-optical image matching, termed as SPPD. Unlike the pixel-by-pixel searching matching method, a correspondence is directly obtained by an outputted matching position probability distribution. First, multiscale pyramidal features are created for each pixel in the SAR and optical images by using two weight-sharing ResNet-50 + Feature Pyramid Network (FPN) networks. The features containing high-level semantic information are then embedded into the proposed image Position Attention Module to obtain the spatial position dependencies between two images. Then, we present a loss function for semantic position matching to optimize the network from both semantic information and pixel alignment perspectives, converting the probability distribution of semantic matching positions into a point-to-point matching problem. In this paper, the SAR and optical images are set as the sensed and reference images. The effects of different image sizes, training label types, and loss function weights on matching accuracy are explored to obtain the optimal parameter settings for matching. The experimental results show that the proposed method is insensitive to image deformation and achieves cross-modal matching for SAR-optical images with high accuracy compared with the best matching method on different scene images, with several orders of magnitude faster inferences time.
Liangzhi Li 0002, Kyle Gao, Hongjie He 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Joint Self-Attention for Remote Sensing Image Matching
abstract
We propose a semantic mapping-based remote sensing image matching method, which aims to obtain the matching positions of candidate patches containing keypoints directly on the reference image, avoiding the use of cost-volume search pixel by pixel. First, a global context-fusing attention structure is created to fuse global semantic information for candidate patches with the entire sensed image. Then, a self-attention layer with semantic dependencies is proposed to extract the semantic dependencies on the reference image for cross-modal representation. The global receptive field provided by self-attention enables the proposed method to obtain the semantic mapping of candidate patches on the reference image. The experimental results show that the proposed method is insensitive to image distortion and achieves cross-modal matching of SAR-optical images with high accuracy, while still running several orders of magnitude faster. This ensures increased speed in remote sensing image analysis and pipeline processing while promoting new directions in learning-based registration.
Liangzhi Li 0002, Hongye Cao, Huijuan Hu
IEEE Geosci. Remote. Sens. Lett.1
2022 Remote Sensing Image Registration Based on Deep Learning Regression Model
abstract
We propose a novel remote sensing image registration method based on the deep learning regression network. Different from the traditional methods of feature extraction and feature matching, we pair the image blocks from sensed and reference images, and then directly learn the displacement parameters of the four corners of the sensed image block relative to the reference image. In addition, we develop the dual deep learning network with weight sharing to fully extract the registration pair image features. The proposed method is tested on different period Landsat-7 and WorldView-3 images and compared with scale-invariant feature transform (SIFT), fast and rotated brief (ORB), and other deep learning methods. The proposed method outperforms all the comparing methods.
Liangzhi Li 0002, Mingtao Ding, Hongye Cao
IEEE Geosci. Remote. Sens. Lett.1