Weibin Li 0002

dblp:186/4512-2 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
19since 2021 · last 2027
0000-0003-0047-8955ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2027 RobustTableEval: How visual degradation affects image-based table recognition capabilities of multimodal large language models
Weibin Li 0002, Lina Hou, Yutong Guo, Haijun Xue
Expert Syst. Appl.2
2026 Like Human Rethinking: Contour Transformer AutoRegression for Referring Remote Sensing Interpretation
abstract
Referring remote sensing interpretation holds significant application value in various scenarios such as ecological protection, resource exploration, and emergency management. However, referring remote sensing expression comprehension and segmentation (RRSECS) faces critical challenges, including micro-target localization drift problem caused by insufficient extraction of boundary features in existing paradigms. Moreover, when transferred to remote sensing domains, polygon-based methods encounter issues such as contour-boundary misalignment and multi-task co-optimization conflicts problems. In this paper, we propose SeeFormer, a novel contour autoregressive paradigm specifically designed for RRSECS, which accurately locates and segments micro, irregular targets in remote sensing imagery. We first introduce a brain-inspired feature refocus learning (BIFRL) module that progressively attends to effective object features via a coarse-to-fine scheme, significantly boosting small-object localization and segmentation. Next, we present a language-contour enhancer (LCE) that injects shape-aware contour priors, and a corner-based contour sampler (CBCS) to improve mask-polygon reconstruction fidelity. Finally, we develop an autoregressive dual-decoder paradigm (ARDDP) that preserves sequence consistency while alleviating multi-task optimization conflicts. Extensive experiments on RefDIOR, RRSISD, and OPTRSVG datasets under varying scenarios, scales, and task paradigms demonstrate transformative performance gains: compared to the baseline PolyFormer, our proposed SeeFormer improves oIoU and mIoU by 27.58% and 39.37% for referring image segmentation and by 18.94% and 28.90% for visual grounding on the RefDIOR dataset.
Jinming Chai, Licheng Jiao, Xiaoqiang Lu, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Weibin Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.9
2026 DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
abstract
Although significant advances have been achieved in SAR land-cover classification, recent methods remain predominantly focused on supervised learning, which relies heavily on extensive labeled datasets. This dependency not only limits scalability and generalization but also restricts adaptability to diverse application scenarios. In this paper, a general-purpose foundation model for SAR land-cover classification is developed, serving as a robust cornerstone to accelerate the development and deployment of various downstream models. Specifically, a Dynamic Instance and Contour Consistency Contrastive Learning (DI3CL) pre-training framework is presented, which incorporates a Dynamic Instance (DI) module and a Contour Consistency (CC) module. DI module enhances global contextual awareness by enforcing local consistency across different views of the same region. CC module leverages shallow feature maps to guide the model to focus on the geometric contours of SAR land-cover objects, thereby improving structural discrimination. Additionally, to enhance robustness and generalization during pre-training, a large-scale and diverse dataset named SARSense, comprising 460,532 SAR images, is constructed to enable the model to capture comprehensive and representative features. To evaluate the generalization capability of our foundation model, we conducted extensive experiments across a variety of SAR land-cover classification tasks, including SAR land-cover mapping, water detection, and road extraction. The results consistently demonstrate that the proposed DI3CL outperforms existing methods. Our code and pre-trained weights are publicly available at: https://github.com/SARpre-train/DI3CL.
Zhongle Ren, Kai Wang 0053, Biao Hou, Xingyu Luo, Weibin Li 0002, Licheng Jiao
IEEE Trans. Image Process.6
2026 Scale-Aware Prompting With Optimal Transport for Remote Sensing Image Captioning
abstract
Remote sensing image captioning is a multimodal foundation task for fine-grained understanding of remote sensing images. However, remote sensing images contain complex scenes and rich objects, it is very challenging to accurately describe the objects in the scene with their attributes and dependencies. To address these issues, the article proposes a novel scale-aware prompting with optimal transport (SPOT) to learn effective multiscale features under diverse scenes, and to build fine-grained cross-modal alignment between semantic features and linguistic words during caption generation. Specifically, a scale-aware prompt extractor is constructed to explore feature integrations in complex scenes through learning prompts that query multi-scale features, and to enhance the representation of attributes and dependencies for objects by embedding positional relations. Besides, a fine-grained cross-modal alignment is designed to dynamically match image feature representations and textual semantics through optimal transport. Through the above manner, the model can learn effective language-aligned feature representations for caption generation. Finally, a caption Transformer with causal self-attention is introduced to generate accurate captions for remote sensing scenes. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on three public datasets, with the superiority of the proposed method further demonstrated by ablating the role of each component.
Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jiawei Ning, Kai Wang 0053, Weibin Li 0002, Licheng Jiao
IEEE Trans. Image Process.6
2025 TableGPT: a novel table understanding method based on table recognition and large language model collaborative enhancement
Weibin Li 0002, Wei Li 0318, Wenbo Ji
Appl. Intell.3
2025 Spatiotemporal proximity analysis of heterogeneous and heterochronous time-geographical entities
abstract
Time geography is an elegant framework for analyzing human activity–travel behaviors across temporal and spatial dimensions. At the core of this framework are the concepts of space–time paths, which represent historical trajectories, and space–time prisms, which define potential future activity spaces. However, most studies examine space–time paths and prisms in isolation, overlooking the integrative nature of time geography as a theoretical framework that incorporates the past, present, and future within a continuous temporal dimension. This study addresses this gap by developing novel methods for measuring and querying the spatiotemporal proximity of these heterogeneous and heterochronous time-geographic entities. To operationalize these methods, a GIS tool was implemented to support the local density analysis of time-geographic entities. Comprehensive computational experiments were conducted to validate the developed methods using large-scale, network-constrained, and time-geographic datasets. The developed methods exhibited high computational efficiency (approximately 2 s) in computing the local density for each path within extensive prism collections. The experimental results demonstrate the effectiveness of the developed methods in analyzing and visualizing the local density of historical paths in the near future. The methods offer new insights into space–time interactions, advancing research in human mobility forecasting and spatial decision-making.
Yu Bo Luo, Bi Yu Chen, Weibin Li 0002, Junli Liu, Xuefei Wu, Qingquan Li 0001
Int. J. Geogr. Inf. Sci.3
2025 Space-time tree: a spatiotemporal construct for efficient similarity matrix calculations among network-constrained trajectories
abstract
Data mining of network-constrained trajectories has broad applications in the GIScience field. The calculation of a complete trajectory similarity matrix is a key step in various data mining algorithms. However, computing this matrix is computationally intensive for large datasets, as it involves numerous point-to-point shortest-path (PPSP) queries. To tackle this issue, we propose a new spatiotemporal construct called the space-time tree, which directly delineates the network distance from a query trajectory to any network space-time point. By constructing the space-time tree, we can efficiently compute the trajectory similarity matrix without additional PPSP queries. The space-time tree supports several similarity metrics, including closest pair distance, furthest pair distance, longest common subsequence (LCSS), and distance-weighted LCSS. It can further integrate with advanced spatiotemporal query techniques for scalable partial trajectory similarity matrix calculations. A case study using real datasets was conducted to apply the space-time tree in the trajectory clustering application. The results show that the space-time tree completed the clustering task on 0.5 million trajectories within 49 minutes, achieving a nearly 147-fold speedup compared to state-of-the-art methods.
Yu Bo Luo, Bi Yu Chen, Yu Zhang 0019, Weibin Li 0002, Jianya Gong, Qingquan Li 0001
Int. J. Geogr. Inf. Sci.4
2025 Hierarchical Prototype Learning With Uncertainty-Aware Adaptation for Cross-Domain Semantic Segmentation of Remote Sensing Images
abstract
Prominent domain discrepancies in Remote Sensing Images (RSIs), such as sensor types, geographical patterns, and land usage, significantly hinder the research and practical applications of cross-scene land classification. Unsupervised Domain Adaptation (UDA) fully exploits domain invariance between labelled source and unlabelled target domain, which alleviates the challenge of inaccurate land classification due to lack of labels in RSIs. However, most existing UDA methods for RSIs semantic segmentation are insufficient in exploring cross-domain features and have difficulty in modelling fine-grained domain-invariant features between inter-class. In this paper, we propose a novel self-training UDA method named Hierarchical Prototype Learning (HPL), which learns the inherent domain-wise consistency and class-wise invariance through progressive exploration of the prototypy, significantly alleviating the ambiguity in uncertain regions. HPL mainly consists of Domain-wise Progressive Prototype Interaction (DPPI) and Class-wise Dynamic Prototype Collaboration (CDPC). DPPI and CDPC specialize in hierarchically building prototype interaction architectures tailored to domain-wise alignment and class-wise calibration, respectively. This design not only mitigates the sensitivity to cross-domain scenarios but also allows for precise correction of uncertain regions. Furthermore, CDPC exhibits the capacity for pixel-level category restoration and promotes the correct and fine-grained updating of pseudo-labels. Extensive Experiments on two public datasets and a private self-build dataset demonstrate the superiority of HPL over other state-of-the-art methods for UDA semantic segmentation of RSIs.
Jiawei Ning, Zhongle Ren, Biao Hou, Runnong Jiang, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2025 Deep Geospatio-Semantic Guided Network With Pseudo-Label Consistency for Domain-Adaptive Remote Sensing Segmentation
abstract
Domain-adaptive Remote Sensing Images (RSIs) semantic segmentation mitigates the overfitting problem that affects the effectiveness of segmentation, which results from the scarcity of high-quality labels and the cross-domain styles of ground objects. The effectiveness of domain adaptive segmentation remains suboptimal in complex scenarios due to inadequate exploitation of latent geographic knowledge. Consequently, inter-class ambiguity and boundary agnostic are further exacerbated under cross-domain transfer scenarios. To address this issue, we first devise a deep geospatio-semantic guided network named DSSAL, which comprehensively investigates the potential spatial relationship and semantic correlation between classes of RSIs by geospatial aware interaction and geosemantic aware interaction, respectively. To mitigate class-wise cognitive deviation in the unlabeled domain, DSSAL-DA is developed to further enhance the segmentation effect with the spatio-semantic domain alignment module in manifold cross domain tasks. Furthermore, a pseudo-labels consistency filter is developed for DSSAL-DA to ensure reliability in self-training through cross-view consistency verification. Extensive experiments on two public datasets and a private dataset demonstrate the superiority of DSSAL and DSSAL-DA over the state-of-the-art methods for UDA semantic segmentation of RSIs.
Jiawei Ning, Zhongle Ren, Biao Hou, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 Self-Supervised Learning of Contrast-Diffusion Models for Land Cover Classification in SAR Images
abstract
Deep learning methods has been widely applied to synthetic aperture radar (SAR) land cover classification. The complexity of SAR data and the limited availability of labeled samples greatly constrain the feature learning and the generalization ability of the model. Inspired by the excellent generative performance of Denoising Diffusion Probabilistic Models(DDPM) on complex data distributions, a self-supervised learning framework based on contrast-diffusion models (CDM) is proposed to expand the applicability to multiple broad scenarios with complex and varying imaging conditions under limited annotated data conditions. Specially, The proposed framework consists of the upstream CDM pre-training on all unlabeled samples and the downstream land cover classification with few labeled samples in each test scene. Concretely, in the upstream task, the features are captured through the generative learning of the DDPM. Following this, the Dimensionality Reduction and Resolution Expansion (DRRE) module is designed and embedded to reduce feature redundancy and align the feature granularity between layers and the input image. Finally, contrastive learning is employed to enforce semantic feature consistency across different steps. In the downstream task, the feature in pre-trained CDM is efficiently delivered in a single-step reverse diffusion process and then fine-tuned with few labeled samples from each test scene and finally output the predictions. Compared with several supervised and self-supervised methods, the proposed framework achieves superior classification and generalization performance on multiple broad scenes with complex and varying imaging conditions. For example, based on the average results from six test scenes, CDM shows improvements in overall accuracy (OA) of 41.42%, 34.40%, 7.70% ,8.53% and 9.60% compared to Deeplabv3+, CCNR, Segformer, MAE and DDPM respectively. The code is available at https://github.com/gosling123456/CDM.git.
Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 Interactive Concept Network Enhanced Transformer for Remote Sensing Image Captioning
abstract
Remote sensing image captioning plays an important role in advancing remote sensing image understanding with natural language generation. However, it is difficult to generate accurate semantic descriptions of crucial objects and their relationships, due to large coverage and abundant information in remote sensing images. To address these issues, this article proposes a novel interactive concept network enhanced transformer (ICNET) for remote sensing image captioning. First, multilevel visual features are extracted within a local and global feature extraction module. To comprehensively capture key objects in the local features, a concept mapping network (CMN) is constructed to project multiscale local features onto high-level semantic concepts of the objects. This allows for the integration of the relevant feature vectors in the visual feature mapping into multiple relatively independent word features, thus bridging the gap between visual features and semantic concepts. Subsequently, a global feature enhancement (GFE) module is introduced to boost the discrimination of global relationships and filter irrelevant content. Finally, to aggregate semantic concepts and global features, a transformer equipped with a concept interaction module (CIM) is designed to facilitate feature alignment and generate captions with proper categories and relationships. The experimental results on three remote sensing image captioning datasets demonstrate the superiority of the proposed method.
Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jianhua Meng, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2025 Adaptive Scale-Aware Semantic Memory Network for Remote Sensing Image Captioning
abstract
Remote sensing image captioning between visual images and natural language remains a long-standing challenge in the remote sensing community. Due to the wide coverage and large amount of information in remote sensing images, existing methods struggle to effectively utilize the relevant semantic information about objects and their attributes at different scales across samples to generate descriptions. To address these issues, the article proposes a novel Adaptive Scale-aware Semantic Memory Network (ASSMN) for remote sensing image captioning. First, to fully extract the semantic information in remote sensing images, multilevel feature enhancement is constructed to improve the feature representation extracted from the CLIP pre-training model. Subsequently, a scale-aware attention aggregator is introduced to further integrate the enhanced multi-scale image features into the high-level semantics of remote sensing images. Then, to fully exploit the semantic information of the joint observed samples, a semantic memory reinforcement is designed to strengthen the semantic representation of the current scene through the relevant semantics obtained from other training samples. Finally, a captioning decoder is performed to generate a comprehensive scene caption with accurate objects and attributes. In the experiments, the performance of the proposed ASSMN on three remote sensing image captioning datasets is evaluated and compared with well-designed baselines and state-of-the-art methods, and the superiority of the proposed method is demonstrated by ablating the role of each proposed component. The code will be available at https://github.com/zcsisiyao/ASSMN.
Cheng Zhang 0028, Zhongle Ren, Biao Hou, Changhui Xu, Jianhua Meng, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 A survey for table recognition based on deep learning
Weibin Li 0002, Wei Li 0318, Ruochen Liu 0006, Biao Hou, Licheng Jiao
Neurocomputing2
2024 Single Domain Generalization Method for Remote Sensing Image Segmentation via Category Consistency on Domain Randomization
abstract
Single Domain Generalization (SDG) is a more realistic setting than Domain Generalization (DG) and Domain Adaptation (DA). It aims to train a domain-agnostic model in the presence of a single source domain to perform well on arbitrary unseen target domains. To facilitate the practical deployment of remote sensing image segmentation in the real world, we propose a novel SDG method, termed Category Consistency on Domain Randomization (CCDR). To expand the coverage of a single source domain, CCDR implements a simple yet effective data generation module to perform domain randomization of texture and style information, since texture divergences from different geographical environments or phenological periods, and style divergences from different illumination or weather conditions, are the main causes of the domain shift in remote sensing. And in addition to emphasizing inter- and intra-class relationships within each domain via multi-domain supervised learning, CCDR further draws inspiration from the triple loss to enhance the semantic correlation across the source domain and the generated auxiliary domain, which makes the segmentation model more sensitive to class discriminative information and better adapted to unseen target domains. In comparison to other state-of-the-art SDG methods, CCDR can synthesize the more effective auxiliary domain, and perform more reliable classification on the unseen target domains without the sophisticated training pipeline and cumbersome data generation process. And the experimental results on two public remote sensing datasets demonstrate that CCDR has remarkable advantages in remote sensing image segmentation tasks. And the code is publicly available at: https://github.com/LCB1970/CCDR.
Chenbin Liang, Weibin Li 0002, Yunyun Dong, Wenlin Fu
IEEE Trans. Geosci. Remote. Sens.2
2024 Self-Supervised Learning Guided by SAR Image Factors for Terrain Classification
abstract
Effective feature representation is the key to SAR image terrain classification. Limited by the abstract appearance and the scarcity of high-quality labeled data in this field, the features learned by current methods, especially deep learning models, do not have enough directivity and applicability, which hampers the performance. This paper proposes Multi-image Factor Self-Supervised Learning(MFSSL) to achieve directional feature learning and obtain generalized features with few patch-level labeled data. The framework consists of an upstream multi-factor image style transfer task and a downstream terrain classification task. In the upstream task, the goal of feature learning is first set up by multiple SAR image factors, including the observation region, the terrain category, and the imaging parameters. And then, different styles of SAR terrain images are generated and reconstructed under this goal. Through this bidirectional generative learning, the low-level external appearance of the terrain is removed, while the essential and discriminative feature representation is retained and shared across different factors. Finally, the downstream model inherits the general feature from the upstream model and implements the terrain classification task using a small amount of labeled data. Experiments conducted on three broad SAR scenes with different image factors demonstrate that the proposed framework can improve pixel-level terrain classification only with a few patch-level labeled data.
Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Hao Zhu 0009, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2023 Lightweight Landslide Detection Method Based On Depth Separable Convolution And Double Self-Attention Mechanism *
abstract
The landslide detection methods using remote sensing images are mostly based on the traditional convolutional neural network model with high depth and complexity. The paper proposes a lightweight method based on Depth Separable Convolution and Double Self-Attention Mechanism (DSC-DSAM) for detecting landslides in remote sensing images. This method aims to reduce storage space and improve detection speed while maintaining accuracy. In our model, it starts with using a lightweight convolutional neural network model. Then, the dual self-attention mechanism is applied to improve the accuracy. The proposed method is compared with other existing classification models, and it is shown to have advantages in memory space and detection speed while maintaining accuracy.
Weibin Li 0002, Yuhui Kong, Rongfang Wang, Chunlei Huo
IGARSS1
2023 Construction and Analysis of Dali Water Segmentation Dataset of SAR Images
abstract
Flood disasters last for a long time and are destructive, so it is necessary to obtain the submerged area in a timely and effective manner, which is very important for reducing disaster losses and monitoring floods. The main contribution of this paper is to construct a dataset for training and validation of deep learning algorithms for flood detection for Gaofen-3. To overcome the scarcity of SAR datasets for water segmentation, this paper constructs a refined water segmentation dataset named Dali Water Segmentation (Dali-WS) based on the Gaofen-3 satellite in China. The dataset provides abundant rural waters in Dali County, and 1776 chips were hand-labeled for further research. We also report extensive performance for the state-of-the-art segmentation algorithms. Additionally, comprehensive evaluations of state-of-the-art segmentation algorithms are presented, demonstrating the challenging nature of the dataset and its potential for driving further advancements in flood detection. The findings of this study are expected to contribute to the progress of flood detection and recognition research.
Weibin Li 0002, Rongfang Wang, Yanhua Hu
IGARSS1
2023 A Multi-Branch U-Net for Water Area Segmentation with Multi-Modality Remote Sensing Images
abstract
Water area segmentation in remote sensing images is of great importance for flood monitoring. Convolutional neural networks have been successfully applied to various computer vision tasks. Among them, a U-shaped CNN known as U-Net achieves state-of-the-art performance on various types of image segmentation, including remote sensing images. However, there are still some difficulties in the water area segmentation of remote sensing images, such as complex backgrounds, cloud shading, and rough edges. In this work, we propose a multi-branch fusion U-Net (MFU-Net) method for water area segmentation with multi-modality remote sensing images. The experimental results showed that our MFU-Net can effectively and efficiently segment water area from Sentinel-1 and Sentinel-2 images, which F1, IoU and PA on the Sen1Floods11 dataset are 91.462%, 84.598% and 98.123%, respectively.
Rongfang Wang, Weibin Li 0002, Chunlei Huo
IGARSS4
2023 Multi-Objective Multi-Factorial Evolutionary Algorithm for Container Placement
abstract
The use of containerization technology in microservice architecture has become widespread, owing to its potential to support fast deployment of web applications and improve the resource utilization in cloud data centers. Evolutionary algorithms (EAs) have been performed promising on the deployment of applications created using microservices. However, with the growing demand for microservice application, the existing EAs fail to solve the large-scale container placement problem due to the high time complexity and poor scalability. A multi-factorial evolutionary algorithm (MFEA) is proposed in this article, which can evolve multiple optimization problems simultaneously for the container placement problem in heterogeneous cluster environments. First, a system model integrated the heterogeneous clusters, microservices, containers, and four optimization objectives is presented. Then, embedded with local search strategy, a multi-objective container placement MFEA (MOCP-MFEA) algorithm is developed to address the container placement problem. MOCP-MFEA is applied to a variety of container placement problems with different application sizes in heterogeneous cluster environments. Experimental results show that compared with various conventional and evolutionary-based approaches, MOCP-MFEA could shorten optimization time significantly and offer a competitive placement solution for the container placement problem and show a good elasticity in heterogeneous cluster environments. Moreover, the deployment scheme of container to physical machines is crucial to lowering resources wastage.
Ruochen Liu 0006, Haoyuan Lv, Weibin Li 0002
IEEE Trans. Cloud Comput.4
2019 Pansharpening with support vector transform and semi-nonnegative matrix factorization
Hong Li 0005, Weibin Li 0002, Shuying Liu
Multim. Tools Appl.2