VLDB 2026 Research / reviewers in the wild / expert
Guangbiao Wang
dblp:365/5900
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0007-6634-4534ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised Domain Adaptation for VHR Urban Scene Segmentation via Prompted Foundation Model-Based Hybrid Training Joint-Optimized NetworkabstractUnsupervised Domain Adaptation for Remote Sensing Semantic Segmentation (UDA-RSSeg) is to adapt a model trained on the source domain data to the target domain samples, thereby minimizing the need for annotated data across diverse remote sensing scenes. In urban planning and monitoring, the task of UDA-RSSeg on Very-High-Resolution (VHR) images has garnered significant research interest. While recent deep learning techniques have demonstrated huge success in tackling the UDA-RSSeg task for VHR urban scenes, a persistent challenge in addressing the domain shift issue remains. Specifically, there are two primary problems: (1) severe inconsistencies in feature representation across diverse domains, characterized by notably differing data distributions, and (2) the domain gap problem due to the representation bias of the source domain patterns when translating features to predictive logits. To solve these problems, we propose a prompted foundation model based hybrid training joint-optimized network (PFM-JONet) for UDA-RSSeg on VHR urban scene. Our approach integrates the notable “Segment Anything Model” (SAM) as prompted foundation model to leverage its robust generalized representation capabilities, thereby alleviating feature inconsistencies. Based on the feature extracted by SAM-Encoder, we introduce a mapping decoder designed to convert SAM-Encoder features into predictive logits. Additionally, a prompted segmentor is employed to generate class-agnostic maps, which guide the mapping decoder’s feature representations. To efficiently optimize the entire network in an end-to-end manner, we design a hybrid training scheme that integrates feature-level and logits-level adversarial training strategies alongside a self-training mechanism. This scheme enhances the model from diverse, compatible perspectives. To evaluate the performance of our proposed PFM-JONet, we conduct extensive experiments on urban scene benchmark datasets, including ISPRS (Potsdam/Vaihingen) and CITY-OSM (Paris/Chicago). On ISPRS dataset, PFM-JONet surpasses previous SOTA methods by 1.60% in mean IoU value across four adaptation tasks. For CITY-OSM’s adaptation task, it outperforms SOTA by 4.84% in mean IoU value. These results demonstrate the effectiveness of our method. Furthermore, visualization and analysis reinforce the method’s interpretability. The code of this paper is available at https://github.com/CV-ShuchangLyu/PFM-JONet. Shuchang Lyu, Qi Zhao 0037, Yaxuan Sun, Yiwei He, Guangbiao Wang, Jinchang Ren, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Learn From Past to Future: Exploiting Self-Training and Curriculum Learning in Remote Sensing Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation focuses on updating the segmentation model with only new-class samples. Catastrophic forgetting and background shift are the two prevalent challenges. We identify two additional issues in remote sensing data that worsen these problems: significant class distribution variability and error accumulation-induced model degradation. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). First, we introduce a self-training auxiliary branch to complement the frozen last-step model, integrating cross-step knowledge to mitigate rapid forgetting. Then, a gradient-oriented Dynamic Refined Loss is proposed to assess under-learned classes and mitigate class imbalance. Furthermore, class-balanced curriculum learning is embedded to alleviate performance degradation throughout incremental training. Extensive experiments on benchmark datasets, including DeepGlobe, iSAID, ISPRS Potsdam, and Vaihingen, demonstrate that the proposed STCL-DRNet achieves state-of-the-art (SOTA) performance. In the 1-1s setting of the DeepGlobe dataset, STCL-DRNet exceeds previous SOTA methods by 11.6% in mIoU. For the iSAID 10-1s setting, it outperforms the previous SOTA by 12.76% in mIoU. As for ISPRS Potsdam and Vaihingen, our STCL-DRNet surpasses the SOTA by 5%-8% in all settings. Visualization and analysis further validate its interpretability. Our code is available at https://github.com/cv516Buaa/STCL-DRNet. Ruimin Ren, Hongbo Zhao 0001, Shuchang Lyu, Guangbiao Wang, Qi Zhao 0037, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | High-Precision Wave-Parameter Perception via Spatiotemporal Coupling of Sea-Clutter Imagery and Ship-Motion ResponsesabstractAccurate detection of near-wave field parameters is crucial for safe navigation and efficient offshore operations. However, estimating precise wave parameter remains challenging due to strong nonlinear wave dynamics and inherent measurement uncertainties. Most of existing deep learning methods relying on single modality data face limitations in insufficient accuracy. Motivated by advances in multimodal fusion algorithms, we propose a novel maritime multimodal fusion inversion model, MR-FuNet. The proposed model integrates a Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) in parallel, enabling effective fusion of spatial information from X-band radar sea clutter images and temporal patterns from ship motion data at the feature level. A multi-dimensional attention strategy was employed to enable the model to dynamically calibrate and integrate heterogeneous information across modalities and feature domains which can significantly improve the accuracy of inversion for significant wave height and characteristic wave period. To address the lack of comprehensive and high-quality public datasets in this research area, this study constructs and releases a large-scale and multimodal datasetRadar Images and Ship Motion Dataset(RSD). RSD is a large-scale multimodal dataset generated via numerical simulation and covers 99 representative sea states. It provides a valuable benchmark for future research. Extensive experiments validate the effectiveness of the proposed model, demonstrating substantial performance gains across various metrics. Compared to traditional Artificial Neural Networks (ANN), the proposed multimodal fusion model achieves notable reductions in RMSE by 61.6% and 62.8% for the characteristic period and significant wave height inversion tasks, respectively. Dataset is available at https://github.com/felixfelixXu/Radar-Images-and-Ship-Motion-Dataset. Guangbiao Wang, Zihang Xu, Limin Huang, Shuchang Lyu, Jingjun Li, Linzhou Tang, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | WEA-DINO: An Improved DINO With Word Embedding Alignment for Remote Scene Zero-Shot Object DetectionabstractRemote sensing scene zero-shot object detection aims to detect and recognize both seen and unseen catrgories of landscape elements with the guidance of the word embeddings. In this task, two primary challenges are identified. Firstly, there exists considerable variability within categories of landscape elements, causing a misalignment between visual features and word embeddings, particularly noticeable for unseen categories. Secondly, existing detection models struggle to provide accurate localization predictions, greatly impacting overall performance. To address these two issues, we propose WEA-DINO (Word Embedding Alignment-DINO). Based on the original DINO structure, our WEA-DINO-Head is specifically designed to align the hidden features of “matching queries” with word embedding features, effectively addressing the misalignment issue between visual features and word embeddings. Furthermore, aligning the hidden features of “denoising queries” with word embedding features enables the translation of localization capabilities from known categories to previously unseen ones. Through extensive experimentation on the DIOR benchmark dataset, our method demonstrates state-of-the-art performance. The code is available at https://github.com/cv516Buaa/WEA-DINO. Guangbiao Wang, Hongbo Zhao 0001, Qing Chang 0003, Shuchang Lyu, Huojin Chen |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | SWIN-TOD: Smooth Wasserstein Distance and Instance-Level Neighboring Enhancement for Remote Sensing Tiny Object DetectionabstractThe advancement of deep neural network has propelled the widespread application of remote sensing target detection. However, compared to natural scenes, remote sensing targets possess inherent characteristics such as weak features and small scale, leading to a significant performance gap in traditional detection methods. To address these challenges, we undertake a systematic analysis of existing approaches, focusing on two key aspects: inadequate extraction of discriminative features and inappropriate regression measurement metrics. To tackle the first issue, an instance-level neighboring enhancement network (INEN) is proposed, enhancing the network’s feature extraction capability through inter-object feature aggregation. To address the second issue, a novel metric, smooth Wasserstein loss (SWL), is devised. Building upon these principles, a new tiny object detection (TOD) network for remote sensing images is developed. Extensive experiments on AI-TOD v1/v2 and DOTA v2 remote sensing tiny target detection datasets demonstrate that our approach achieves state-of-the-art (SOTA) performance. Codes are available athttps://github.com/sevenwgb/SWIN-TOD. Guangbiao Wang, Hongbo Zhao 0001, Shuchang Lyu, Qing Chang 0003, Wenquan Feng, Qi Zhao 0037, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |