EDBT 2026 Demo / reviewers in the wild / expert
Zhigang Yang 0002
dblp:71/1564-2
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-1049-2283ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InterMamba: A Visual-Prompted Interactive Framework for Dense Object Detection and AnnotationabstractExisting object detection methods is constrained by the high annotation costs, particularly in remote sensing due to the diversity of targets and the large scale of data. Visual-Prompted Interactive Object Detection can enhance the efficiency of data annotation by leveraging user-provided visual prompts to iteratively refine detection results. However, current interactive annotation frameworks are hindered by their reliance on simple feature fusion strategies, which limit their ability to capture fine-grained semantic relationships. Moreover, more advanced fusion methods face computational complexity challenges, making them unsuitable for high-resolution feature spaces commonly encountered in remote sensing imagery. To address these limitations, we propose InterMamba, an efficient framework for interactive object detection in remote sensing images. InterMamba integrates the VMamba backbone and a novel Cross Vision Selective Scan Module (Cross-VSSM) to achieve linear-complexity multi-scale feature fusion, reducing memory consumption while capturing fine-grained details in high-resolution feature spaces. To further enhance interaction flexibility and detection precision, a hybrid Gaussian heatmap generation method is proposed to encodes user-provided point and bounding box annotations. Meanwhile, a User Interaction Loss function further optimizes detection accuracy in dense scenarios by aligning localization and classification with user guidance. Our experiments demonstrate that InterMamba consistently outperforms existing methods in mean Average Precision (mAP). In terms of enhancing precision and reducing annotation costs, InterMamba establishes a robust solution for interactive remote sensing object detection. Code will be available at https://github.com/lsjhaha/InterMamba. Shanji Liu, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Semantic-Guided Multiview Stereo Reconstruction for Aerial ImageabstractThe application of learning-based Multi-view Stereo (MVS) depth estimation methods has achieved significant results in large-scale 3D reconstruction benchmarks. However, adjacent terrains in aerial image interfere with depth estimation along building edges during matching process, leading to inaccurate results. To address these challenges, we propose a new end-to-end MVS network, named FuS-MVSNet, which fuses monocular depth probability as a semantic guidance into the multi-view geometry-based MVS framework. By combining the strengths of geometric consistency and local semantics, FuS-MVSNet achieves notable enhancements in both accuracy and robustness. Specifically, we first construct a monocular branch based on the pre-trained Depth Anything model to perform monocular metric depth estimation. The non-shared parameters ensure that the depth estimation process is independent of multi-view branch, focusing exclusively on semantic depth inference. Subsequently, to incorporate monocular features into the multi-view network, we introduce a volume adaptive fusion module, which adaptively integrates monocular feature volumes into the standard cost volume via an attention mechanism and guides the cost volume regularization. Finally, confidence-based dynamic selection between the two prediction branches ensures the selection of the more robust branch result under challenging conditions. Qualitative and quantitative results indicate that we achieve competitive performance on multiple benchmarks, including the WHU and LuoJia-MVS datasets. Wei Zhang 0250, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Parameter-Efficient Transfer Learning for Remote Sensing Image CaptioningabstractRemote sensing image captioning (RSIC) aims to generate accurate and concise textual descriptions for remote sensing (RS) images. It plays a significant role in the analysis of earth observation data. The success of Vision-and-Language Pre-training (VLP) models provides the foundation for their transfer to the RSIC task. To reduce the cost of transferring VLP models to downstream tasks, numerous Parameter-Efficient Transfer Learning (PETL) techniques have been proposed. However, most of them focus on fine-tuning general-purpose foundation models without fully considering the unique characteristics of remote sensing data. In this paper, we introduce PE-RSIC, a novel PETL framework tailored for RSIC. Specifically, the framework builds on a pre-trained BLIP-2 model while further designing a lightweight Cross-modal RS adapter (CRS-Adapter) and a Class Prompt. During training, all parameters of the pre-trained model remain frozen, and the newly added CRS-Adapter modules are updated to efficiently transfer vision-and-language knowledge from the natural domain to the RS domain. The Class Prompt is obtained by projecting the vision-encoded [CLS] token into the decoder, guiding the model to generate more accurate captions. This approach enables the model to capture critical RS class features that might be lost during the query decoding process, with only a minimal increase in parameters. Extensive experiments show that our PE-RSIC framework outperforms full fine-tuning while utilizing only 5% of the trainable parameters. Xuezhi Zhao, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Road Extraction From Remote Sensing Images via Channel Attention and Multilayer Axial TransformerabstractRemote sensing images contain many objects that resemble road structures, making it diffcult to distinguish roads from the background. Moreover, road extraction is affected by many factors, such as lighting conditions, noise, occlusions, etc., resulting in incomplete and discontinuous road extraction. Learning discriminative road features from remote sensing images is a highly challenging task. In this paper, a novel road extraction model is proposed for remote sensing images under encoder and decoder U-Net like architecture. An axial Transformer module (ATM) is designed to learn global road features in the deepest layer with linear computational complexity regarding image size. And a multilayer attention fusion module (MLAF) is also presented to fuse multiple layers of Transformer features, obtaining more comprehensive and richer semantic information. In the skip connection, a channel attention module (CAM) is designed to weight the feature maps along the channel dimension, with the goal of improving the capability of feature representation. Extensive experiments are conducted on the DeepGlobe and Massachusetts road datasets. Compared with other methods, our proposed method in this paper realized road extraction from remote sensing images with higher accuracy and less computational cost, e.g., achieving a intersection over union (IoU) of 81.71% (1.02% improvement) and a 22.38% reduction in convergence time over the latest TransRoadNet on the Massachusetts road dataset. Ablation experiments also demonstrate the effectiveness of the designed model. Qingliang Meng, Daoxiang Zhou, Zhigang Yang 0002, Zehua Chen 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Edge-Guided Perceptual Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) plays a critical role in applications such as night navigation and fire rescue. Its primary purpose is to extract small targets from cluttered backgrounds. While deep learning-based methods have made great advancements in this field, there are still some limitations. One common issue is that the detected target shape tends to be smooth, and extremely small targets may not be effectively detected due to background interference. This article proposes an edge-guided perception network (EGPNet) for IRSTD to alleviate this trouble. To maintain the information of small targets, EGPNet utilizes a multiscale feature progressive fusion (MFPF) encoder to extract features. This progressive fusion manner enhances semantic information and contextual correlation. Considering that the detected target shapes may result in smoothing effect, an edge-guided image refinement module (EIRM) is incorporated to improve the integrity of the target shape. Moreover, we introduce a local target amplifier (LTA) to boost the visibility and representation of targets, while suppressing the clutter background interference. The experimental results illustrate that the proposed model can detect the targets with small and weak in different scenes well. Our code is publicly available athttps://github.com/qianngli/EGPNet. Qiang Li 0042, Zhigang Yang 0002, Yuan Yuan 0026, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Semantic-Spatial Collaborative Perception Network for Remote Sensing Image CaptioningabstractImage captioning is a fundamental vision-language task with wide-ranging applications in daily life. The existing methods often struggle to accurately interpret the semantic information in remote sensing images due to the complexity of backgrounds. Target region masks can effectively reflect the shape characteristics of targets and their potential interrelationships. Therefore, incorporating and fully integrating these features can significantly improve the quality of generated captions. However, researchers are hindered by the lack of relevant datasets that contain corresponding object masks. It is natural to ask the following: how to efficiently introduce and utilize object masks? In this article, we provide potential target masks for the publicly available remote sensing image caption (RSIC) datasets, enabling models to utilize the regional features of targets for RSIC. Meanwhile, a novel RSIC algorithm is proposed that combines regional positional features with fine-grained semantic information, abbreviated as$\text {S}^{2}$CPNet. To effectively capture the semantic information from image and position relationship from mask, respectively, the semantic and spatial feature enhancement submodules are introduced at the ends of encoder branches, respectively. Furthermore, the cross-view feature fusion module is designed to integrate regional features and semantic information efficiently. Then, a target recognition decoder is developed to enhance the ability of model to identify and describe critical targets in images. Finally, we improve the caption generation decoder by adaptively merging textual information with visual features to generate more accurate descriptions. Our model achieves satisfactory results on three RSIC datasets compared with the existing method. The related datasets and code will be open-sourced inhttps://github.com/CVer-Yang/SSCPNet. Qi Wang 0009, Zhigang Yang 0002, Weiping Ni, Junzheng Wu, Qiang Li 0042 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | HCNet: Hierarchical Feature Aggregation and Cross-Modal Feature Alignment for Remote Sensing Image CaptioningabstractRemote sensing image captioning aims to describe the crucial objects from remote sensing images in the form of natural language. The inefficient utilization of object texture and semantic features in images, along with the ineffective cross-modal alignment between image and text features, are the primary factors that impact the model to generate high-quality captions. To alleviate this trouble, this paper presents a network for remote sensing image captioning, namely HCNet, including hierarchical feature aggregation and cross-modal feature alignment. Specifically, a hierarchical feature aggregation module is proposed to obtain a comprehensive representation of vision features, which is beneficial for producing accurate descriptions. Considering the disparities between different modal features, we design a cross-modal feature interaction module in the decoder to facilitate feature alignment. It can fully utilize cross-modal features to localize critical objects. Besides, a cross-modal feature align loss is introduced to realize the alignment between image and text features. Extensive experiments show our HCNet can achieve satisfactory performance. Especially, we demonstrate significant performance improvements of +14.15% CIDEr score on NWPU datasets compared to existing approaches. The source code is publicly available at https://github.com/CVer-Yang/HCNet. Zhigang Yang 0002, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | C²Net: Road Extraction via Context Perception and Cross Spatial-Scale Feature InteractionabstractRoad extraction from remote sensing images (RSIs) holds significant application value in various aspects of daily scenarios. However, it is still challenging to extract high-quality road results from RSIs due to the interference of objects sharing similar structures with roads in the background and the occlusion caused by surroundings. To alleviate these problems, a road extraction network based on the global-local Context perception and Cross spatial-scale feature interaction is proposed ($\text {C}^{2}$Net). First, a global-local context perception module (GLCPM) is incorporated to capture the overall topology features of the road, which aims to improve the ability of the model to discriminate between roads and similar objects. Then, the cross spatial-scale feature interaction module is designed in the skip connection to effectively aggregate full-scale features without loss of feature information, which can provide rich and accurate road structural features for the decoder. Experiments conducted on public road datasets demonstrate that$\text {C}^{2}$Net outperforms existing methods in terms of comprehensive metrics such as intersection over union (IoU) and the$F1$-score. The results indicate that$\text {C}^{2}$Net can produce road results with superior connectivity and quality. The source code will be publicly available athttps://github.com/CVer-Yang/CCNet. Zhigang Yang 0002, Wei Zhang 0250, Qiang Li 0042, Weiping Ni, Junzheng Wu, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Road Extraction From Satellite Imagery by Road Context and Full-Stage FeatureabstractRoad extraction from satellite imagery is vital in a broad range of applications. However, extracting complete roads is challenging due to road occlusions caused by the surroundings. This letter proposed an improved encoder–decoder network via extracting road context and integrating full-stage features from satellite imagery, dubbed as RCFSNet. A multiscale context extraction (MSCE) module is designed to enhance inference capabilities by introducing adequate road context. Multiple full-stage feature fusion (FSFF) modules in the skip connection are devised to provide accurate road structure information, and we devise a coordinate dual-attention mechanism (CDAM) to strengthen the representation of road features. Extensive experiments are carried out on two public datasets, and as a result, our RCFSNet outperforms other state-of-the-art methods. The results indicate that the road labels extracted by our method have preferable connectivity. The source code will be available athttps://github.com/CVer-Yang/RCFSNet. Zhigang Yang 0002, Daoxiang Zhou, Zehua Chen 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | A Context-Aware Road Extraction Method for Remote Sensing Imagery Based on Transformer NetworkabstractIn remote sensing images, roads are usually in complex shapes and can be partially occluded by buildings, trees, and other surroundings. To extract a complete and continuous road network is still a challenging job. This paper proposes a context-aware road extraction method for remote sensing imagery based on Transformer network, in which, a foreground feature enhancement module (FFEM) is designed to further extract detailed road features such as contours from the shallowest feature map; Dual-attention module (DAM) is constructed and applied at different skip connections to make the model focus more on road features in the different level of feature maps; A Swin Transformer-based contextual information extraction module (CIEM) is built between the encoder and decoder modules to capture the global and local road contextual information so as to recover the occluded roads information as much as possible. Furthermore, a multi-scale decoder (M-Decoder) is designed to improve the feature map recovery ability of the decoder module. Experiments on the DeepGlobe road dataset are conducted to verify the efficiency of the proposed method. Xianzhi Ma, Zhigang Yang 0002, Xilin Liu 0003, Zehua Chen 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | TransRoadNet: A Novel Road Extraction Method for Remote Sensing Images via Combining High-Level Semantic Feature and ContextabstractRoad extraction is a significant research hotspot in the area of remote sensing images. Extracting an accurate road network from remote sensing images is still challenging, because some objects in the images are similar to the road, and some results are discontinuous due to the occlusion. Recently, convolutional neural networks (CNNs) have shown their power in a road extraction process. However, the contextual information cannot be captured effectively by those CNNs. Based on CNNs, combining with high-level semantic features and foreground contextual information (FCI), a novel road extraction method for remote sensing images is proposed in this letter. First, the position attention (PA) mechanism is designed to enhance the expression ability for the road feature. Then, the contextual information extraction module (CIEM) is constructed to capture the road contextual information in the images. At last, an FCI supplement module (FCISM) is proposed to provide foreground context information at different stages of the decoder, which can improve the inference ability for the occluded area. Extensive experiments on the DeepGlobal road dataset showed that the proposed method outperforms the existing methods in accuracy, intersection over union (IoU), precision, and$F1$score and yields competitive recall results, which demonstrated the efficiency of the new model. Zhigang Yang 0002, Daoxiang Zhou, Zehua Chen 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |