Wei Zhang 0250

dblp:10/4661-250 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-4655-356XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MSDP-Net: Multi-scale distribution perception network for rotating object detection in remote sensing
Wei Zhang 0250, Qiang Li 0042, Qi Wang 0009
Pattern Recognit.3
2026 RAPTOR: Rotational Adaptive Parallel Topology for Object Detection in remote sensing
Wei Zhang 0250, Qiang Li 0042, Qi Wang 0009
Pattern Recognit.3
2025 Multibranch Mutual-Guiding Learning for Infrared Small Target Detection
abstract
At present, many infrared target detection approaches focus on designing modules that address the two key characteristics of targets: their weak signals and small size. However, these approaches often fail to fully leverage guided learning for weak and small target content, resulting in sub-optimal detection performance, particularly in terms of shape preservation and target positioning. To tackle this challenge, this paper proposes a multi-branch mutual-guiding learning network (MMLNet) that enhances the accuracy of infrared target detection, even in the absence of clear morphological and textural features in images. The method consists of three branches: edge, positioning, and detection, each of which is designed with a specialized module from a unique perspective. In the detection branch, we introduce a multi-dimensional lossless encoder optimized through a downsampling strategy and multi-level feature fusion to mitigate feature loss in small targets. In the positioning branch, a target positioning strategy is proposed to explicitly identify candidate targets from the image by means of a learnable multi-kernel pattern. In the edge branch, a simple architecture is adopted to enhance the ability of the model to preserve the target shape. To effectively utilize the knowledge of different branches, a mutual-guiding fusion module is developed to adjust information within and between branches. The manner adaptively utilizes the specific knowledge from each input branch. Experiment results demonstrate that the proposed method achieves comparable performance, and the visualization results show the advantages of our method in shape preservation and positioning of the targets. Our code is publicly available at https://github.com/qianngli/MMLNet.
Qiang Li 0042, Wei Zhang 0250, Wanxuan Lu, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.2
2025 Refined Cascade Cost Volume for Multiview Remote Sensing Image Reconstruction
abstract
Research on remote sensing multi-view stereo has significantly advanced the development of large-scale 3D urban reconstruction. However, existing frameworks encounter challenges with blurred edge details when processing aerial image, which impedes the accuracy of depth estimation. To address these limitations, we propose RC-MVS, the deep estimation network specifically tailored for remote sensing multi-view stereo tasks. This network aims to enhance the geometric details within the view space while effectively reducing match noise, achieving high-precision depth estimation. Specifically, we introduce a refined cascade framework that integrates geometric details with semantic information, ensuring both global structural consistency and local feature expressiveness. During the feature extraction phase, we redesign the feature space construction process and introduce a denoising feature pyramid module. This module reduces feature inconsistency and employs multiple denoising strategies to purify feature representations, thereby enhancing the accuracy of the matching process. Furthermore, to achieve progressive optimization of the depth range, we propose a progressive cross-layer fusion module. This module progressively fuses low-resolution cost volumes, reducing domain shifts between different data dimensions, thereby enhancing the understanding of fine structures within the depth map and the broader context. Experimental results show that the RC-MVS model performs exceptionally well on the LuoJia-MVS and WHU datasets, achieving superior quantitative and qualitative performance.
Wei Zhang 0250, Qiang Li 0042, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2025 Semantic-Guided Multiview Stereo Reconstruction for Aerial Image
abstract
The application of learning-based Multi-view Stereo (MVS) depth estimation methods has achieved significant results in large-scale 3D reconstruction benchmarks. However, adjacent terrains in aerial image interfere with depth estimation along building edges during matching process, leading to inaccurate results. To address these challenges, we propose a new end-to-end MVS network, named FuS-MVSNet, which fuses monocular depth probability as a semantic guidance into the multi-view geometry-based MVS framework. By combining the strengths of geometric consistency and local semantics, FuS-MVSNet achieves notable enhancements in both accuracy and robustness. Specifically, we first construct a monocular branch based on the pre-trained Depth Anything model to perform monocular metric depth estimation. The non-shared parameters ensure that the depth estimation process is independent of multi-view branch, focusing exclusively on semantic depth inference. Subsequently, to incorporate monocular features into the multi-view network, we introduce a volume adaptive fusion module, which adaptively integrates monocular feature volumes into the standard cost volume via an attention mechanism and guides the cost volume regularization. Finally, confidence-based dynamic selection between the two prediction branches ensures the selection of the more robust branch result under challenging conditions. Qualitative and quantitative results indicate that we achieve competitive performance on multiple benchmarks, including the WHU and LuoJia-MVS datasets.
Wei Zhang 0250, Zhigang Yang 0002, Qiang Li 0042, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1
2024 C²Net: Road Extraction via Context Perception and Cross Spatial-Scale Feature Interaction
abstract
Road extraction from remote sensing images (RSIs) holds significant application value in various aspects of daily scenarios. However, it is still challenging to extract high-quality road results from RSIs due to the interference of objects sharing similar structures with roads in the background and the occlusion caused by surroundings. To alleviate these problems, a road extraction network based on the global-local Context perception and Cross spatial-scale feature interaction is proposed ($\text {C}^{2}$Net). First, a global-local context perception module (GLCPM) is incorporated to capture the overall topology features of the road, which aims to improve the ability of the model to discriminate between roads and similar objects. Then, the cross spatial-scale feature interaction module is designed in the skip connection to effectively aggregate full-scale features without loss of feature information, which can provide rich and accurate road structural features for the decoder. Experiments conducted on public road datasets demonstrate that$\text {C}^{2}$Net outperforms existing methods in terms of comprehensive metrics such as intersection over union (IoU) and the$F1$-score. The results indicate that$\text {C}^{2}$Net can produce road results with superior connectivity and quality. The source code will be publicly available athttps://github.com/CVer-Yang/CCNet.
Zhigang Yang 0002, Wei Zhang 0250, Qiang Li 0042, Weiping Ni, Junzheng Wu, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.2
2024 Visual Consistency Enhancement for Multiview Stereo Reconstruction in Remote Sensing
abstract
Learnable multiview stereo (MVS) aerial image depth estimation has obtained great success in 3-D digital urban reconstruction. Currently, most depth estimation methods in the large-scale sense heavily involve adapting the general MVS framework. However, these methods often overlook the cross-view interval and limited viewpoint inherent in aerial images data. In this article, we introduce an learning-based MVS method for aerial image depth estimation, which enhances visual consistency to address the insufficient accuracy caused by the characteristics of aerial image data, namely, AggrMVS. First, an optical flow-guided feature extraction module is introduced to map the dynamic relationship between reference and source images. It explicitly captures edge information of different depth components to guide the cost volume regularization. Second, a cross-view volume fusion module is proposed to enhance the interaction among reference volumes, further improving the aggregation ability of the source volume. Furthermore, AggrMVS achieves refined aerial image depth estimation results with a lightweight cascade architecture. Since low-altitude oblique aerial datasets currently lack, we reconstruct a multicategory synthetic aerial scene benchmark from general MVS datasets. The benchmark dataset is available athttps://github.com/ToscW/BlendedUAV. Experiments on public and proposed datasets confirm that AggrMVS outperforms other MVS depth estimation methods in terms of qualitative and quantitative aspects.
Wei Zhang 0250, Qiang Li 0042, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.1