VLDB 2026 Research / reviewers in the wild / expert
Fahong Zhang 0001
dblp:211/1952-1
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-0209-8841ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Global Collinearity-Aware Polygonizer for Polygonal Building Mapping in Remote SensingabstractThis paper addresses the challenge of mapping polygonal buildings from remote sensing images and introduces a novel algorithm, the Global Collinearity-aware Polygonizer (GCP). GCP, built upon an instance segmentation framework, processes binary masks produced by any instance segmentation model. The algorithm begins by collecting polylines sampled along the contours of the binary masks. These polylines undergo a refinement process using a transformer-based regression module to ensure they accurately fit the contours of the targeted building instances. Subsequently, a collinearity-aware polygon simplification module simplifies these refined polylines and generate the final polygon representation. This module employs dynamic programming technique to optimize an objective function that balances the simplicity and fidelity of the polygons, achieving globally optimal solutions. Furthermore, the optimized collinearity-aware objective is seamlessly integrated into network training, enhancing the cohesiveness of the entire pipeline. The effectiveness of GCP has been validated on three public benchmarks for polygonal building mapping. Further experiments reveal that applying the collinearity-aware polygon simplification module to arbitrary polylines, without prior knowledge, enhances accuracy over traditional methods such as the Douglas-Peucker algorithm. This finding underscores the broad applicability of GCP. The code for the proposed method will be made available at https://github.com/zhu-xlab/GCP. Fahong Zhang 0001, Yilei Shi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | One for All: Toward Unified Foundation Models for Earth VisionabstractFoundation models characterized by extensive parameters and trained on large-scale datasets have demonstrated remarkable efficacy across various downstream tasks for remote sensing data. Current remote sensing foundation models typically specialize in a single modality or a specific spatial resolution range, limiting their versatility for downstream datasets. While there have been attempts to develop multi-modal remote sensing foundation models, they typically employ separate vision encoders for each modality or spatial resolution, necessitating a switch in backbones contingent upon the input data. To address this issue, we introduce a simple yet effective method, termed OFA-Net (One-For-All Network): employing a single, shared Transformer backbone for multiple data modalities with different spatial resolutions. Using the masked image modeling mechanism, we pre-train a single Transformer backbone on a curated multi-modal dataset with this simple design. Then the backbone model can be used in different downstream tasks, thus forging a path towards a unified foundation backbone model in Earth vision. The proposed method is evaluated on 12 distinct downstream tasks and demonstrates promising performance. Zhitong Xiong, Yi Wang 0072, Fahong Zhang 0001, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2024 | Few-Shot Object Detection in Remote Sensing: Lifting the Curse of Incompletely Annotated Novel ObjectsabstractObject detection is an essential and fundamental task in computer vision and satellite image processing. Existing deep learning methods have achieved impressive performance thanks to the availability of large-scale annotated datasets. Yet, in real-world applications the availability of labels is limited. In this context, few-shot object detection (FSOD) has emerged as a promising direction, which aims at enabling the model to detect novel objects with only few of them annotated. However, many existing FSOD algorithms overlook a critical issue: when an input image contains multiple novel objects and only a subset of them are annotated, the unlabeled objects will be considered as background during training. This can cause confusions and severely impact the model’s ability to recall novel objects. To address this issue, we propose a self-training-based FSOD (ST-FSOD) approach, which incorporates the self-training mechanism into the few-shot fine-tuning process. ST-FSOD aims to enable the discovery of novel objects that are not annotated, and take them into account during training. On the one hand, we devise a two-branch region proposal networks (RPN) to separate the proposal extraction of base and novel objects, On another hand, we incorporate the student-teacher mechanism into RPN and the region of interest (RoI) head to include those highly confident yet unlabeled targets as pseudo labels. Experimental results demonstrate that our proposed method outperforms the state-of-the- art in various FSOD settings by a large margin. The codes will be publicly available at https://github.com/zhu-xlab/ST-FSOD. Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Pseudo Features-Guided Self-Training for Domain Adaptive Semantic Segmentation of Satellite ImagesabstractSemantic segmentation is a fundamental and crucial task that is of great importance to real-world satellite image-based applications. Yet a widely acknowledged issue that occurs when applying the semantic segmentation models to unseen scenery is that the model will perform much poorer than when it was applied to scenery similar to the training data. This phenomenon is usually termed as the domain shift problem. To tackle it, this article presents a self-training-based unsupervised domain adaptation (UDA) method. Different from the previous self-training approaches which focus on rectifying and improving the quality of the pseudo labels, we instead seek to exploit feature-level relation among neighboring pixels to structure and regularize the prediction of the adapted model. Based on the assumption that spatial topological relation is maintained despite the impact of the domain shift, we propose a novel self-training mechanism to perform DA by exploiting local relation in the feature space spanned by the teacher model, from which the pseudo labels are generated. Quantitative experiments on four different public benchmarks demonstrate that the proposed method can outperform the other UDA methods. Besides, analytical experiments also intuitively verify the proposed assumption. Codes will be publicly available athttps://github.com/zhu-xlab/PFST. Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Wei Huang 0068, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Domain-Agnostic Domain Adaption for Building Footprint ExtractionabstractFor global range satellite imaging mission, images captured from different areas may have large distribution biases due to different illuminations, shooting angles and atmospheric conditions. A straightforward idea to mitigate this problem is to categorize the images into different domains according the cities they belong to, and apply domain adaptation approaches. However, categorization by cities becomes unreasonable with the increase of the city number, and the emergence of inter-city similarity and intra-city discrepancy. With such consideration, this paper proposes a novel domain adaptation method named domain-agnostic domain adaptation (DADA) to reduce the distribution biases without explicitly defining the domain each image belongs to. To implement this, we augment the images to the styles of different domains by Generative Adversarial Networks (GAN) and contrastive learning to increase the generalizability of down-stream tasks. Experiments on Planetscope building footprint extraction datasets verify the effectiveness of our method. Fahong Zhang 0001, Yilei Shi, Xiao Xiang Zhu 0001 |
IGARSS | 1 |