Haotian Zhang 0010

dblp:83/4184-10 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0005-1296-3984ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 11 since 2021
YearPublicationVenuePosition
2025 A Late-Stage Bitemporal Feature Fusion Network for Semantic Change Detection
abstract
Semantic change detection (SCD) is an important task in geoscience and Earth observation. By producing a semantic change map for each temporal phase, both the land use land cover (LULC) categories and change information can be interpreted. Recently some multitask learning-based SCD methods have been proposed to decompose the task into semantic segmentation (SS) and binary change detection (BCD) subtasks. However, previous works comprise triple branches in an entangled manner, which may not be optimal and hard to adopt foundation models. Besides, lacking explicit refinement of bitemporal features during fusion may cause low accuracy. In this letter, we propose a novel late-stage bitemporal feature fusion network to address the issue. Specifically, we propose local–global attentional aggregation module to strengthen feature fusion, and propose local global context enhancement module to highlight pivotal semantics. Comprehensive experiments are conducted on two public datasets, including SECOND and Landsat-SCD. Quantitative and qualitative results show that our proposed model achieves new state-of-the-art performance on both datasets.
Chenyao Zhou, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 Efficient Semantic Splatting for Remote Sensing Multiview Segmentation
abstract
Remote sensing multi-view image segmentation is essential for achieving accurate and consistent stereoscopic perception of target scenes. This task involves processing RGB images from multiple viewpoints to generate high-accuracy, view-consistent semantic segmentation across all views. Traditional training-based methods struggle with maintaining cross-view consistency, while optimization-driven approaches using implicit neural networks improve view consistency but suffer from slow parameter optimization and inference. To overcome these limitations, we propose a novel Gaussian Splatting-based semantic segmentation framework. Our method efficiently projects the color attributes and semantic features of 3D Gaussians onto the image plane, enabling the simultaneous generation of both RGB images and segmentation outputs. By leveraging explicit spatial structures and a splatting rendering strategy, our approach significantly enhances optimization efficiency and rendering speed. Additionally, we incorporate SAM2 to generate pseudo-labels for boundary regions, addressing the lack of supervision in sparsely labeled views (e.g., 3%). To further enforce cross-view consistency and feature coherence of 3D Gaussians, we introduce a two-level aggregation loss that operates at both the 2D feature map and 3D spatial levels. Extensive experiments across nine datasets demonstrate the superiority of our method, achieving competitive segmentation quality with limited supervisory views. Notably, our approach reduces rendering (inference) times by 90%, while improving the average mIoU by up to 3.5%.
Zipeng Qi, Hao Chen 0045, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 CDMamba: Incorporating Local Clues Into Mamba for Remote Sensing Image Binary Change Detection
abstract
Recently, the Mamba architecture based on state-space models has demonstrated remarkable performance in a series of natural language processing tasks and has been rapidly applied to remote sensing change detection (CD) tasks. However, most methods enhance the global receptive field by directly modifying the scanning mode of Mamba, neglecting the crucial role that local information plays in dense prediction tasks (e.g., binary CD). In this article, we propose a model called CDMamba, which effectively combines global and local features for handling binary CD tasks. Specifically, the scaled residual ConvMamba (SRCM) block is proposed to utilize the ability of Mamba to extract global features and convolution to enhance the local details, to alleviate the issue that current Mamba-based methods lack detailed clues and are difficult to achieve fine detection in dense prediction tasks. Furthermore, considering the characteristics of bi-temporal feature interaction required for CD, the adaptive global–local guided fusion (AGLGF) block is proposed to dynamically facilitate the bi-temporal interaction guided by other temporal global/local features. Our intuition is that more discriminative change features can be acquired with the guidance of other temporal features. Extensive experiments on five datasets demonstrate that our proposed CDMamba is comparable to the current methods (such as the F1/intersection over union (IoU) scores are improved by 2.10%/3.00%, 2.44%/2.91%, on LEVIR+CD and CLCD, respectively). Our code is open-sourced athttps://github.com/zmoka-zht/CDMamba.
Haotian Zhang 0010, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 FoBa: A Foreground-Background Co-Guided Method and New Benchmark for Remote Sensing Semantic Change Detection
abstract
Despite the remarkable progress achieved in remote sensing semantic change detection (SCD), two major challenges remain. At the data level, existing SCD datasets suffer from limited change categories, insufficient change types, and a lack of fine-grained class definitions, making them inadequate to fully support practical applications. At the methodological level, most current approaches underutilize change information, typically treating it as a post-processing step to enhance spatial consistency, which constrains further improvements in model performance. To address these issues, we construct a new benchmark for remote sensing SCD, LevirSCD. Focused on the Beijing area, the dataset covers 16 change categories and 210 specific change types, with more fine-grained class definitions (e.g., roads are divided into unpaved and paved roads). Furthermore, we propose a foreground-background co-guided SCD (FoBa) method, which leverages foregrounds that focus on regions of interest and backgrounds enriched with contextual information to guide the model collaboratively, thereby alleviating semantic ambiguity while enhancing its ability to detect subtle changes. Considering the requirements of bi-temporal interaction and spatial consistency in SCD, we introduce a gated interaction fusion (GIF) module along with a simple consistency loss to further enhance the model’s detection performance. Extensive experiments on three datasets (SECOND, JL1, and the proposed LevirSCD) demonstrate that FoBa achieves competitive results compared to current SOTA methods, with improvements of 1.48%, 3.61%, and 2.81% in the SeK metric, respectively. Our code and dataset are available at https://github.com/zmoka-zht/FoBa.
Haotian Zhang 0010, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Time Travelling Pixels: Bitemporal Features Integration with Foundation Model for Remote Sensing Image Change Detection
abstract
Change detection, a prominent research area in remote sensing, is pivotal in observing and analyzing surface transformations. Despite significant advancements achieved through deep learning-based methods, executing high-precision change detection in spatiotemporally complex remote sensing scenarios still presents a substantial challenge. The recent emergence of foundation models, with their powerful universality and generalization capabilities, offers potential solutions. However, bridging the gap of data and tasks remains a significant obstacle. In this paper, we introduce Time Travelling Pixels (TTP), a novel approach that integrates the latent knowledge of the SAM foundation model into change detection. TTP can effectively address the domain shift in general knowledge transfer and the challenge of expressing homogeneous and heterogeneous characteristics of multi-temporal images. The state-of-the-art results obtained on the LEVIR-CD underscore the efficacy of the TTP. The code has been made publicly available at https://github.com/KyanChen/TTP.
Keyan Chen 0001, Chengyang Liu, Wenyuan Li 0002, Hao Chen 0045, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IGARSS6
2024 Pixel-Level Change Detection Pseudo-Label Learning For Remote Sensing Change Captioning
abstract
The existing Remote Sensing Image Change Captioning (RSICC) methods perform well in simple scenes but exhibit poorer performance in complex scenes. This limitation is primarily attributed to the model’s constrained visual ability to distinguish and locate changes. Acknowledging the inherent correlation between change detection (CD) and RSICC tasks, we believe pixel-level CD is significant for describing the differences between images through language. Regrettably, the current RSICC dataset lacks readily available pixel-level CD labels. To address this deficiency, we leverage a model trained on existing CD datasets to derive CD pseudo-labels. We propose an innovative network with an auxiliary CD branch, supervised by pseudo-labels. Furthermore, a semantic fusion augment (SFA) module is proposed to fuse the feature information extracted by the CD branch, thereby facilitating the nuanced description of changes. Experiments demonstrate that our method achieves state-of-the-art performance and validate that learning pixel-level CD pseudo-labels significantly contributes to change captioning.
Keyan Chen 0001, Zipeng Qi, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IGARSS5
2024 RSCaMa: Remote Sensing Image Change Captioning With State Space Model
abstract
Remote Sensing Image Change Captioning (RSICC) aims to describe surface changes between multi-temporal remote sensing images in language, including the changed object categories, locations, and dynamics of changing objects (e.g., added or disappeared). This poses challenges to spatial and temporal modeling of bi-temporal features. Despite previous methods progressing in the spatial change perception, there are still weaknesses in joint spatial-temporal modeling. To address this, in this paper, we propose a novel RSCaMa model, which achieves efficient joint spatial-temporal modeling through multiple CaMa layers, enabling iterative refinement of bi-temporal features. To achieve efficient spatial modeling, we introduce the recently popular Mamba (a state space model) with a global receptive field and linear complexity into the RSICC task and propose the Spatial Difference-aware SSM (SD-SSM), overcoming limitations of previous CNN- and Transformer-based methods in the receptive field and computational complexity. SD-SSM enhances the model’s ability to capture spatial changes sharply. In terms of efficient temporal modeling, considering the potential correlation between the temporal scanning characteristics of Mamba and the temporality of the RSICC, we propose the Temporal-Traversing SSM (TT-SSM), which scans bi-temporal features in a temporal cross-wise manner, enhancing the model’s temporal understanding and information interaction. Experiments validate the effectiveness of the efficient joint spatial-temporal modeling and demonstrate the outstanding performance of RSCaMa and the potential of the Mamba in the RSICC task. Additionally, we systematically compare three different language decoders, including Mamba, GPT-style decoder, and Transformer decoder, providing valuable insights for future RSICC research. The code will be available at https://github.com/Chen-Yang-Liu/RSCaMa.
Keyan Chen 0001, Bowen Chen 0002, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation Based on Visual Foundation Model
abstract
Leveraging the extensive training data from SA-1B, the Segment Anything Model (SAM) demonstrates remarkable generalization and zero-shot capabilities. However, as a category-agnostic instance segmentation method, SAM heavily relies on prior manual guidance, including points, boxes, and coarse-grained masks. Furthermore, its performance in remote sensing image segmentation tasks remains largely unexplored and unproven. In this paper, we aim to develop an automated instance segmentation approach for remote sensing images, based on the foundational SAM model and incorporating semantic category information. Drawing inspiration from prompt learning, we propose a method to learn the generation of appropriate prompts for SAM. This enables SAM to produce semantically discernible segmentation results for remote sensing images, a concept we have termed RSPrompter. We also propose several ongoing derivatives for instance segmentation tasks, drawing on recent advancements within the SAM community, and compare their performance with RSPrompter. Extensive experimental results, derived from the WHU building, NWPU VHR-10, and SSDD datasets, validate the effectiveness of our proposed method. The code for our method is publicly available at https://kychen.me/RSPrompter.
Keyan Chen 0001, Hao Chen 0045, Haotian Zhang 0010, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and Analysis
abstract
Monitoring changes in the Earth’s surface is crucial for understanding natural processes and human impacts, necessitating precise and comprehensive interpretation methodologies. Remote sensing (RS) satellite imagery offers a unique perspective for monitoring these changes, leading to the emergence of RS image change interpretation (RSICI) as a significant research focus. Current RSICI technology encompasses change detection and change captioning, each with its limitations in providing comprehensive interpretation. To address this, we propose an interactive Change-Agent, which can follow user instructions to achieve comprehensive change interpretation and insightful analysis, such as change detection and change captioning, change object counting, and change cause analysis. The Change-Agent integrates a multilevel change interpretation (MCI) model as the eyes and a large language model (LLM) as the brain. The MCI model contains two branches of pixel-level change detection and semantic-level change captioning, in which the BI-temporal iterative interaction (BI3) layer is proposed to enhance the model’s discriminative feature representation capabilities. To support the training of the MCI model, we build the LEVIR-MCI dataset with a large number of change masks and captions of changes. Experiments demonstrate the state-of-the-art (SOTA) performance of the MCI model in achieving both change detection and change description simultaneously and highlight the promising application value of our Change-Agent in facilitating comprehensive interpretation of surface changes, which opens up a new avenue for intelligent RS applications. To facilitate future research, we will make our dataset and codebase publicly available athttps://github.com/Chen-Yang-Liu/Change-Agent.
Keyan Chen 0001, Haotian Zhang 0010, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 BiFA: Remote Sensing Image Change Detection With Bitemporal Feature Alignment
abstract
Despite the success of deep learning-based change detection methods, their existing insufficiency in temporal (channel, spatial) and multi-scale alignment have rendered them insufficient capability in mitigating external factors (illumination changes and perspective differences, etc.) arising from different imaging conditions during change detection. In this paper, a Bi-temporal Feature Alignment (BiFA) model is proposed to produce a precise change detection map in a lightweight manner by reducing the impact of irrelevant factors. Specifically, for the temporal alignment, the Bi-temporal Interaction (BI) module is proposed to realize the alignment of the bi-temporal image channel level. Our intuition is introducing the bi-temporal interaction in the feature extraction stage may benefit suppressing the interference, such as illumination changes. Simultaneously, the Alignment module based on Differential Flow Field (ADFF) is proposed to explicitly estimate the offset of the bi-temporal image and realize their spatial level alignment to mitigate the inadequate registration resulting from different perspectives. Furthermore, for the multi-scale alignment, we introduce the Implicit Neural alignment Decoder (IND) to produce more refined prediction maps achieving precise alignment of multi-scale features by learning continuous image representations in coordinate space. Our BiFA outperforms other state-of-the-art methods on six datasets (such as the F1/IoU scores are improved by 2.70%/3.91%, 2.01%/2.94% on LEVIR+-CD and SYSU-CD, respectively) and displays greater robustness in cross-resolutions change detection. Our code is available at https://github.com/zmoka-zht/BiFA.
Haotian Zhang 0010, Hao Chen 0045, Chenyao Zhou, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Continuous Cross-Resolution Remote Sensing Image Change Detection
abstract
Most contemporary supervised Remote Sensing (RS) image Change Detection (CD) approaches are customized for equal-resolution bitemporal images. Real-world applications raise the need for cross-resolution change detection, aka, CD based on bitemporal images with different spatial resolutions. Given training samples of a fixed bitemporal resolution difference (ratio) between the high-resolution (HR) image and the low-resolution (LR) one, current cross-resolution methods may fit a certain ratio but lack adaptation to other resolution differences. Toward continuous cross-resolution CD, we propose scale-invariant learning to enforce the model consistently predicting HR results given synthesized samples of varying resolution differences. Concretely, we synthesize blurred versions of the HR image by random downsampled reconstructions to reduce the gap between HR and LR images. We introduce coordinate-based representations to decode per-pixel predictions by feeding the coordinate query and corresponding multi-level embedding features into an MLP that implicitly learns the shape of land cover changes, therefore benefiting recognizing blurred objects in the LR image. Moreover, considering that spatial resolution mainly affects the local textures, we apply local-window self-attention to align bitemporal features during the early stages of the encoder. Extensive experiments on two synthesized and one real-world different-resolution CD datasets verify the effectiveness of the proposed method. Our method significantly outperforms several vanilla CD methods and two cross-resolution CD methods on the three datasets both in in-distribution and out-of-distribution settings. The empirical results suggest that our method could yield relatively consistent HR change predictions regardless of varying bitemporal resolution ratios. Our code will be public.
Hao Chen 0045, Haotian Zhang 0010, Keyan Chen 0001, Chenyao Zhou, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2