Tianyang Zhang 0002

dblp:160/9939-2 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0001-9079-7970ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Tiny object detection based on dynamic scale-awareness label assignment and contextual enhancement
Tianyang Zhang 0002, Xiangrong Zhang, Chaozhuo Hua, Guanchun Wang, Xiao Han 0012, Licheng Jiao
Pattern Recognit.1
2025 RegionMatch: Pixel-Region Collaboration for Semi-Supervised Semantic Segmentation in Remote Sensing Images
abstract
Semi-supervised semantic segmentation (S4) has shown significant promise in reducing the burden of labor-intensive data annotation. However, existing methods mainly rely on pixel-level information, neglecting the strong region consistency inherent in remote sensing images (RSIs), which limits their effectiveness in handling the complex and diverse backgrounds of RSIs. To address this, we propose RegionMatch, a novel approach that leverages unlabeled data from a fresh object-level perspective, which is more tailored to the nature of semantic segmentation. We design the Pixel-Region Synergy Pseudo-Labeling strategy, which explicitly injects object-level contextual information into the S4 pipeline and promotes knowledge collaboration between pixel and region perspectives for generating high-quality pseudo-labels. In addition, we propose the Region Structure-Aware Correlation Consistency, which models object-level relationships by establishing inter-region correlations across images and pixel correlations within regions, providing more effective supervision signals for unlabeled data. Experimental results demonstrate that RegionMatch outperforms state-of-the-art methods on multiple authoritative remote sensing datasets, highlighting its superiority in the RSIs.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Chaowei Fang, Xu Tang 0004, Licheng Jiao
IJCAI3
2025 ACMamba: Fast Unsupervised Anomaly Detection via An Asymmetrical Consensus State Space Model
abstract
Unsupervised anomaly detection in hyperspectral images (HSI), aiming to detect unknown targets from backgrounds, is challenging for earth surface monitoring. However, current studies are hindered by steep computational costs due to the high-dimensional property of HSI and dense sampling-based training paradigm, constraining their rapid deployment. Our key observation is that, during training, not all samples within the same homogeneous area are indispensable, whereas ingenious sampling can provide a powerful substitute for reducing costs. Motivated by this, we propose an Asymmetrical Consensus State Space Model (ACMamba) to significantly reduce computational costs without compromising accuracy. Specifically, we design an asymmetrical anomaly detection paradigm that utilizes region-level instances as an efficient alternative to dense pixel-level samples. In this paradigm, a low-cost Mamba-based module is introduced to discover global contextual attributes of regions that are essential for HSI reconstruction. Additionally, we develop a consensus learning strategy from the optimization perspective to simultaneously facilitate background reconstruction and anomaly compression, further alleviating the negative impact of anomaly reconstruction. Theoretical analysis and extensive experiments across eight benchmarks verify the superiority of ACMamba, demonstrating a faster speed and stronger performance over the state-of-the-art. Code is released at https://github.com/PURE-melo/ACMamba.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Licheng Jiao
ACM Multimedia5
2025 OraL: An Observational Learning Paradigm for Unsupervised Hyperspectral Change Detection
abstract
Unsupervised hyperspectral change detection (UHCD), detecting subtle changes between bi-temporal images without manual annotations, is an essential but challenging task in the earth observation community. The current modus operandi often performs it in a feature comparison manner, which is limited by variations in imaging conditions. We observe that fully supervised paradigms using limited annotations are capable of overcoming this challenge. Based on this, we introduce a novel Observational Learning Paradigm (OraL) for UHCD by mimicking fully supervised paradigms. OraL comprises two sequential stages: Observation, which designs a spatial-temporal observation strategy (STO) that records the learning consistency of pixels under different training steps and views, to obtain reliable pseudo-labels. Reproduction, which retrains the model with these pseudo-labels and introduces a distribution-aware spectral learning strategy (DSL) to adaptively increase their learning difficulty according to spectral distributions, enhancing the robustness and generalization of the model. Extensive experiments on several public hyperspectral image datasets demonstrate its state-of-the-art performance and pluggability for previous unsupervised methods. Code will be made available.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Shunli Tian, Tianyang Zhang 0002, Xu Tang 0004, Licheng Jiao
IEEE Trans. Circuits Syst. Video Technol.5
2025 S2Mamba: A Spatial-Spectral State Space Model for Hyperspectral Image Classification
abstract
The land cover analysis using hyperspectral images (HSIs) remains an open problem due to their low spatial resolution and complex spectral information. Recent studies are primarily dedicated to designing Transformer-based architectures for spatial-spectral long-range dependencies modeling, which is computationally expensive with quadratic complexity. Selective structured state space model (SSM; Mamba), which is efficient for modeling long-range dependencies with linear complexity, has recently shown promising progress. However, its potential in HSI processing that requires handling numerous spectral bands has not yet been explored. In this article, we innovatively propose S2Mamba, a spatial-spectral SSM for HSI classification, to excavate spatial-spectral contextual features, resulting in more efficient and accurate land cover analysis. In S2Mamba, two selective structured SSMs through different dimensions are designed for feature extraction, one for spatial, and the other for spectral, along with a spatial-spectral mixture gate (SMG) for optimal fusion. More specifically, S2Mamba first captures spatial contextual relations by interacting each pixel with its adjacent through a patch cross scanning (PCS) module and then explores semantic information from continuous spectral bands through a bidirectional spectral scanning (BSS) module. Considering the distinct expertise of the two attributes in homogenous and complicated texture scenes, we realize the SMG by a group of learnable matrices, allowing for the adaptive incorporation of representations learned across different dimensions. Extensive experiments conducted on HSI classification benchmarks demonstrate the superiority and prospect of S2Mamba. The code will be made available at:https://github.com/PURE-melo/S2Mamba.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 Beyond Single Pixel: Context Priors Guided Semi-Supervised Building Footprint Segmentation
abstract
Automated building footprint segmentation is crucial in remote sensing with widespread applications in various fields. Semi-supervised semantic segmentation methods are gaining traction in the remote sensing community as they significantly reduce the need for labor-intensive pixel-level annotations in training segmentation models. However, these methods typically rely on individual pixel-level supervision for unlabeled data, neglecting the contextual relationships between pixels. This limits their potential to exploit unlabeled data. To bridge this gap, this paper proposes a novel Context Priors Guided Semi-Supervised Building Footprint Segmentation method that leverages contextual relationships among numerous unlabeled pixels to build supervisory signals that extend beyond individual pixel-level guidance for learning on unlabeled data. The CPG comprises two main components: Spatial context priors-guided pseudo-label regularization and semantic context priors-guided representation learning. By integrating contextual knowledge from both pixel spatial locations and semantic representation spaces, these components capture comprehensive class semantic attributes and enable the model to learn complete shapes of building footprints. Our approach achieves state-of-the-art performance on three publicly available building footprint segmentation datasets, validating its effectiveness.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xiao Han 0012, Licheng Jiao, Lianchao Zhang
IEEE Trans. Geosci. Remote. Sens.3
2025 Temporal-Feedback Self-Training for Semi-Supervised Object Detection in Remote Sensing Images
abstract
Although modern Remote Sensing Object Detection (RSOD) methods have achieved advanced performance, they heavily rely on a large amount of annotated data. This paper explores semi-supervised RSOD to mitigate annotation costs, leveraging recent extensive research in generic Semi-Supervised Object Detection (SSOD) based on the self-training paradigm. Current SSOD methods encounter challenges in adapting to remote sensing images due to the complexity and variability of RSIs. Two key issues remain underexplored: the noise in pseudo-labels caused by model instability and the difficulty in distinguishing similar categories. This paper introduces the Temporal-Feedback Self-Training (TST) framework, a novel approach to tackle these challenges in semi-supervised RSOD. TST consists of two components: Temporal Consistency Based Pseudo-labels Certainty Estimation (TCE) and Temporal Self-Feedback Feature Refinement (TSF). TCE addresses pseudo-label noise during training by evaluating the stability of pseudo-label classification and localization over time series to assess the quality of pseudo-labels. On the other hand, TSF enhances pseudo-label quality by dynamically identifying the models confusing categories as feedback for feature refinement. Both components facilitate the progression of the self-training-based RSOD during training. We conducted extensive experiments on two challenging public datasets, DOTA and DIOR. The results demonstrate that the proposed TST and TCE components significantly improve the baseline models performance, surpassing the state-of-the-art generic SSOD method. This suggests that our approach is more effective than generic SSOD methods in addressing the challenges posed by remote sensing images.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2025 Negative Deterministic Information-Based Multiple Instance Learning for Weakly Supervised Object Detection and Segmentation
abstract
Weakly supervised object detection (WSOD) and semantic segmentation with image-level annotations have attracted extensive attention due to their high label efficiency. Multiple instance learning (MIL) offers a feasible solution for the two tasks by treating each image as a bag with a series of instances (object regions or pixels) and identifying foreground instances that contribute to bag classification. However, conventional MIL paradigms often suffer from issues, e.g., discriminative instance domination and missing instances. In this article, we observe that negative instances usually contain valuable deterministic information, which is the key to solving the two issues. Motivated by this, we propose a novel MIL paradigm based on negative deterministic information (NDI), termed NDI-MIL, which is based on two core designs with a progressive relation: NDI collection and negative contrastive learning (NCL). In NDI collection, we identify and distill NDI from negative instances online by a dynamic feature bank. The collected NDI is then utilized in a NCL mechanism to locate and punish those discriminative regions, by which the discriminative instance domination and missing instances issues are effectively addressed, leading to improved object- and pixel-level localization accuracy and completeness. In addition, we design an NDI-guided instance selection (NGIS) strategy to further enhance the systematic performance. Experimental results on several public benchmarks, including PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO, show that our method achieves satisfactory performance. The code is available at: https://github.com/GC-WSL/NDI.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.4
2024 Multistage Enhancement Network for Tiny Object Detection in Remote Sensing Images
abstract
With the rapid advances in deep learning techniques, remote sensing object detection has achieved remarkable achievements in recent years. However, tiny object detection remains unsatisfactory and suffers from two main drawbacks, including (1) the high sensitivity of IoU for location deviation in tiny objects and (2) the poor-quality feature representations of tiny objects. To address the aforementioned problems, we propose a Multi-stage Enhancement Network (MENet) that achieves the instance-level and feature-level enhancement of tiny objects from different stages of the detector. Since the IoU-based label assignment drastically deteriorates the positive samples for tiny objects, we first propose a Central Region-based (CR) label assignment to substitute it in the Region Proposal Network (RPN). The CR label assignment regards the anchors that fall into the central region of ground-truth boxes as positive samples, which provides more positive samples for tiny objects. Then, we design a Gated Context Aggregation (GCA) module that selectively aggregates valuable context information to enhance the feature representation of tiny objects. Additionally, we devise a positive RoI feature (pRoI) generator in the Region Convolutional Neural Network (R-CNN) to generate a rich diversity of high-quality positive RoI features for tiny objects. We conduct extensive experiments on AI-TOD and SODA-A datasets, and the results demonstrate the effectiveness of our proposed method.
Tianyang Zhang 0002, Xiangrong Zhang, Xiaoqian Zhu, Guanchun Wang, Xiao Han 0012, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2023 Semantics and Contour Based Interactive Learning Network for Building Footprint Extraction
abstract
Building footprint extraction plays an important role in the analysis of remote sensing images and has an extensive range of applications. Obtaining precise boundaries of buildings remains a challenge in existing building extraction methods. Some previous works have made notable efforts to address this concern. However, most of these methods require cumbersome and expensive post-processing steps. Moreover, they ignored the correlation between building semantics and contours, which we believe is crucial for building footprint extraction. To mitigate this issue, our paper presents an intuitive and effective framework that explores semantic and contour cues of buildings and fully excavates their correlation. Specifically, we construct an interactive dual-stream decoder. The Intermediate connections within this decoder interactively transmit features between branches, contributing to learning correlations between semantics and contours. We propose the Semantic Collaboration Module (SCM) to strengthen the connection between the two branches. To further boost performance, we build the Multi-Scale Semantic Context Fusion Module (MSCF) to fuse semantic information from the higher and lower layers of the network, allowing the network to obtain superior feature representations. The experimental results on the WHU, INRIA, and Massachusetts building datasets demonstrate the superior performance of our method.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 Mask Decoupled Head for Instance Segmentation in Remote Sensing Images
abstract
Instance segmentation predicts the categories of all instances and locates them using pixel-level masks. Although existing methods have shown exemplary performance, the poor boundaries due to the lack of fine-grained information re-mains a challenge for RSIs instance segmentation. In this paper, to address the problem, we propose a novel instance segmentation branch, namely Mask Decoupled Head, which is mainly composed of a Feature Enhance Module (FEM) and a Feature Decoupled Module (FDM). FEM enhances the rep-resentation of the body features through the low-frequency component of images. FDM decouples the segmentation task by supervising body and edge separately and leverages fine-grained information to complement the boundary details. We performed comprehensive experiments on NWPU VHR -10 and HRSID datasets to evaluate the effectiveness of our pro-posed method and achieved good performance.
Xiangrong Zhang, Tianyang Zhang 0002, Xiaoqian Zhu, Xu Tang 0004, Licheng Jiao
IGARSS3
2022 Semantic Attention and Scale Complementary Network for Instance Segmentation in Remote Sensing Images
abstract
In this article, we focus on the challenging multicategory instance segmentation problem in remote sensing images (RSIs), which aims at predicting the categories of all instances and localizing them with pixel-level masks. Although many landmark frameworks have demonstrated promising performance in instance segmentation, the complexity in the background and scale variability instances still remain challenging, for instance, segmentation of RSIs. To address the above problems, we propose an end-to-end multicategory instance segmentation model, namely, the semantic attention (SEA) and scale complementary network, which mainly consists of a SEA module and a scale complementary mask branch (SCMB). The SEA module contains a simple fully convolutional semantic segmentation branch with extra supervision to strengthen the activation of interest instances on the feature map and reduce the background noise's interference. To handle the undersegmentation of geospatial instances with large varying scales, we design the SCMB that extends the original single mask branch to trident mask branches and introduces complementary mask supervision at different scales to sufficiently leverage the multiscale information. We conduct comprehensive experiments to evaluate the effectiveness of our proposed method on the iSAID dataset and the NWPU Instance Segmentation dataset and achieve promising performance.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Cybern.1
2022 Remote Sensing Image Super-Resolution via Dual-Resolution Network Based on Connected Attention Mechanism
abstract
Limited by hardware conditions and complex degradation processes, the obtained remote sensing images (RSIs) are often low-resolution (LR) data with insufficient high-frequency information. Image super-resolution (SR) aims to improve the spatial resolution of images and add reasonable detailed information. Although existing convolutional neural network (CNN)-based methods achieve good performance by adding residual structure and attention mechanism to the network, simply stacking the residual structure and embedding the attention module directly on the residual branch lead to localized use of features and information loss. To address the above problems, we propose a dual-resolution connected attention network (DRCAN). Specifically, a high-resolution (HR) learning branch is constructed to complement the mapping learning between LR images and HR images, and a connected attention module with residual learning is introduced to make full use of the different levels of intermediate layer features. Besides, we collect data at different resolutions from Google Earth to form a dataset named XD IPIU for RSIs SR. Extensive experiments demonstrate the effectiveness of the proposed model and DRCAN shows the state-of-the-art performance in terms of quantitative evaluation and visual quality.
Xiangrong Zhang, Tianyang Zhang 0002, Fengsheng Liu, Xu Tang 0004, Puhua Chen, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 Foreground Refinement Network for Rotated Object Detection in Remote Sensing Images
abstract
Object detection has been a fundamental task in the field of remote sensing and has made considerable progress in recent years. However, the high background complexity in remote sensing images (RSIs) remains challenging. In this article, we propose a refined rotation detector, namely, the Foreground Refinement Network (FoRDet), to alleviate the above problem by leveraging the information of foreground regions from the perspectives of feature and optimization. Specifically, we propose a foreground relation module (FRL) that aggregates the foreground-contextual representations from the coarse stage and improves the discrimination of foreground regions on feature maps in the refined stage. Besides, considering the risk of the potential foreground anchors being overwhelmed in the training phase, we design a foreground anchor reweighting (FRW) loss that integrates the classification confidence and localization accuracy of each foreground anchor from the coarse stage to dynamically regulate their contributions in the refined stage, which highlights the potential foreground anchors. The comprehensive experimental results on three public datasets for rotated object detection DOTA, HRSC2016, and UCAS-AOD demonstrate the effectiveness of our proposed method.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Puhua Chen, Xu Tang 0004, Chen Li 0011, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2021 Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation has been continuously investigated in the last ten years, and majority of the established technologies are based on supervised models. In recent years, image-level weakly supervised semantic segmentation (WSSS), including single- and multi-stage process, has attracted large attention due to data labeling efficiency. In this paper, we propose to embed affinity learning of multi-stage approaches in a single-stage model. To be specific, we introduce an adaptive affinity loss to thoroughly learn the local pairwise affinity. As such, a deep neural network is used to deliver comprehensive semantic information in the training phase, whilst improving the performance of the final prediction module. On the other hand, considering the existence of errors in the pseudo labels, we propose a novel label reassign loss to mitigate over-fitting. Extensive experiments are conducted on the PASCAL VOC 2012 dataset to evaluate the effectiveness of our proposed approach that outperforms other standard single-stage methods and achieves comparable performance against several multi-stage methods.
Xiangrong Zhang, Zelin Peng, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Huiyu Zhou 0001, Licheng Jiao
ACM Multimedia4
2021 GRS-Det: An Anchor-Free Rotation Ship Detector Based on Gaussian-Mask in Remote Sensing Images
abstract
Ship detection is a significant and challenging task in remote sensing. Due to the arbitrary-oriented property and large aspect ratio of ships, most of the existing detectors adopt rotation boxes to represent ships. However, manual-designed rotation anchors are needed in these detectors, which causes multiplied computational cost and inaccurate box regression. To address the abovementioned problems, an anchor-free rotation ship detector, named GRS-Det, is proposed, which mainly consists of a feature extraction network with selective concatenation module (SCM), a rotation Gaussian-Mask model, and a fully convolutional network-based detection module. First, a U-shape network with SCM is used to extract multiscale feature maps. With the help of SCM, the channel unbalance problem between different-level features in feature fusion is solved. Then, a rotation Gaussian-Mask is designed to model the ship based on its geometry characteristics, which aims at solving the mislabeling problem of rotation bounding boxes. Meanwhile, the Gaussian-Mask leverages context information to strengthen the perception of ships. Finally, multiscale feature maps are fed to the detection module for classification and regression of each pixel. Our proposed method, evaluated on ship detection benchmarks, including HRSC2016 and DOTA Ship data sets, achieves state-of-the-art results.
Xiangrong Zhang, Guanchun Wang, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2020 Adaptive Feature Aggregation Network for Object Detection in Remote Sensing Images
abstract
Object detection in remote sensing images is a challenging task because of the large scale variations across the geospatial objects. The feature pyramid network (FPN) is widely used to alleviate the scale variations problem, however, it only fuses the features from adjacent levels and lacks the information of the entire feature hierarchy. In this paper, we propose a novel and effective feature pyramid aggregation network, called Adaptive Feature Aggregation Network (AFANet). Specifically, we propose the Adaptive Feature Aggregation (AFA) module to adaptively aggregate multi-level features of FPN and introduce the Bottom-up Path to enhance the location information of the entire feature levels. In addition, we use the Receptive Field Block (RFB) module to capture different receptive field features for each level feature map. We evaluate the effectiveness of our AFANet on the DOTA dataset and achieves noticeable performance compared with the baseline.
Wenliang Sun, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004
IGARSS3
2020 Discriminative Feature Pyramid Network For Object Detection In Remote Sensing Images
abstract
Multi-class geospatial object detection in remote sensing images suffer great challenges, such as large scales variability and complex background. Although feature pyramid network (FPN) can alleviate the problem of scale variation to some extent, it causes the loss of spatial and semantic information which is not conducive to object location. To address the above problem, this paper proposes a discriminative feature pyramid network (DFPN) by introducing a global guidance module (GGM) and a feature aggregation module (FAM). Specifically, the global guidance module delivers the high-level semantic information to lower layers, so as to obtain feature maps with stronger semantic information to eliminate the interference caused by complex background. The feature aggregation module enhances the interflow of information between different layers and better captures the discrimination information at each layer. We validate the effectiveness of our method on the NWPU VHR-10 and RSOD datasets, the results outperform baseline by 2.06 and 3.88 points respectively.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011
IJCNN3