Tong Zhang 0028

dblp:07/4227-28 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
19since 2021 · last 2025
0000-0002-1769-9829ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 19 since 2021
YearPublicationVenuePosition
2025 LLaMA-Unidetector: An LLaMA-Based Universal Framework for Open-Vocabulary Object Detection in Remote Sensing Imagery
abstract
Object detection is a crucial task in computer vision for remote sensing applications. However, the reliance of traditional methods on predefined and trained object categories limits their applicability in open-world scenarios. A key challenge in open-vocabulary object detection lies in accurately identifying unseen objects. Existing approaches often focus solely on detecting object locations, struggling to recognize the categories of previously unseen targets. To address this issue, we propose a novel benchmark where models are trained on known base classes and evaluated on their performance in detecting and recognizing unseen or novel classes. To this end, we introduce llama-Unidetector, a universal framework that incorporates textual information into a closed-set detector, enabling the generalization to open-set scenarios. Our llama-Unidetector leverages a decoupled learning strategy that separates localization and recognition. In the first stage, a class-agnostic detector identifies objects, distinguishing only between foreground and background. In the second stage, the detected foreground objects are passed through TerraOV-LLM, a multimodal large language model, for recognition, utilizing the strong generalization capabilities of large language models to infer the correct categories. We propose a self-built Vision Question Answering (VQA) remote sensing dataset, TerraVQA, and conduct extensive experiments on the NWPU-VHR10, DOTA1.0, and DIOR datasets. The llama-Unidetector achieves impressive results, with a performance of 75.46% AP, 50.22% AP and 51.38% AP on the zero-shot detection benchmarks for the NWPU-VHR10, DOTA1.0 and DIOR datasets, respectively. Our source code is available at: https://github.com/ChloeeGrace/LLaMA-Unidetector.
Jianlin Xie, Guanqun Wang, Tong Zhang 0028, Yikang Sun, He Chen 0004, Yin Zhuang, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.3
2025 EarthGPT-X: A Spatial MLLM for Multilevel Multisource Remote Sensing Imagery Understanding With Visual Prompting
abstract
Recent advances in natural-domain multi-modal large language models (MLLMs) have demonstrated effective spatial reasoning through visual and textual prompting. However, their direct transfer to remote sensing (RS) is hindered by heterogeneous sensing physics, diverse modalities, and unique spatial scales. Existing RS MLLMs are mainly limited to optical imagery and plain language interaction, preventing flexible and scalable real-world applications. In this article, EarthGPT-X is proposed, the first flexible spatial MLLM that unifies multi-source RS imagery comprehension and accomplishes both coarse-grained and fine-grained visual tasks under diverse visual prompts in a single framework. Distinct from prior models, EarthGPT-X introduces: 1) a dual-prompt mechanism combining text instructions with various visual prompts (i.e., point, box, and free-form) to mimic the versatility of referring in human life; 2) a comprehensive multi-source multi-level prompting dataset, the model advances beyond holistic image understanding to support hierarchical spatial reasoning, including scene-level understanding and fine-grained object attributes and relational analysis; 3) a cross-domain one-stage fusion training strategy, enabling efficient and consistent alignment across modalities and tasks. Extensive experiments demonstrate that EarthGPT-X substantially outperforms prior natural and RS MLLMs, establishing the first framework capable of multi-source, multi-task, and multi-level interpretation using visual prompting in RS scenarios. The code and dataset are available athttps://github.com/wivizhang/EarthGPT-X.
Wei Zhang 0389, Miaoxin Cai, Yaqian Ning, Tong Zhang 0028, Yin Zhuang, Shijian Lu, He Chen 0004, Jun Li 0009, Xuerui Mao
IEEE Trans. Geosci. Remote. Sens.4
2025 EarthMarker: A Visual Prompting Multimodal Large Language Model for Remote Sensing
abstract
Recent advances in prompt learning have allowed users to interact with artificial intelligence (AI) tools in multiturn dialog, enabling an interactive understanding of images. However, it is difficult and inefficient to deliver information in complicated remote sensing (RS) scenarios using plain language instructions alone, which would severely hinder deep comprehension of the latent content in imagery. Besides, existing prompting strategies in natural scenes are hard to apply to interpret the RS data due to significant domain differences. To address these challenges, the first visual prompting-based multimodal large language model (MLLM) named EarthMarker is proposed in the RS domain. EarthMarker is capable of interpreting RS imagery at the image, region, and point levels by levering visual prompts (i.e., boxes and points). Specifically, a shared visual encoding method is developed to establish the spatial pattern interpretation relationships between the multiscale representations of input images and various visual prompts. Subsequently, the mixed visual-spatial representations are associated with language instructions to construct joint prompts, enabling the interpretation of intricate content of RS imagery. Furthermore, to bridge the domain gap between natural and RS data, and effectively transfer domain-level knowledge from natural scenes to the RS domain, a cross-domain learning strategy is developed to facilitate the RS imagery understanding. In addition, to tackle the lack of RS visual prompting data, a dataset named RSVP featuring multimodal multigranularity visual prompts instruction-following is constructed. Extensive experiments are conducted to demonstrate the competitive performance of the EarthMarker. The proposed EarthMarker represents a significant advance in multigranularity RS imagery interpretation under the visual prompting learning framework. Our code and dataset are available athttps://github.com/wivizhang/EarthMarker.
Wei Zhang 0389, Miaoxin Cai, Tong Zhang 0028, Yin Zhuang, Jun Li 0009, Xuerui Mao
IEEE Trans. Geosci. Remote. Sens.3
2025 A Unified Remote Sensing Object Detector Based on Fourier Contour Parametric Learning
abstract
A unified object detector needs to integrate various abilities for adapting to different remote sensing object detection tasks. However, there is a lack of a feasible way to integrate multigrained object detection requirements i.e., horizontal bounding box (HBB), oriented bounding box (OBB), and instance segmentation (InSeg) into a unified detection way. Then, it often has to design specific parametric learning ways and their corresponding architectures, which cannot be finely adaptive to various kinds of object detection tasks. Therefore, in this article, a new benchmark is set up to integrate multigrained object detection requirements of HBB, OBB, and InSeg into one challenging task of arbitrary-shaped object contour detection. At the same time, a unified object contour detector (UniconDet) is proposed for achieving multigrained object detection from complicated remote sensing scenes. First, a Fourier contour parametric modeling (FCPM) is defined to project arbitrary-shaped object contours from the spatial domain into the frequency domain. Then, it can unify spatial parametric representations of HBB, OBB, and InSeg as frequency coefficient representations, which can be used for realizing a more generic and robust parametric regression. Second, a multiview cross-attention (MVCA) feature extraction way is designed at each scale of the regression layer, which can assist UniconDet in perceiving Fourier contour parameters by exploring the coupled relations between different discrete contour sampling periods of each object. Third, a center-contour enhancing regression layer (C2-ERL) is designed to generate regional guidance and cascade contour propagation, which can ensure a more accurate center point prediction and Fourier contour parameter regression. Finally, extensive experiments are carried out on benchmarks of HBB, OBB, InSeg, and new multigrained object detection, and the results indicate that our proposed UniconDet can obtain superior performance. The source code is available athttps://github.com/ZhAnGToNG1/UniconDet.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, LianLin Li, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.1
2025 Controllable Generative Knowledge-Driven Few-Shot Object Detection From Optical Remote Sensing Imagery
abstract
Few-shot object detection (FSOD) has to learn classification and localization information for unseen object detection under very low-data resource regimes. However, when deficient samples are adopted for model training, it is hard to build powerful location-aware and identification abilities for well coping with agnostic bias from diverse testing scenarios; at the same time, the overfitting phenomenon is easily occurring. Therefore, in this article, a controllable generative knowledge-driven FSOD called CGK-FSOD is proposed for unseen object detection from optical remote sensing imagery. Specifically, to enrich the learnable data space of scarce samples for preventing incomplete agnostic-bias learning, while avoiding the overfitting phenomenon, a visual-textual prompt-based controllable data generation is designed to generate high-quality object detection data based on pretrained foundational models [i.e., the stable diffusion (SD) and contrastive language-image pre-training (CLIP)], which not only can introduce the generalized domain-level knowledge into the remote sensing domain but also sets up an all-round data space to support complete learning of potential agnostic bias. Furthermore, with respect to the denoising generative process of SD, a series of cross-modality generative features in latent representation space are reused for few-shot fine-tuning by the designed cross-modality feature embedding (CMFE), which not only can bring diverse generative abilities into the feature fusion step of the detector but also gracefully sets up feature representation scalability to make the detector better adapt to agnostic bias from diverse testing scenarios of FSOD. Finally, extensive experiments are executed on two public remote sensing datasets (e.g., DIOR and NWPUVHR-10), and the results indicate that the proposed CGK-FSOD is very effective and flexible for FSOD.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, LianLin Li, Jun Li 0009
IEEE Trans. Geosci. Remote. Sens.1
2024 Diffusion-Geo: A Two-Stage Controllable Text-To-Image Generative Model for Remote Sensing Scenarios
abstract
Image generation is a crucial task to facilitate intelligent interpretation in remote sensing domain. Expanding dataset size through image generation can enhance model performance of downtown task. However, current generative models in remote sensing are mostly unconditional or guided by simple text, resulting in generated images lacking spatial and semantic constraints. This lack of control can negatively optimize downstream task models. To tackle these challenges, a two-stage controllable text-image generative model called Diffusion-Geo is presented. In the first stage, an extensive image-text generation dataset called RS-Control is created through prompt engineering of multimodal large language models (MLLMs) and manual prompts for existing datasets, incorporates diverse conditional controls with rich spatial and semantic information. Then RS-Control dataset is utilized to train a universal controllable image generative model. The second stage involves efficient tuning the universal model for different task datasets, minimizing fine-tuning costs while preserving diversity and high-quality features. Experiments conducted on the RSICD caption dataset and WHU change detection dataset demonstrate the superiority of Diffusion-Geo over other state-of-the-art models in image generation.
Miaoxin Cai, Wei Zhang 0389, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Liang Chen 0004, Can Li 0005
IGARSS3
2024 Regression-Guided Positive Sample Refocusing Paradigm for Tiny Object Detection in Aerial Images
abstract
Tiny object detection represents a pivotal challenge in remote sensing intelligent interpretation, necessitating detectors to exhibit heightened precision in object localization. However, typical model optimization strategies cannot release the detector’s potential for precisely localizing objects. And the lack of interpretability in detection box filtering based on object classification scores serves as a constraint on further performance improvement. Therefore, this paper proposed a novel model optimization strategy to thoroughly unleash the potential of the detector for precise localization. Then, the utilization of object comprehensive confidence score enhances the interpretability of the post-processing step for detection boxes. Rigorous experiments on the AI-TOD dataset have demonstrated the effectiveness of our method, achieving state-of-the-art performance.
Lihui Ge, He Chen 0004, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, Fukun Bi, Liang Chen 0004
IGARSS4
2024 Advancing Controllable Diffusion Model for Few-Shot Object Detection in Optical Remote Sensing Imagery
abstract
Few-shot object detection (FSOD) from optical remote sensing imagery has to detect rare objects given only a few annotated bounding boxes. The limited training data is hard to represent the data distribution of realistic remote sensing scenes, restricting the performance of FSOD. Recently, learning conditional controls for text-to-image diffusion model has achieved great progress, which is capable of precisely generating the controllable yet imaginational images by text prompt and spatially localized input conditions. Accordingly, in this work, we aim to explore the potential of diffusion model and propose a solution for few-shot object detection by controllable data generation. Firstly, draw upon a few annotated objects, their bounding boxes and categories are respectively used as the spatial conditions and text prompts, then employ them into large text-to-image diffusion models for controlled image generation. Secondly, based the generated images, in order to adapt to the scale and orientation variances of remote sensing objects, a data transformation is devised for boosting the robustness of model training. Finally, some experiments were conducted on public remote sensing dataset DIOR, and the results proved its effectiveness.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004, Fukun Bi
IGARSS1
2024 Regression-Guided Refocusing Learning With Feature Alignment for Remote Sensing Tiny Object Detection
abstract
Tiny object detection is a formidable challenge in remote sensing intelligent interpretation. Tiny objects are usually fuzzy, densely distributed and highly sensitive to positioning errors, which leads to the mainstream detector usually achieving suboptimal detection performance when facing tiny objects. To address the mismatch of mainstream detector architectures and model optimization strategies in the context of tiny object detection, this paper presents an efficient and interpretable algorithm for tiny object detection, termed the Cross-Attention based Feature Fusion Enhanced tiny object detection Network (CAF2ENet). First, the cross-attention mechanism is introduced to refine the upsampling results of deep features. This refinement improves the precision of multi-scale feature fusion. Second, a training strategy named regression-based refocusing learning is introduced. Deviating from the conventional optimization strategy, our method guides the optimizer to prioritize higher-quality detection boxes by adjusting sample weights. This adjustment significantly amplifies the detector’s potential to achieve superior detection results. Finally, the object composite confidence score is employed for the interpretable filtering of detection boxes. Extensive experiments on Tiny Object Detection in Aerial Images (AI-TOD) and object Detection in Optical Remote sensing images (DIOR) datasets are carried out, and comparison indicate that the proposed CAF2ENet can perform the remarkable performance compared to other state-of-the-art (SOTA) tiny object detection detectors, as it can reach 63.7% Average Precision (AP50) on AI-TOD and 75.4%AP50on DIOR, achieve SOTA performance.
Lihui Ge, Guanqun Wang, Tong Zhang 0028, Yin Zhuang, He Chen 0004, Hao Dong 0003, Liang Chen 0004
IEEE Trans. Geosci. Remote. Sens.3
2024 DECOR: Dynamic Decoupling and Multiobjective Optimization for Long-Tailed Remote Sensing Image Classification
abstract
In the realm of remote sensing, targets of interest span a range of categories. However, their distribution is not always uniform. Certain categories substantially outnumber others, resulting in what’s termed a ‘long-tailed distribution’ in remote sensing imagery. This imbalanced distribution often biases a classifier’s focus toward the more abundant (head) classes, at the detriment of the less-represented (tail) classes. Such biases undermine the classifier’s generalization performance, particularly in the context of remote sensing image classification (RSIC). While existing mitigation approaches such as resampling, reweighting, and transfer learning offer some respite, they often miss out on in-depth knowledge refinement, rendering them less effective for severe long-tailed RSIC scenarios. To counter these challenges, we introduce DECOR, a dynamic decoupling and multi-objective optimization framework. Within DECOR, the feature extractor and classifier are dynamically decoupled, promoting superior feature representation and classifier training. Then, a multi-objective optimization approach is proposed to delve deeper, refining feature representation at the knowledge level using learnable feature centroids coupled with masked world knowledge learning. Moreover, to combat the pronounced effects of sample imbalance on classifier training, we employ a class-balanced re-sampling technique paired with a parameter-efficient adapter, which sharpens the classifier’s decision boundary and bridges the gap between representation and classification. DECOR’s efficacy is validated through comprehensive experiments on several datasets, including the NWPU-RESISC45-LT (NWPU-LT), AID-LT, and our self-built BIT-AFGR50-LT. Experimental results demonstrate DECOR’s marked enhancement in performance on long-tailed datasets. Our source code is available at: https://github.com/ChloeeGrace/DECOR.
Jianlin Xie, Guanqun Wang, Yin Zhuang, Can Li 0005, Tong Zhang 0028, He Chen 0004, Liang Chen 0004, Shanghang Zhang
IEEE Trans. Geosci. Remote. Sens.5
2024 EarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain
abstract
Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant domain gap between natural and remote sensing (RS) images, the development of MLLMs in the RS domain is still in the infant stage. To fill the gap, a pioneer MLLM named EarthGPT integrating various multi-sensor RS interpretation tasks uniformly is proposed in this paper for universal RS image comprehension. Firstly, a visual-enhanced perception mechanism is constructed to refine and incorporate coarse-scale semantic perception information and fine-scale detailed perception information. Secondly, a cross-modal mutual comprehension approach is proposed, aiming at enhancing the interplay between visual perception and language comprehension and deepening the comprehension of both visual and language content. Finally, a unified instruction tuning method for multi-sensor multi-task in the RS domain is proposed to unify a wide range of tasks including scene classification, image captioning, region-level captioning, visual question answering (VQA), visual grounding, object detection, etc. More importantly, a dataset named MMRS-1M featuring large-scale multi-sensor multi-modal RS instruction-following is constructed, comprising over 1M image-text pairs based on 34 existing diverse RS datasets and including multi-sensor images such as optical, synthetic aperture radar (SAR), and infrared. The MMRS-1M dataset addresses the drawback of MLLMs on RS expert knowledge and stimulates the development of MLLMs in the RS domain. Extensive experiments are conducted, demonstrating the EarthGPT’s superior performance in various RS visual interpretation tasks compared with the other specialist models and MLLMs, proving the effectiveness of the proposed EarthGPT and offering a versatile paradigm for open-set reasoning tasks. Our code and dataset are available at https://github.com/wivizhang/EarthGPT.
Wei Zhang 0389, Miaoxin Cai, Tong Zhang 0028, Yin Zhuang, Xuerui Mao
IEEE Trans. Geosci. Remote. Sens.3
2024 Heterogeneous Prototype Distillation With Support-Query Correlative Guidance for Few-Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification (FSRSSC) aims to identify unseen classes only relying on very limited training samples. However, scarce training samples are insufficient to support a robust classwise representation, which is easily influenced by agnostic biases from diverse testing scenarios. Fortunately, there are abundant spatial contextual clues that exist in very limited training samples to have an enormous potential to establish discriminative and transferable concepts. Thus, in this article, a hybrid architecture called ProtoConViT is proposed to learn a powerful classwise representation based on spatial contextual clues for FSRSSC promotion. First, support-query correlative guidance is designed to generate more stable spatial connections among support and query data based on intermediate convolution neural network (CNN) feature maps, which not only can be embedded into each episodic training task to reduce redundant spatial contextual representation learning space of vision transformer (ViT) but also can assist it in rapidly capturing critical spatial contextual clues to classify query data into one of classes from support set. Second, followed by the designed support-query correlative guidance, a novel heterogeneous prototype distillation is proposed to integrate the advantages of CNN and ViT for heterogeneous prototype construction, which can rapidly set up discriminative and transferable concepts for FSRSSC. Third, corresponding to the proposed ProtoConViT, a joint loss is designed to make the model rapid convergence based on meta-learning. Finally, extensive experiments are carried out on three FSRSSC benchmarks, and comparative results indicate that the proposed ProtoConViT can achieve a superior FSRSSC performance.
Yin Zhuang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li
IEEE Trans. Geosci. Remote. Sens.3
2023 Multi-Grained Global-Local Semantic Feature Fusion for Few Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification aims to classify unseen scenes by using only a few labeled samples. Hence, how to set up a more effective feature description according to a few labeled samples, becomes an important issue. In this paper, in view of more complicated remote sensing scenes containing several hierarchical and coupled spatial relations (e.g., internal and external spatial contexts), which severely hinder the feature extraction under few-shot learning scenarios, a multi-grained global-local semantic feature fusion (MGGL-SFF) method is proposed for few-shot remote sensing scene classification, which can better combine the global discriminative spatial semantic features with local transferable fragment features to set a powerful prototype representation up for few shot learning. Finally, experiments are carried out on defined few-shot remote sensing scene classification benchmark, and results proved the proposed MGGL-SFF can achieve a new state-of-the-art performance.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, He Chen 0004
IGARSS2
2023 Contour Modeling Arbitrary-Oriented Ship Detection From Very High-Resolution Optical Remote Sensing Imagery
abstract
Under the multiscale distribution, due to dramatic aspect ratio variance leading to prominent arbitrary-oriented character of ships in very high-resolution (VHR) optical remote sensing imagery, how to generate accurate oriented bounding box (OBB) becomes a hot research topic for arbitrary-oriented ship detection. Consequently, in this letter, a concise and effective one-stage anchor-free contour modeling detector called CMDet is proposed for accurate arbitrary-oriented ship detection. Different from currently existed methods via carefully decoupling several independently characteristic parameters for OBB modeling and regression, we resolve the OBB modeling by jointly regressing the contour information. Specifically, the contour information is expressed as a series of Fourier transform coefficients, which are generated by setting up the mapping relation of 1-D Fourier contour coefficients and spatial OBB contour. In addition, a new inherent geometry loss is designed to make detector better learn the geometry information in training phase. After that, the proposed CMDet only needs to predict the correct center point of ships and regress the corresponding entire 1-D Fourier contour coefficients to generate accurate OBB for ship detection. Finally, extensive experiments are carried out on two public OBB ship detection datasets (e.g., HRSC2016 and DIOR-ship), and comparison results demonstrate that the proposed CMDet can obtain the competitive result than the state-of-the-art (SOTA) detectors.
Yin Zhuang, Yuqun Liu, Tong Zhang 0028, He Chen 0004
IEEE Geosci. Remote. Sens. Lett.3
2023 Full Semantic Constructed Network for Urban Use Classification From Very High-Resolution Optical Remote Sensing Imagery
abstract
Recently, semantic segmentation technology has been a research hotspot in optical remote sensing urban use classification. However, because of coupled semantic relations in very high-resolution and complex urban scenes, a more effective semantic description for pixelwise urban use interpretation has become a challenge. Then, aiming to set up a more effective semantic description, the effective receptive field (ERF) is analyzed in general convolutional neural networks. The unreasonable ERF distribution in the stacked convolutional layers of the encoder would lead to a large amound of small ERFs and fewer not large enough ERFs that form a naive semantic description in decoder. Therefore, in this article, a novel full semantic constructed network (FSCNet) is proposed to improve the naive semantic description and set up an effective semantic description. First, to avoid noise from shallow feature layers, a residual refinement convolution is designed to optimize the full-scale skip connections based on the U-shaped encoder–decoder. Second, an interscale fusion module is newly designed for multiscale feature fusion, which can generate three initial semantic modalities that are prepared for redefining the full semantic description. Third, a multiscale local context spatial attention module and boundary supervision are designed for an initial shallow semantic modality to capture the pure boundary information, and then, pyramid spatial pooling is employed for an initial deep semantic modality to further enlarge the ERF and obtain more abstract global information. Next, a self-calibration convolution combined with the atrous spatial pyramid pooling is designed to rectify and enrich an initial middle semantic modality, which can improve the naive semantic description and bridge the semantic gap between the redefined shallow and deep semantic modalities to advance the full semantic feature fusion. Finally, extensive experiments are carried out on three benchmarks (e.g., ISPRS Vaihingen, Potsdam, and DLRSD), and comparative results show that the proposed FSCNet can get remarkable performance compared to state-of-the-art (SOTA) methods. Besides, the code is available athttps://github.com/DorisCV/FSCNet.
Shan Dong, Yin Zhuang, He Chen 0004, Tong Zhang 0028, LianLin Li
IEEE Trans. Geosci. Remote. Sens.4
2023 Posterior Instance Injection Detector for Arbitrary-Oriented Object Detection From Optical Remote-Sensing Imagery
abstract
Arbitrary-oriented object detection (AOOD) from optical remote sensing imagery has to correctly generate delicate oriented boundary boxes (OBBs) and meanwhile identify their specific categories. However, how to make detectors learn delicate parameters of OBBs, especially for the crucial orientation information, and identify object category from complex background becomes a challenge task. Therefore, in this article, for exploring a better way to guide the detector to learn specific category and parametric information of OBBs, a novel one-stage anchor-free detector called Posterior Instance Injection Detector (PIIDet) is proposed for AOOD. First, as the anchor-free manner lacks prior information, an object-aware posterior guidance (OAPG) structure is proposed to generate specific-category instances used for conditioning on OBB prediction. This structure can assist the proposed PIIDet in better learning the relative parametric information of OBBs corresponding to their specific categories. Besides, to guarantee a high quality injection of specific-category instances, a new hierarchical feature fusion module is developed to establish a suitable multi-scale feature mapping space. Second, considering the negative optimization of angle regression, which is caused by the boundary discontinuity of angular periods and sudden shifts of the relation between width and height in training phase, a novel binary classification embedded angle regression space (BCE-RegSpace) is devised for providing continuous angle regression space and stable relation between width and height. Finally, extensive experiments are executed on three AOOD benchmarks (e.g., DOTA, DIOR-R and HRSC2016), and results proved that the proposed concise one-stage anchor-free PIIDet can reach the state-of-the-art (SOTA) performance and meanwhile have an impressive inference speed.
Tong Zhang 0028, Yin Zhuang, He Chen 0004, Guanqun Wang, Lihui Ge, Liang Chen 0004, Hao Dong 0003, LianLin Li
IEEE Trans. Geosci. Remote. Sens.1
2022 Adaptive Local Context Embedding for Small Vehicle Detection from Aerial Optical Remote Sensing Images
abstract
Small vehicle detection is one of the remaining challenging task because the ambiguous appearance is against complex background interference. Consequently, in order to improve the performance of small vehicle detection from aerial optical remote sensing images, a novel adaptive local context (ALC) embedding way is designed and further introduced into an anchor free detection manner which is called ALC-Net, and in ALC-Net, it can adaptively set up the effective local context feature to improve keypoint description of small vehicles and boost the detection performance without adding extra prior information. Finally, several experiments are carried out on two widely used datasets (e.g., UCAS-AOD [1] and VEDAI [2]) and the results indicate that the proposed ALC-Net can exhibit the competitive small vehicle detection performance than other detectors.
Shanjunyu Liu, Yin Zhuang, Hao Dong 0003, Peng Gao 0007, Guanqun Wang, Tong Zhang 0028, Liang Chen 0004, He Chen 0004, LianLin Li
IGARSS6
2022 FSoD-Net: Full-Scale Object Detection From Optical Remote Sensing Imagery
abstract
Object detection is an essential task in computer vision. Recently, several convolution neural network (CNN)-based detectors have achieved a great success in natural scenes. However, for optical remote sensing images with a large scale of view, lower proportion of foreground target pixels and drastic differences in object scale present considerable challenges. To address these problems, we propose a novel one-stage detector called the full-scale object detection network (FSoD-Net) which consists of proposed multiscale enhancement network (MSE-Net) backbone cascaded with scale-invariant regression layers (SIRLs). First, MSE-Net provides the multiscale description enhancement by integrated the Laplace kernel with fewer parallel multiscale convolution layers. Second, SIRLs contain three different isolated regression branch layers (i.e., corresponding to small, medium, and large scales), which make default discrete scale bounding boxes (bboxes) cover full-scale object information in regression procedure. A novel specific scale joint loss is also designed that uses the softmax function combined with a strong$L_{1}$-norm constraint in each regression branch layer. It can further speed up the convergence and improve the classification scores of predicted bboxes. Finally, extensive experiments are carried on challenge data sets of large-scale dataset for object detection in aerial images (DOTA) and object detection in optical remote sensing images (DIOR) which contain multiple instances from different imaging platforms, and these results demonstrate that FSoD-Net can achieve better performance than other state-of-the-art one-stage detectors, and it can reach a mean average precision (mAP) of 75.33% on DOTA and 71.80% mAP on DIOR, respectively. Especially, the average precision (AP) of tiny object detection can improve 10%–20% approximately.
Guanqun Wang, Yin Zhuang, He Chen 0004, Tong Zhang 0028, LianLin Li, Shan Dong, Qianbo Sang
IEEE Trans. Geosci. Remote. Sens.5
2022 Multiscale Semantic Fusion-Guided Fractal Convolutional Object Detection Network for Optical Remote Sensing Imagery
abstract
Optical remote sensing object detection is a challenging task, because of the complex background interference, ambiguous appearances of tiny objects, densely arranged circumstances, and multiclass object with vaster scale variances and irregular aspect ratios. The performance of object detection is seriously restricted. Thus, in this article, inspired by the anchor-free object detection framework, and aiming to solve these difficulties to improve the optical remote sensing object detection performance, a powerful one-stage detector of multiscale semantic fusion-guided fractal convolution network (MSFC-Net) is proposed. First, facing these strong-coupled semantic relations in each complex scene, a compound semantic feature fusion (CSFF) way is designed for generating an effective semantic description, which is a benefit to pixel-wise object center point interpretation. In addition, it can be easily extended into a semantic segmentation task. Second, in view of accurate multiclass pixel-wise center point predictions based on an effective compound semantic description, a novel fractal convolution (FC) regression layer is designed, which adaptively achieves the regression of multiscale bounding boxes (bboxes) with irregular aspect ratio under no priori information. Third, related to the set up FC regression layer, a specific hybrid loss is designed to make the proposed MSFC-Net converge better. Finally, the extensive experiments on challenge data sets of large-scale dataset for object detection in aerial images (DOTA) and object detection in optical remote sensing images (DIOR) datasets are carried out, and comparisons indicate that the proposed MSFC-Net can perform the remarkable performance than other state-of-the-art one-stage detectors, as it can reach 80.26% mean average precision (mAP) and 79.33% mF1 on DOTA and 70.08% mAP and 73.45% mF1 on DIOR. Then, our work is available athttps://github.com/ZhAnGToNG1/MSFC-Net.
Tong Zhang 0028, Yin Zhuang, Guanqun Wang, Shan Dong, He Chen 0004, LianLin Li
IEEE Trans. Geosci. Remote. Sens.1
2020 Feature Enhanced Centernet for Object Detection in Remote Sensing Images
abstract
Multi-scale object detection in optical remote sensing imagery is a challenging task due to the varied object scales. Existed state-of-art object detection methods have achieved significant growth. However, most of the methods are based on default anchors, which need to be predefined. The multi-scale object detection accuracy still needs to be improved, especially for small and dense objects. To improve the robustness of the detection algorithm and the performance of multi-scale object detection, a novel anchor-free multi-scale object detection method Feature Enhanced CenterNet is proposed in this paper. First, we use the “encoder-decoder” structure and introduce horizontal connections to enhance feature representation capabilities. Second, an context-aware up-sampling method is proposed to obtain feature maps with suitable scale. To demonstrate the performance of the proposed method, we perform abundant experiments on the public remote sensing datasets. The experimental results demonstrate the robustness and effectiveness of the proposed method.
Tong Zhang 0028, Guanqun Wang, Yin Zhuang, He Chen 0004, Hao Shi 0006, Liang Chen 0004
IGARSS1