Caiguang Zhang

dblp:268/1360 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-0321-9900ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 EOOD: End-to-end oriented object detection
Caiguang Zhang, Zilong Chen, Boli Xiong, Kefeng Ji, Gangyao Kuang
Neurocomputing1
2023 TCD: Task-Collaborated Detector for Oriented Objects in Remote Sensing Images
abstract
Oriented object detection (OOD) in remote sensing image interpretation is challenging due to the difficulty of locating objects with arbitrary orientations. Existing methods have made considerable progress based on oriented heads or anchors. However, most of them follow the classical detection paradigm, such as assigning samples based on Intersection-over-Unions (IoU) and predicting through two independent tasks. These fixed strategies impair the consistency between classification and localization predictions, resulting in the prediction with optimal localization accuracy being suppressed by the nonoptimal ones during nonmaximum suppression (NMS). To address this problem, a task-collaborated detector (TCD) is proposed. Compared with current single-stage methods, its improvements include two aspects: task-collaborated assignment (TCA) and task-collaborated head (TCH). Specifically, to better pull closer the best anchors for two tasks, TCA introduces classification and localization confidence into sample assignment and tends to select the anchors with accurate and consistent predictions as positive during training. TCH provides a better balance for learning interactive and discriminative features. It can flexibly adjust the spatial feature distribution of classification and localization tasks by learning the joint features from the aggregation layer. Extensive experiments are conducted on HRSC2016, DOTA, and DIOR-R, and the proposed TCD achieves the state-of-the-art performance [90.60, 80.89, and 65.04 mean average precision (mAP), respectively]. Consistency analysis also demonstrates that TCD can significantly improve prediction consistency.
Caiguang Zhang, Boli Xiong, Xiao Li 0017, Gangyao Kuang
IEEE Trans. Geosci. Remote. Sens.1
2022 Aspect-Ratio-Guided Detection for Oriented Objects in Remote Sensing Images
abstract
Although existing oriented object detection methods have made considerable progress based on oriented heads or anchors, the training process itself is not perfect. In this letter, we point out the inconsistency problem between the fixed network setting and varying aspect ratios, which greatly limits the performance. For example, the fixed parameters in label assignment and regression loss cannot fit the changes of aspect ratios and, thus, are harmful to the training process. Considering the prior information about objects’ aspect ratios, the aspect-ratio-guided (ARG) methods are proposed. Specifically, the ARG label assignment is used to adjust the label assignment criteria (intersection over union (IoU) threshold) automatically, and the ARG IoU loss can change the weights of angle regression dynamically. This ARG design makes better use of training samples and pushes the detector more robust to the change of aspect ratios. With no additional cost, our method improves upon the ResNet-50-feature pyramid network (FPN) baseline with 3.99% AP50 and 6.09% AP75 on HRSC2016.
Caiguang Zhang, Boli Xiong, Xiao Li 0017, Gangyao Kuang
IEEE Geosci. Remote. Sens. Lett.1
2022 Dense Adaptive Grouping Distillation Network for Multimodal Land Cover Classification With Privileged Modality
abstract
Multimodal land cover classification (MLCC) is a fundamental problem in remote sensing interpretation, which can obtain excellent performance on account of the complementary information between the optical and SAR modalities. However, it is usually impossible to obtain multimodal data at the same time, due to the restriction of imaging conditions. When one of the modalities data is completely missing during test phase, classical multimodal learning methods might not be able to handle the MLCC task with privileged modality. In this paper, we propose an efficient Dense Adaptive Grouping Distillation Network (DAGDNet), which learns privileged information from available modalities in the train sets, and improves the classification performance in the test sets when one modality data is scarce. More specifically, to relieve the heterogeneous gaps between different modalities and then transfer the privileged information, we propose an Interactive Gated-based Feature Grouping Module (IG-FGM), which decomposes multimodal features into modalities-shared and modality-specific components to realize the decoupling of multimodal features and grouping distillation. Furthermore, the IG-FGM is inserted into different layers of the “teacher" network to implement progressive blending of multi-modalities. Then, to adaptively highlight the importance of hierarchical features distillation and grouping distillation, we propose a Multi-stage Adaptive Distillation Learning (MS-ADL) strategy so that the weights of different distillation losses are required to change continuously along with the training process. Finally, we evaluate the superior performances of our model on representative co-registered optical and SAR datasets.
Xiao Li 0017, Lin Lei, Caiguang Zhang, Gangyao Kuang
IEEE Trans. Geosci. Remote. Sens.3
2022 Multimodal Semantic Consistency-Based Fusion Architecture Search for Land Cover Classification
abstract
Multimodal Land Cover Classification (MLCC) using the optical and Synthetic Aperture Radar (SAR) modalities has resulted in outstanding performances over using only unimodal data due to their complementary information on land properties. Previous multimodal deep learning (MDL) methods have relied on handcrafted multi-branch convolutional neural networks (CNN) to extract the features of different modalities and merged them for land cover classification. However, natural images-oriented handcrafted CNN models may not the optimal strategies to handle Remote Sensing (RS) image interpretation problems, due to the huge difference in terms of imaging angles and imaging ways. Furthermore, few MDL methods have analyzed optimal combinations of hierarchical features from different modalities. In this article, we propose an efficient multimodal architecture search framework, namely Multimodal Semantic Consistency-Based Fusion Architecture Search (M2SC-FAS) in continuous search space with the gradient-based optimization method, which can not only discover optimal optical- and SAR-specific architectures according to the different characteristics of the optical and SAR images, respectively, but also realizes the search of optimal multimodal dense fusion architecture. Specifically, the semantic-consistency constraint is introduced to guarantee dense fusion between hierarchical optical and SAR features with high semantic consistency and then capture the complementary performance on land properties. Finally, the basis of curriculum learning strategy is adopted on the M2SC-FAS. Extensive experiments show superior performances of our work on three broad co-registered optical and SAR datasets.
Xiao Li 0017, Lin Lei, Caiguang Zhang, Gangyao Kuang
IEEE Trans. Geosci. Remote. Sens.3
2021 Ship Detection and Recognition in Optical Remote Sensing Images Based on Scale Enhancement Rotating Cascade R-CNN Networks
abstract
Ship detection and recognition in remote sensing images have important significance in military and civilian applications. Traditional methods have insufficient generalization ability in complicated scenes. The Faster R-CNN-based methods cannot predict the orientation of the ship. The R2CNN-based methods can predict the orientation of the ship but not considerate the scale of object in classification. In order to solve the problems mentioned above, this article proposes a scale enhancement rotating Cascade R-CNN network (SER-Cascade). Using the multistage network of rotating Cascade R-CNN, the output of the previous stage is fed to current stage, which can effectively regress the orientation of the ship. To improve the recognition performance of multi-class ships, a novel RoI pooling method is proposed in this article, in which the scale information is enhanced and context information is reserved. To evaluate the proposed networks, a dataset named HR-SHIP-15 that currently contains 15 categories of ship targets has been produced for ship recognition. Experiments are conducted on HR-SHIP-15 dataset, and the results verify that the proposed method has state-of-the-art performance.
Caiguang Zhang, Boli Xiong, Gangyao Kuang
IGARSS1