Yue Gao 0008

dblp:33/3099-8 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-5020-586XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 PETDet: Proposal Enhancement for Two-Stage Fine-Grained Object Detection
abstract
Fine-grained object detection (FGOD) extends object detection with the capability of fine-grained recognition. In recent two-stage FGOD methods, the region proposal serves as a crucial link between detection and fine-grained recognition. However, current methods overlook that some proposal-related procedures inherited from general detection are not equally suitable for FGOD, limiting the multitask learning from generation, representation, to utilization. In this article, we present a proposal enhancement for two-stage FGOD (PETDet) to better handle the subtasks in two-stage FGOD methods. First, an anchor-free quality-oriented proposal network (QOPN) is proposed with dynamic label assignment and attention-based decomposition to generate high-quality-oriented proposals. In addition, we present a bilinear channel fusion network (BCFN) to extract independent and discriminative features of the proposals. Furthermore, we designed a novel adaptive recognition loss (ARL) that offers guidance for the region-based convolutional neural networks (R-CNNs) head to focus on high-quality proposals. Extensive experiments validate the effectiveness of PETDet. Quantitative analysis reveals that PETDet with ResNet50 reaches state-of-the-art performance on various FGOD datasets, including FAIR1M-v1.0 (42.96 AP), FAIR1M-v2.0 (48.81 AP), MAR20 (85.91 AP), and ShipRSImageNet (74.90 AP). The proposed method also achieves superior compatibility between accuracy and inference speed. Our code and models will be released athttps://github.com/canoe-Z/PETDet.
Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Panoptic Perception: A Novel Task and Fine-Grained Dataset for Universal Remote Sensing Image Interpretation
abstract
Current remote-sensing interpretation models often focus on a single task such as detection, segmentation, or caption. However, the task-specific designed models are unattainable to achieve the comprehensive multi-level interpretation of images. The field also lacks support for multi-task joint interpretation datasets. In this paper, we propose Panoptic Perception: a novel task and a new fine-grained dataset (FineGrip) to achieve a more thorough and universal interpretation for RSIs. The new task: 1) integrates pixel-level, instance-level, and image-level information for universal image perception, 2) captures image information from coarse to fine granularity, achieving deeper scene understanding and description, and 3) enables various independent tasks to complement and enhance each other through multi-task learning. By emphasizing multi-task interactions and the consistency of perception results, this task enables the simultaneous processing of fine-grained foreground instance segmentation, background semantic segmentation, and global fine-grained image captioning. Concretely, the FineGrip dataset includes 2,649 remote sensing images, 12054 fine-grained instance segmentation masks belonging to 20 foreground things categories, and 7599 background semantic masks for 5 stuff classes. Furthermore, we propose a joint optimization-based panoptic perception model. Experimental results on FineGrip demonstrate the feasibility of the panoptic perception task and the beneficial effect of multi-task joint optimization on individual tasks. The dataset will be publicly available.
Danpei Zhao, Bo Yuan 0009, Tian Li 0009, Zhuoran Liu 0006, Yue Gao 0008
IEEE Trans. Geosci. Remote. Sens.7
2023 Classification Matters More: Global Instance Contrast for Fine-Grained SAR Aircraft Detection
abstract
Since significant intraclass differences and inconspicuous interclass variations, fine-grained aircraft detection in synthetic aperture radar (SAR) images is challenging. Also, the inherent lack of detailed features and severe noise interference in SAR images make it difficult to learn class-specific feature representations. Current detection approaches focus more on localization accuracy and ignore classification performance, which is more critical in fine-grained detection. To address the above challenges, we present GICNet: global instance contrast (GIC) for fine-grained SAR aircraft detection a global instance-level contrast module is proposed to improve interclass divergences and intraclass compactness. With a specially constructed global instance set, GICNet can contrast a large number of different aircraft targets while keeping a small batch size. Furthermore, we design a novel quality-aware focal loss (QAFL) to facilitate the accurate classification of well-localized aircraft targets. Meanwhile, to maintain localization performance, we develop a new edge-aware bounding-box refinement (EABR) module to refine predicted coarse bounding boxes. Experimental results show that our GICNet outperforms current advanced detectors and achieves a new state-of-the-art performance on the GaoFen-3 SAR aircraft detection dataset. In particular, GICNet also has advantages in reducing misclassification and recognizing well-located targets.
Danpei Zhao, Yue Gao 0008, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Semantic Segmentation of Remote Sensing Image Based on Regional Self-Attention Mechanism
abstract
In remote sensing images (RSIs), accurate semantic segmentation faces more challenges because of small targets, unbalanced categories, and complex scenes. Restricted by local receptive field of convolution layers, the traditional semantic segmentation models cannot use global information of RSIs. According to the characteristics of RSIs, we propose an RSANet based on regional self-attention mechanism. Our model is no longer limited by the locality of convolution, but transfers the information flow in the whole image. It can mine out the relationship between pixels in the surrounding areas, which is more logical for understanding images content. Moreover, compared with the traditional self-attention mechanism, RSANet can effectively reduce the noise of feature maps and the interference of redundant features. Our model can get better semantic segmentation results than other current models on the DroneDeploy data set and the Chreos semantic segmentation data set. The experiments show that our RSANet achieves 2% higher mean intersection over union (mIoU) than the baseline model, especially in terms of fineness, edge integrity, and classification accuracy.
Danpei Zhao, Chenxu Wang 0017, Yue Gao 0008, Zhenwei Shi 0001, Fengying Xie
IEEE Geosci. Remote. Sens. Lett.3
2022 UGCNet: An Unsupervised Semantic Segmentation Network Embedded With Geometry Consistency for Remote-Sensing Images
abstract
In remote-sensing image (RSI) semantic segmentation, the dependence on large-scale and pixel-level annotated data has been a critical factor restricting its development. In this letter, we propose an unsupervised semantic segmentation network embedded with geometry consistency (UGCNet) for RSIs, which imports the adversarial-generative learning strategy into a semantic segmentation network. The proposed UGCNet can be trained on a source-domain dataset and achieve accurate segmentation results on a different target-domain dataset. Furthermore, for refining the remote-sensing target geometric representation such as densely distributed buildings, we propose a geometry-consistency (GC) constraint that can be embedded in both image-domain adaptation process and semantic segmentation network. Therefore, our model could achieve cross-domain semantic segmentation with target geometric property preservation. The experimental results on Massachusetts and Inria buildings datasets prove that the proposed unsupervised UGCNet could achieve a very comparable segmentation accuracy with the fully supervised model, which validates the effectiveness of the proposed method.
Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Xinhu Qi, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.3
2019 Unsupervised Oil Tank Detection by Shape-Guide Saliency Model
abstract
In this letter, a novel oil tank detection framework based on a shape-guide saliency (SGS) model is proposed. Beyond the low-level visual stimuli, SGS focuses more on simulating the selective visual searching, which is dominated by the goal in human minds. Using a top–down strategy, SGS breaks the limitation of the low-level visual features and introduces the high-level task concept to measure saliency. For the oil tank detection, SGS model skillfully extracts the contour shape cue (CSC) as the target-oriented information and uses CSC to guide the selective saliency value calculation. Specifically, a sparse reconstruction with the target-specific dictionary is implemented to generate the saliency map. This saliency map only assigns high values to oil tank regions instead of highlighting all high-contrast regions. Consequently, SGS model is capable of accurately locating oil tanks and eliminating the interferences of high-contrast backgrounds. Experimental results on a remote sensing data set demonstrate that the proposed SGS model outperforms five class-independent saliency models. Comparisons with the state-of-the-art oil tank detection approaches demonstrate the effectiveness of the proposed method.
Minhao Jing, Danpei Zhao, Yue Gao 0008, Zhiguo Jiang 0001, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.4
2009 Maneuvering Target Tracking in Cluttered Background Based on Color Invariance and Support Vector Machine
abstract
Maneuvering targets tracking in cluttered environment is a challenging problem in computer vision because of the difficulty of distinguishing the target from the background. In this paper, we treat tracking as a binary classification problem and employ support vector machine to suppress the background. In order to enhance the robustness against illumination changes, we propose to combine color invariance with traditional RGB values to train the SVM. First, we use expectation maximization algorithm to extract the target from the environment; then, RGB and color invariance values are used to train SVM. In the incoming frames, pixels in regions of interest are classified by SVM and the confidence map is produced, which will afterward be used by traditional tracking approach to track the target, in this paper, we employ particle filter. Experimental results on challenging sequences validate the effectiveness of the proposed method in cluttered background target tracking.
Gang Meng, Zhiguo Jiang 0001, Danpei Zhao, Yue Gao 0008
ICIG4