EDBT 2026 Demo / reviewers in the wild / expert
Danpei Zhao
dblp:41/7950
· DBLP profile ↗
32ranked-venue papers
10as first author
22since 2021 · last 2025
0000-0001-6701-0471ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images With Autonomous AgentsabstractCurrent methods for disaster scene interpretation in remote sensing images (RSIs) mostly focus on isolated tasks such as segmentation, detection, or visual question-answering (VQA). However, these methods often fail to provide comprehensive and actionable insights, particularly in scenarios that demand the integration of multiple perception methods and specialized tools to address complex, multilayered challenges in geophysical disaster analysis. To fill this gap, this article introduces adaptive disaster interpretation (ADI), a novel task designed to solve requests by planning and executing multiple sequentially correlative interpretation tasks to provide a comprehensive analysis of disaster scenes. To facilitate research and application in this area, we present a new dataset named RescueADI, which contains high-resolution RSIs with annotations for three connected aspects: planning, perception, and recognition. The dataset includes 4044 RSIs, 16949 semantic masks, 14483 object bounding boxes, and 13424 interpretation requests across nine challenging request types. Moreover, we propose a new disaster interpretation method employing autonomous agents driven by large language models (LLMs) for task planning and execution, proving its efficacy in handling complex disaster interpretations. The proposed agent-based method solves various complex interpretation requests such as counting, area calculation, and path finding without human intervention, which traditional single-task approaches cannot handle effectively. Experimental results on RescueADI demonstrate the feasibility of the proposed task and show that our method achieves an accuracy 9% higher than existing VQA methods, highlighting its advantages over conventional disaster interpretation approaches. Zhuoran Liu 0006, Danpei Zhao, Bo Yuan 0009, Zhiguo Jiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Reconciling Semantic Controllability and Diversity for Remote Sensing Image Synthesis With Hybrid Semantic EmbeddingabstractSignificant advancements have been made in semantic image synthesis in remote sensing. However, existing methods still face formidable challenges in balancing semantic controllability and diversity. In this article, we present a hybrid semantic embedding guided generative adversarial network (HySEGGAN) for controllable and efficient remote sensing image synthesis. Specifically, HySEGGAN leverages hierarchical information from a single source. Motivated by feature description, we propose a hybrid semantic embedding method that coordinates fine-grained local semantic layouts to characterize the geometric structure of remote sensing objects without extra information. In addition, a semantic refinement network (SRN) is introduced, incorporating a novel loss function to ensure fine-grained semantic feedback. The proposed approach mitigates semantic confusion and prevents geometric pattern collapse. Experimental results indicate that the method strikes an excellent balance between semantic controllability and diversity. Furthermore, HySEGGAN significantly improves the quality of synthesized images and achieves state-of-the-art performance as a data augmentation technique across multiple datasets for downstream tasks. Junde Liu, Danpei Zhao, Bo Yuan 0009, Tian Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing ImagesabstractContinual learning (CL) breaks off the one-way training manner and enables a model to adapt to new data, semantics and tasks continuously. However, current CL methods mainly focus on single tasks. Besides, CL models are plagued by catastrophic forgetting and semantic drift since the lack of old data, which often occurs in remote-sensing interpretation due to the intricate fine-grained semantics. In this paper, we propose Continual Panoptic Perception (CPP), a unified continual learning model that leverages multi-task joint learning covering pixel-level classification, instance-level segmentation and image-level perception for universal interpretation in remote sensing images. Concretely, we propose a collaborative cross-modal encoder (CCE) to extract the input image features, which supports pixel classification and caption generation synchronously. To inherit the knowledge from the old model without exemplar memory, we propose a task-interactive knowledge distillation (TKD) method, which leverages cross-modal optimization and task-asymmetric pseudo-labeling (TPL) to alleviate catastrophic forgetting. Furthermore, we also propose a joint optimization mechanism to achieve end-to-end multi-modal panoptic perception. Experimental results on the fine-grained panoptic perception dataset validate the effectiveness of the proposed model, and also prove that joint optimization can boost sub-task CL efficiency with over 13% relative improvement on panoptic quality. The project page is available at https://github.com/YBIO/CPP. Bo Yuan 0009, Danpei Zhao, Zhuoran Liu 0006, Tian Li 0009 |
ACM Multimedia | 2 |
| 2024 | YOLO-Parallel: Positive Gradient Modeling for Long-Tail Remote Sensing Object DetectionabstractThe long-tail distribution problem is widely prevalent in remote sensing images (RSIs), posing significant challenges to object detection tasks. Most existing methods for long-tail detection are designed for two-stage models. Such approaches of suppressing negative gradients tend to increase false alarms in one-stage detectors, resulting in a decline in overall performance and an increase in post-processing time. This paper presents a novel long-tail loss with broad applicability in diverse You Only Look Once (YOLO) networks. We present a novel Positive Gradient Loss (PGLoss) that effectively enhances the accuracy of tail categories while preserving the accuracy of head categories. Furthermore, to address the performance degradation caused by the pseudo-residual structure, we create Parallel Block with efficient computation and superior feature extraction abilities. We designed and trained the network named YOLO-Parallel to verify the effectiveness of PGLoss and Parallel Block. Extensive experiments were conducted on two large-scale optical remote sensing datasets, DIOR and DOTA, which are severely affected by the long-tail problem. The results powerfully demonstrate the superiority of our algorithm. YOLO-Parallel, with only 33.3% of the parameters of YOLOX, achieved a comparable detection performance of 96.9% on DIOR. On DOTA dataset, PGLoss achieved mean Average Precision (mAP) improvements of around 1.5% for YOLO-Parallel, YOLOv5n, and YOLOv7-tiny without increasing NMS processing time. Xiangyi Gao, Danpei Zhao, Zhichao Yuan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and ApplicationabstractContinual learning, also known as incremental learning or life-long learning, stands at the forefront of deep learning and AI systems. It breaks through the obstacle of one-way training on close sets and enables continuous adaptive learning on open-set conditions. In the recent decade, continual learning has been explored and applied in multiple fields especially in computer vision covering classification, detection and segmentation tasks. Continual semantic segmentation (CSS), of which the dense prediction peculiarity makes it a challenging, intricate and burgeoning task. In this paper, we present a review of CSS, committing to building a comprehensive survey on problem formulations, primary challenges, universal datasets, neoteric theories and multifarious applications. Concretely, we begin by elucidating the problem definitions and primary challenges. Based on an in-depth investigation of relevant approaches, we sort out and categorize current CSS models into two main branches including data-replay and data-free sets. In each branch, the corresponding approaches are similarity-based clustered and thoroughly analyzed, following qualitative comparison and quantitative reproductions on relevant datasets. Besides, we also introduce four CSS specialities with diverse application scenarios and development tendencies. Furthermore, we develop a benchmark for CSS encompassing representative references, evaluation results and reproductions. We hope this survey can serve as a reference-worthy and stimulating contribution to the advancement of the life-long learning field, while also providing valuable perspectives for related fields. Bo Yuan 0009, Danpei Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Learning at a Glance: Towards Interpretable Data-Limited Continual Semantic Segmentation via Semantic-Invariance ModellingabstractContinual semantic segmentation (CSS) based on incremental learning (IL) is a great endeavour in developing human-like segmentation models. However, current CSS approaches encounter challenges in the trade-off between preserving old knowledge and learning new ones, where they still need large-scale annotated data for incremental training and lack interpretability. In this paper, we present Learning at a Glance (LAG), an efficient, robust, human-like and interpretable approach for CSS. Specifically, LAG is a simple and model-agnostic architecture, yet it achieves competitive CSS efficiency with limited incremental data. Inspired by human-like recognition patterns, we propose a semantic-invariance modelling approach via semantic features decoupling that simultaneously reconciles solid knowledge inheritance and new-term learning. Concretely, the proposed decoupling manner includes two ways, i.e., channel-wise decoupling and spatial-level neuron-relevant semantic consistency. Our approach preserves semantic-invariant knowledge as solid prototypes to alleviate catastrophic forgetting, while also constraining sample-specific contents through an asymmetric contrastive learning method to enhance model robustness during IL steps. Experimental results in multiple datasets validate the effectiveness of the proposed method. Furthermore, we introduce a novel CSS protocol that better reflects realistic data-limited CSS settings, and LAG achieves superior performance under multiple data-limited conditions. Bo Yuan 0009, Danpei Zhao, Zhenwei Shi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | PETDet: Proposal Enhancement for Two-Stage Fine-Grained Object DetectionabstractFine-grained object detection (FGOD) extends object detection with the capability of fine-grained recognition. In recent two-stage FGOD methods, the region proposal serves as a crucial link between detection and fine-grained recognition. However, current methods overlook that some proposal-related procedures inherited from general detection are not equally suitable for FGOD, limiting the multitask learning from generation, representation, to utilization. In this article, we present a proposal enhancement for two-stage FGOD (PETDet) to better handle the subtasks in two-stage FGOD methods. First, an anchor-free quality-oriented proposal network (QOPN) is proposed with dynamic label assignment and attention-based decomposition to generate high-quality-oriented proposals. In addition, we present a bilinear channel fusion network (BCFN) to extract independent and discriminative features of the proposals. Furthermore, we designed a novel adaptive recognition loss (ARL) that offers guidance for the region-based convolutional neural networks (R-CNNs) head to focus on high-quality proposals. Extensive experiments validate the effectiveness of PETDet. Quantitative analysis reveals that PETDet with ResNet50 reaches state-of-the-art performance on various FGOD datasets, including FAIR1M-v1.0 (42.96 AP), FAIR1M-v2.0 (48.81 AP), MAR20 (85.91 AP), and ShipRSImageNet (74.90 AP). The proposed method also achieves superior compatibility between accuracy and inference speed. Our code and models will be released athttps://github.com/canoe-Z/PETDet. Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | See, Perceive, and Answer: A Unified Benchmark for High-Resolution Postdisaster Evaluation in Remote Sensing ImagesabstractVisual-language generation for remote sensing image (RSI) is an emerging and challenging research area that requires multi-task learning to achieve a comprehensive understanding. However, most existing models are limited to single-level tasks and do not leverage the advantages of the Visual-Language Pre-training (VLP) model. In this paper, we present a unified benchmark that learns multiple tasks, including interpretation, perception, and question answering. Specifically, a model is designed to perform semantic segmentation, image captioning, and visual question answering for high-resolution RSIs simultaneously. Our model not only attains pixel-level segmentation accuracy and global semantic comprehension, but also responds to user-defined queries of interest. Moreover, to address the challenges of multi-task perception, we construct a novel multi-task data set called FloodNet+, which provides a new solution for the comprehensive post-disaster assessment. The experimental results demonstrate that our approach surpasses existing methods or baseline in all three tasks. This is the first attempt to simultaneously consider multiple remote sensing perception tasks in an integrated framework, which lays a solid foundation for future research in this area. Our data set are publicly available at: https://github.com/LDS614705356/FloodNet-plus. Danpei Zhao, Jiankai Lu, Bo Yuan 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Panoptic Perception: A Novel Task and Fine-Grained Dataset for Universal Remote Sensing Image InterpretationabstractCurrent remote-sensing interpretation models often focus on a single task such as detection, segmentation, or caption. However, the task-specific designed models are unattainable to achieve the comprehensive multi-level interpretation of images. The field also lacks support for multi-task joint interpretation datasets. In this paper, we propose Panoptic Perception: a novel task and a new fine-grained dataset (FineGrip) to achieve a more thorough and universal interpretation for RSIs. The new task: 1) integrates pixel-level, instance-level, and image-level information for universal image perception, 2) captures image information from coarse to fine granularity, achieving deeper scene understanding and description, and 3) enables various independent tasks to complement and enhance each other through multi-task learning. By emphasizing multi-task interactions and the consistency of perception results, this task enables the simultaneous processing of fine-grained foreground instance segmentation, background semantic segmentation, and global fine-grained image captioning. Concretely, the FineGrip dataset includes 2,649 remote sensing images, 12054 fine-grained instance segmentation masks belonging to 20 foreground things categories, and 7599 background semantic masks for 5 stuff classes. Furthermore, we propose a joint optimization-based panoptic perception model. Experimental results on FineGrip demonstrate the feasibility of the panoptic perception task and the beneficial effect of multi-task joint optimization on individual tasks. The dataset will be publicly available. Danpei Zhao, Bo Yuan 0009, Tian Li 0009, Zhuoran Liu 0006, Yue Gao 0008 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Decoupled Instances Distillation for Remote Sensing Object DetectionabstractAs a significant field in object detection, knowledge distillation is plagued by coarse pixel division and complex object form. Although some methods currently use ground truth boxes for pixel division, they ignore consideration of teacher knowledge, resulting in inferior or even erroneous knowledge not being distinguished. In this paper, we propose a distillation method named Decoupled Instances Distillation, which solves the problem of imprecise pixel division and complex forms of remote sensing objects. First, we use the predicted boxes of the teacher network to distinguish different object instances and discard instances containing wrong knowledge. Second, considering the performance of classification and regression, we design the concentration function to differentiate object instances. Experimental results show that our method is powerful on remote sensing image datasets. Decoupled Instances Distillation improves the performance of YOLOv3 by 3.1% mAP on the DIOR test set, outperforming existing methods. Xiangyi Gao, Danpei Zhao |
IGARSS | 2 |
| 2023 | Inherit With Distillation and Evolve With Contrast: Exploring Class Incremental Semantic Segmentation Without Exemplar MemoryabstractAs a front-burner problem in incremental learning, class incremental semantic segmentation (CISS) is plagued by catastrophic forgetting and semantic drift. Although recent methods have utilized knowledge distillation to transfer knowledge from the old model, they are still unable to avoid pixel confusion, which results in severe misclassification after incremental steps due to the lack of annotations for past and future classes. Meanwhile data-replay-based approaches suffer from storage burdens and privacy concerns. In this paper, we propose to address CISS without exemplar memory and resolve catastrophic forgetting as well as semantic drift synchronously. We present Inherit with Distillation and Evolve with Contrast (IDEC), which consists of a Dense Knowledge Distillation on all Aspects (DADA) manner and an Asymmetric Region-wise Contrastive Learning (ARCL) module. Driven by the devised dynamic class-specific pseudo-labelling strategy, DADA distils intermediate-layer features and output-logits collaboratively with more emphasis on semantic-invariant knowledge inheritance. ARCL implements region-wise contrastive learning in the latent space to resolve semantic drift among known classes, current classes, and unknown classes. We demonstrate the effectiveness of our method on multiple CISS tasks by state-of-the-art performance, including Pascal VOC 2012, ADE20K and ISPRS datasets. Our method also shows superior anti-forgetting ability, particularly in multi-step CISS tasks. Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | A Hierarchical Decoder Architecture for Multilevel Fine-Grained Disaster DetectionabstractAs a cutting-edge challenge in the field of disaster evaluation, the detection of disasters in remote sensing images is crucial. However, most existing approaches to disaster detection simply solve the problem as a naive multi-class change detection, lacking accurate damage-level classification. In this paper, we propose a new approach to disaster detection called multi-level disaster detection (MLDD) that focuses on fine-grained damage-level classification. Our proposed approach tackles MLDD through hierarchical-correlation modeling and presents a universal disaster detection architecture. Specifically, we summarize two existing applicative methods, one-step training and pre-training, which are compatible with our proposed architecture. In addition, we propose two novel hierarchical approaches, namely the multi-task (MT) based and graph-encoding (GE) based approaches. The MT approach resolves MLDD through layer-wise learning in a progressive manner, building explicit multi-stage and implicit joint models to probe into the coarse-to-fine correlation for damage-level evaluation. The GE approach enhances hierarchical relationships by encoding multifold messaging directions and probabilities using a graph neural network. Furthermore, all four hierarchical paradigms can be embedded in our hierarchical MLDD architecture, which outperforms state-of-the-art methods on the xBD dataset, particularly in fine-grained damage-level classification. Overall, our proposed approach represents a significant improvement over existing disaster detection methods and has the potential to advance the field of disaster evaluation. Chenxu Wang 0017, Danpei Zhao, Xinhu Qi, Zhuoran Liu 0006, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Classification Matters More: Global Instance Contrast for Fine-Grained SAR Aircraft DetectionabstractSince significant intraclass differences and inconspicuous interclass variations, fine-grained aircraft detection in synthetic aperture radar (SAR) images is challenging. Also, the inherent lack of detailed features and severe noise interference in SAR images make it difficult to learn class-specific feature representations. Current detection approaches focus more on localization accuracy and ignore classification performance, which is more critical in fine-grained detection. To address the above challenges, we present GICNet: global instance contrast (GIC) for fine-grained SAR aircraft detection a global instance-level contrast module is proposed to improve interclass divergences and intraclass compactness. With a specially constructed global instance set, GICNet can contrast a large number of different aircraft targets while keeping a small batch size. Furthermore, we design a novel quality-aware focal loss (QAFL) to facilitate the accurate classification of well-localized aircraft targets. Meanwhile, to maintain localization performance, we develop a new edge-aware bounding-box refinement (EABR) module to refine predicted coarse bounding boxes. Experimental results show that our GICNet outperforms current advanced detectors and achieves a new state-of-the-art performance on the GaoFen-3 SAR aircraft detection dataset. In particular, GICNet also has advantages in reducing misclassification and recognizing well-located targets. Danpei Zhao, Yue Gao 0008, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Focal and Global Knowledge Distillation for DetectorsabstractKnowledge distillation has been applied to image classification successfully. However, object detection is much more sophisticated and most knowledge distillation methods have failed on it. In this paper, we point out that in object detection, the features of the teacher and student vary greatly in different areas, especially in the foreground and background. If we distill them equally, the uneven differences between feature maps will negatively affect the distillation. Thus, we propose Focal and Global Distillation (FGD). Focal distillation separates the foreground and background, forcing the student to focus on the teacher's critical pixels and channels. Global distillation rebuilds the relation between different pixels and transfers it from teachers to students, compensating for missing global information in focal distillation. As our method only needs to calculate the loss on the feature map, FGD can be applied to various detectors. We experiment on various detectors with different backbones and the results show that the student detector achieves excellent mAP improvement. For example, ResNet-50 based RetinaNet, Faster RCNN, RepPoints and Mask RCNN with our distillation method achieve 40.7%, 42.0%, 42.0% and 42.1% mAP on COCO2017, which are 3.3, 3.6, 3.4 and 2.9 higher than the baseline, respectively. Our codes are available at https://github.com/yzd-v/FGD. Zhendong Yang, Xiaohu Jiang, Yuan Gong 0002, Zehuan Yuan, Danpei Zhao, Chun Yuan 0003 |
CVPR | 6 |
| 2022 | Semantic Segmentation of Remote Sensing Image Based on Regional Self-Attention MechanismabstractIn remote sensing images (RSIs), accurate semantic segmentation faces more challenges because of small targets, unbalanced categories, and complex scenes. Restricted by local receptive field of convolution layers, the traditional semantic segmentation models cannot use global information of RSIs. According to the characteristics of RSIs, we propose an RSANet based on regional self-attention mechanism. Our model is no longer limited by the locality of convolution, but transfers the information flow in the whole image. It can mine out the relationship between pixels in the surrounding areas, which is more logical for understanding images content. Moreover, compared with the traditional self-attention mechanism, RSANet can effectively reduce the noise of feature maps and the interference of redundant features. Our model can get better semantic segmentation results than other current models on the DroneDeploy data set and the Chreos semantic segmentation data set. The experiments show that our RSANet achieves 2% higher mean intersection over union (mIoU) than the baseline model, especially in terms of fineness, edge integrity, and classification accuracy. Danpei Zhao, Chenxu Wang 0017, Yue Gao 0008, Zhenwei Shi 0001, Fengying Xie |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | UGCNet: An Unsupervised Semantic Segmentation Network Embedded With Geometry Consistency for Remote-Sensing ImagesabstractIn remote-sensing image (RSI) semantic segmentation, the dependence on large-scale and pixel-level annotated data has been a critical factor restricting its development. In this letter, we propose an unsupervised semantic segmentation network embedded with geometry consistency (UGCNet) for RSIs, which imports the adversarial-generative learning strategy into a semantic segmentation network. The proposed UGCNet can be trained on a source-domain dataset and achieve accurate segmentation results on a different target-domain dataset. Furthermore, for refining the remote-sensing target geometric representation such as densely distributed buildings, we propose a geometry-consistency (GC) constraint that can be embedded in both image-domain adaptation process and semantic segmentation network. Therefore, our model could achieve cross-domain semantic segmentation with target geometric property preservation. The experimental results on Massachusetts and Inria buildings datasets prove that the proposed unsupervised UGCNet could achieve a very comparable segmentation accuracy with the fully supervised model, which validates the effectiveness of the proposed method. Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Xinhu Qi, Zhenwei Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Birds of a Feather Flock Together: Category-Divergence Guidance for Domain Adaptive SegmentationabstractUnsupervised domain adaptation (UDA) aims to enhance the generalization capability of a certain model from a source domain to a target domain. Present UDA models focus on alleviating the domain shift by minimizing the feature discrepancy between the source domain and the target domain but usually ignore the class confusion problem. In this work, we propose an Inter-class Separation and Intra-class Aggregation (ISIA) mechanism. It encourages the cross-domain representative consistency between the same categories and differentiation among diverse categories. In this way, the features belonging to the same categories are aligned together and the confusable categories are separated. By measuring the align complexity of each category, we design an Adaptive-weighted Instance Matching (AIM) strategy to further optimize the instance-level adaptation. Based on our proposed methods, we also raise a hierarchical unsupervised domain adaptation framework for cross-domain semantic segmentation task. Through performing the image-level, feature-level, category-level and instance-level alignment, our method achieves a stronger generalization performance of the model from the source domain to the target domain. In two typical cross-domain semantic segmentation tasks, i.e., GTA 5→ Cityscapes and SYNTHIA → Cityscapes, our method achieves the state-of-the-art segmentation accuracy. We also build two cross-domain semantic segmentation datasets based on the publicly available data, i.e., remote sensing building segmentation and road segmentation, for domain adaptive segmentation. Our code, models and datasets are available at https://github.com/HibiscusYB/BAFFT. Bo Yuan 0009, Danpei Zhao, Shuai Shao 0005, Zehuan Yuan, Changhu Wang |
IEEE Trans. Image Process. | 2 |
| 2021 | V2RNet: An Unsupervised Semantic Segmentation Algorithm for Remote Sensing Images via Cross-Domain Transfer LearningabstractThe dependence on large-scale pixel-level annotations brings great challenge to semantic segmentation task for remote sensing images (RSIs). To alleviate this issue, we propose V2RNet, an unsupervised semantic segmentation method which introduces adversarial learning into segmentation network. Our method creatively transfers the segmentation model from the synthetic GTA-V data to the real optical remote sensing data via domain adaptation. Additionally, to unify the source domain semantic structures and target domain image style, we design a semantic segmentation discriminator as auxiliary to optimize the domain adaptation efficiency. Thus the proposed method is effective on typical remote sensing targets such densely arranged, intertwined road. Experimental results on Massachusetts Road data set demonstrate our unsupervised semantic segmentation model achieves comparable segmentation accuracy, which also validates the effectiveness of the proposed method. Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001 |
IGARSS | 1 |
| 2021 | Cross-Domain Transfer for Ship Instance Segmentation in SAR ImagesabstractConsidering insufficient data and difficulty of labeling in Synthetic Aperture Radar (SAR) images, we propose a method for SAR ship instance segmentation based on cross-domain transfer learning. Compared with optical images, transfer learning in SAR images faces the difficulties of insufficient data to pre-train and lacking detail features. The proposed method, containing sample transfer module and knowledge transfer module, simulates images from optics to SAR and pre-train the ship detection part of the instance segmentation network with simulation images. In addition, we design a Res-Pyramid network to prevent the deep network from being unable to extract efficient features of SAR images. The method proposed combines the content of the optics and the style of the SAR and incorporates multiscale features in backbone, which improves performance in ship instance segmentation in SAR images. Experiments show that it has achieved 1.3 and 1.1 points higher Average Precision (AP) in detection and segmentation tasks on SAR dataset of HRSID when using cross-domain transfer learning, which has exceeded state-of-the-art methods. Chunbo Zhu, Danpei Zhao, Xinhu Qi, Zhenwei Shi 0001 |
IGARSS | 2 |
| 2021 | Selective focus saliency model driven by object class-awarenessabstractAbstract Current many salient object detection (SOD) models only focus on highlighting visual conspicuous region but fail to make saliency detection for specific targets. In this paper, a selective focus saliency model driven by object class‐awareness (SF‐OCA) to run saliency detection is proposed. The framework consists of a visual saliency detection flow, a segmentation‐classification flow, and a class‐awareness selection module. It combines bottom‐up visual perception with a top‐down task‐driven manner, which is capable of detecting specific category salient targets and eliminating the interference from other saliency areas, providing a new idea for saliency detection. Experimental results show that the method achieves comparable performance with state‐of‐the‐art models on four public saliency datasets. In addition, a new dataset was also built to test the proposed framework for the selective focus saliency detection. Compared with other SOD methods, the method not only highlights visual saliency regions but can choose more important or more noteworthy targets in a class‐awareness manner. The method also shows better robustness under a variety of conditions including multi‐targets, small targets and complex background. Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001, Zhiguo Jiang 0001 |
IET Image Process. | 1 |
| 2021 | Single-shot weakly-supervised object detection guided by empirical saliency model
Danpei Zhao, Zhichao Yuan, Zhenwei Shi 0001, Fengying Xie |
Neurocomputing | 1 |
| 2021 | Spatiotemporal module for video saliency prediction based on self-attention
Zhuoran Liu 0006, Yibo Xia, Chunbo Zhu, Danpei Zhao |
Image Vis. Comput. | 5 |
| 2020 | Multi-Scale Remote Sensing Targets Detection with Rotated Feature PyramidabstractFor solving the difficult problem of multi-scale and multi-class target detection in complex environments of remote sensing, a target detection network is proposed based on rotated feature pyramid (RFP) and multi-scale context. Proposed method can overcome the interference caused by widely dispersed range in scale and terrain background. By extracting rotated anchors in four feature layers, the RFP module gains ample direction information to enhance plying-up target's contour. Through rotating anchors with a certain angle, RFP can decrease feature information of non-target area and avoid big scale anchor regression. Furthermore, we construct an anchor optimization method using multi-scale context which adjusts the anchor size proportion between different scales to improve the anchor selection accuracy. Experimental results on DIOR dataset demonstrate that the proposed network outperforms six state-of-the-art methods with 4.2% average precision higher. Beyond applicable to different backbones, our network has better performance for multi-class remote sensing targets. Yinan Mao, Hongkun Dou, Danpei Zhao |
IGARSS | 4 |
| 2020 | Hierarchical Attention for Ship Detection in SAR ImagesabstractConsidering the difficulty of ship detection in Synthetic Aperture Radar (SAR) images lacking color and texture details, we propose a method for SAR ship detection based on hierarchical attention mechanism. Compared with the optical images, the detection methods based on deep-learning for SAR images are aiming at designing a network that is sensitive to high-level features. The proposed method, containing Global Attention Module (GAM) and Local Attention Module (LAM), presents a hierarchical attention strategy respectively from the image level and the target level. GAM in both spatial and channel domain is constructed to highlight target characteristics. Anchor generation is guided by the LAM for locating more accurate candidate regions. Hierarchical attention improves the performance of extracting significant features of SAR images, which makes the algorithm more efficient in detecting small and multiple ship targets. Experiments results show that our method has achieved 0.7 to 4.1 points higher Average Precision (AP) than several state-of-the-art detection methods on SAR datasets of GF-3 and Sentinel-1. Chunbo Zhu, Danpei Zhao, Yinan Mao |
IGARSS | 2 |
| 2019 | Unsupervised Oil Tank Detection by Shape-Guide Saliency ModelabstractIn this letter, a novel oil tank detection framework based on a shape-guide saliency (SGS) model is proposed. Beyond the low-level visual stimuli, SGS focuses more on simulating the selective visual searching, which is dominated by the goal in human minds. Using a top–down strategy, SGS breaks the limitation of the low-level visual features and introduces the high-level task concept to measure saliency. For the oil tank detection, SGS model skillfully extracts the contour shape cue (CSC) as the target-oriented information and uses CSC to guide the selective saliency value calculation. Specifically, a sparse reconstruction with the target-specific dictionary is implemented to generate the saliency map. This saliency map only assigns high values to oil tank regions instead of highlighting all high-contrast regions. Consequently, SGS model is capable of accurately locating oil tanks and eliminating the interferences of high-contrast backgrounds. Experimental results on a remote sensing data set demonstrate that the proposed SGS model outperforms five class-independent saliency models. Comparisons with the state-of-the-art oil tank detection approaches demonstrate the effectiveness of the proposed method. Minhao Jing, Danpei Zhao, Yue Gao 0008, Zhiguo Jiang 0001, Zhenwei Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Sparsity-constrained probabilistic latent semantic analysis for land cover classificationabstractLand cover classification can be regarded as topic assignment that the pixels can be classified into different kinds of regions (e.g. road, tree, grass) according to the semantics of topics in topic model. In this paper, we present a novel probabilistic latent semantic analysis (pLSA) model based on sparsity constraint for classifying different kinds of land cover. In contrast with conventional topic model which usually assumes each local feature descriptor is only related to one visual word of the dictionary, our method uses sparse coding to characterize the potential relationship between the descriptor and multiple words. Therefore each descriptor can be represented by a small set of words. More importantly, we further apply sparse coding to mine the correlation of documents (i.e. image) in pLSA model. Consequently, our model can generate the more discriminative latent topics and benefit land cover classification. Experimental results on high-resolution remote sensing images demonstrate the excellent superiority of our method. Jun Shi 0006, Xilan Tian, Zhiguo Jiang 0001, Danpei Zhao |
IGARSS | 4 |
| 2016 | Hierarchical reinforcement learning for saliency detection of low-resolution airportsabstractThe traditional airport detection methods usually utilize geometric characteristics, which are limited by large amount of data and low-resolution of the remote sensing images. In this paper, we present a novel hierarchical reinforcement learning (HRL) saliency model for quickly airports detecting in large cover area. In contrast with conventional saliency models which usually are effective for high-resolution nature images, our method learns hierarchically high-level features via multi-scale superpixels segmentation and Least Absolute Shrinkage and Selection Operator (LASSO). More importantly, we introduce back-propagation theory for hierarchical learning to adaptively control and generate saliency map. Therefore our unsupervised saliency model is more simple and effective for low-resolution airport detection. Compared with 18 state-of-the-art saliency models, experimental results demonstrate the excellent performance of our method on the remote sensing image datasets. It is more robust and accurate for long-range airports detection. Danpei Zhao, Zhiguo Jiang 0001 |
IGARSS | 1 |
| 2013 | Local and Non-local Graph Regularized Sparse Coding for Face RecognitionabstractThe recent emerging sparse coding (SC) algorithms do not take local manifold structure of samples into consideration, while graph regularized sparse coding (GraphSC) algorithm only constrains the locality consistency of samples. Furthermore, the graph construction approach based on k-nearest-neighbor usually pre-defines the number of neighbors for all the samples, which may fails to fit the intrinsic structure of each sample. To address these issues, we propose an local and nonlocal graph regularized sparse coding (LN-GraphSC) algorithm. LN-GraphSC incorporates both local and nonlocal information of samples at the same time. On the other hand, to alleviate the problem of neighbor parameter selection, we use average distance of each sample to wisely determine its own local and nonlocal samples. To verify the effectiveness of our proposed method, we evaluate our method on the task of face recognition. The experimental results on ORL and Yale face databases show our method has competitive performance when compared to SC and GraphSC. Danpei Zhao, Jun Shi 0006, Zhiguo Jiang 0001 |
ICIG | 2 |
| 2011 | A Hierarchical Connection Graph Algorithm for Gable-Roof Detection in Aerial ImageabstractIn this letter, we present a hierarchical connection graph (HCG) algorithm based on a self-avoiding polygon (SAP) model for detecting and extracting gable roofs from aerial imagery. The SAP model is a deformable shape model that is capable of representing gable roofs of various shapes and appearances. The model is composed of a sequence of roof-corner templates that are connected into a SAP, which serves as a flexible shape prior. An energy function that combines features from three channels (corner, boundary, and interior area) is defined over the sequence to quantify the variability in appearances of gable roofs. To infer the most probable state of the corner sequence for an input image, we use an efficient algorithm-called HCG algorithm. The algorithm converts the solution space of a SAP model into a directed graph (which we call “HCG”) and searches for the best path using dynamic programming (DP). It is efficient for two reasons: 1) By constructing an HCG, the algorithm can quickly prune out a large amount of invalid solutions using only geometric constraints, which are inexpensive to compute, and 2) by employing DP, the algorithm decomposes the searching problem into smaller overlapping subproblems and reuses energy scores, which are expensive to compute. Experimental results on a set of challenging gable roofs show that our algorithm has good performance and is computationally effective. Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2009 | Gable Roof Description by Self-Avoiding Polygon
Qiongchen Wang, Zhiguo Jiang 0001, Junli Yang, Danpei Zhao, Zhenwei Shi 0001 |
ACCV (3) | 4 |
| 2009 | An architecture of optimised SIFT feature detection for an FPGA implementation of an image matcherabstractThis paper has proposed an architecture of optimised SIFT (scale invariant feature transform) feature detection for an FPGA implementation of an image matcher. In order for SIFT based image matcher to be implemented on an FPGA efficiently, in terms of speed and hardware resource usage, the original SIFT algorithm has been significantly optimised in the following aspects: 1) upsampling has been replaced with downsampling to save the interpolation operation. 2) Only four scales with two octaves are needed for our image matcher with moderate degradation of matching performance. 3) The total dimension of the feature descriptor has been reduced to 72 from 128 of the original SIFT, which leads to significantly simplify the image matching operation. With the optimisation above, the proposed FPGA implementation is able to detect the features of a typical image of 640 × 480 pixels within 31 milliseconds. Therefore, compared with the existing SIFT FPGA implementation, which requires 33 milliseconds for an image of 320 × 240 pixels, a significant improvement has been achieved for our proposed architecture. Lifan Yao, Yiqun Zhu, Zhiguo Jiang 0001, Danpei Zhao, Wenquan Feng |
FPT | 5 |
| 2009 | Maneuvering Target Tracking in Cluttered Background Based on Color Invariance and Support Vector MachineabstractManeuvering targets tracking in cluttered environment is a challenging problem in computer vision because of the difficulty of distinguishing the target from the background. In this paper, we treat tracking as a binary classification problem and employ support vector machine to suppress the background. In order to enhance the robustness against illumination changes, we propose to combine color invariance with traditional RGB values to train the SVM. First, we use expectation maximization algorithm to extract the target from the environment; then, RGB and color invariance values are used to train SVM. In the incoming frames, pixels in regions of interest are classified by SVM and the confidence map is produced, which will afterward be used by traditional tracking approach to track the target, in this paper, we employ particle filter. Experimental results on challenging sequences validate the effectiveness of the proposed method in cluttered background target tracking. Gang Meng, Zhiguo Jiang 0001, Danpei Zhao, Yue Gao 0008 |
ICIG | 3 |