Bo Yuan 0009

dblp:41/1662-9 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0001-7723-4823ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images With Autonomous Agents
abstract
Current methods for disaster scene interpretation in remote sensing images (RSIs) mostly focus on isolated tasks such as segmentation, detection, or visual question-answering (VQA). However, these methods often fail to provide comprehensive and actionable insights, particularly in scenarios that demand the integration of multiple perception methods and specialized tools to address complex, multilayered challenges in geophysical disaster analysis. To fill this gap, this article introduces adaptive disaster interpretation (ADI), a novel task designed to solve requests by planning and executing multiple sequentially correlative interpretation tasks to provide a comprehensive analysis of disaster scenes. To facilitate research and application in this area, we present a new dataset named RescueADI, which contains high-resolution RSIs with annotations for three connected aspects: planning, perception, and recognition. The dataset includes 4044 RSIs, 16949 semantic masks, 14483 object bounding boxes, and 13424 interpretation requests across nine challenging request types. Moreover, we propose a new disaster interpretation method employing autonomous agents driven by large language models (LLMs) for task planning and execution, proving its efficacy in handling complex disaster interpretations. The proposed agent-based method solves various complex interpretation requests such as counting, area calculation, and path finding without human intervention, which traditional single-task approaches cannot handle effectively. Experimental results on RescueADI demonstrate the feasibility of the proposed task and show that our method achieves an accuracy 9% higher than existing VQA methods, highlighting its advantages over conventional disaster interpretation approaches.
Zhuoran Liu 0006, Danpei Zhao, Bo Yuan 0009, Zhiguo Jiang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Reconciling Semantic Controllability and Diversity for Remote Sensing Image Synthesis With Hybrid Semantic Embedding
abstract
Significant advancements have been made in semantic image synthesis in remote sensing. However, existing methods still face formidable challenges in balancing semantic controllability and diversity. In this article, we present a hybrid semantic embedding guided generative adversarial network (HySEGGAN) for controllable and efficient remote sensing image synthesis. Specifically, HySEGGAN leverages hierarchical information from a single source. Motivated by feature description, we propose a hybrid semantic embedding method that coordinates fine-grained local semantic layouts to characterize the geometric structure of remote sensing objects without extra information. In addition, a semantic refinement network (SRN) is introduced, incorporating a novel loss function to ensure fine-grained semantic feedback. The proposed approach mitigates semantic confusion and prevents geometric pattern collapse. Experimental results indicate that the method strikes an excellent balance between semantic controllability and diversity. Furthermore, HySEGGAN significantly improves the quality of synthesized images and achieves state-of-the-art performance as a data augmentation technique across multiple datasets for downstream tasks.
Junde Liu, Danpei Zhao, Bo Yuan 0009, Tian Li 0009
IEEE Trans. Geosci. Remote. Sens.3
2024 Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
abstract
Continual learning (CL) breaks off the one-way training manner and enables a model to adapt to new data, semantics and tasks continuously. However, current CL methods mainly focus on single tasks. Besides, CL models are plagued by catastrophic forgetting and semantic drift since the lack of old data, which often occurs in remote-sensing interpretation due to the intricate fine-grained semantics. In this paper, we propose Continual Panoptic Perception (CPP), a unified continual learning model that leverages multi-task joint learning covering pixel-level classification, instance-level segmentation and image-level perception for universal interpretation in remote sensing images. Concretely, we propose a collaborative cross-modal encoder (CCE) to extract the input image features, which supports pixel classification and caption generation synchronously. To inherit the knowledge from the old model without exemplar memory, we propose a task-interactive knowledge distillation (TKD) method, which leverages cross-modal optimization and task-asymmetric pseudo-labeling (TPL) to alleviate catastrophic forgetting. Furthermore, we also propose a joint optimization mechanism to achieve end-to-end multi-modal panoptic perception. Experimental results on the fine-grained panoptic perception dataset validate the effectiveness of the proposed model, and also prove that joint optimization can boost sub-task CL efficiency with over 13% relative improvement on panoptic quality. The project page is available at https://github.com/YBIO/CPP.
Bo Yuan 0009, Danpei Zhao, Zhuoran Liu 0006, Tian Li 0009
ACM Multimedia1
2024 A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and Application
abstract
Continual learning, also known as incremental learning or life-long learning, stands at the forefront of deep learning and AI systems. It breaks through the obstacle of one-way training on close sets and enables continuous adaptive learning on open-set conditions. In the recent decade, continual learning has been explored and applied in multiple fields especially in computer vision covering classification, detection and segmentation tasks. Continual semantic segmentation (CSS), of which the dense prediction peculiarity makes it a challenging, intricate and burgeoning task. In this paper, we present a review of CSS, committing to building a comprehensive survey on problem formulations, primary challenges, universal datasets, neoteric theories and multifarious applications. Concretely, we begin by elucidating the problem definitions and primary challenges. Based on an in-depth investigation of relevant approaches, we sort out and categorize current CSS models into two main branches including data-replay and data-free sets. In each branch, the corresponding approaches are similarity-based clustered and thoroughly analyzed, following qualitative comparison and quantitative reproductions on relevant datasets. Besides, we also introduce four CSS specialities with diverse application scenarios and development tendencies. Furthermore, we develop a benchmark for CSS encompassing representative references, evaluation results and reproductions. We hope this survey can serve as a reference-worthy and stimulating contribution to the advancement of the life-long learning field, while also providing valuable perspectives for related fields.
Bo Yuan 0009, Danpei Zhao
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Learning at a Glance: Towards Interpretable Data-Limited Continual Semantic Segmentation via Semantic-Invariance Modelling
abstract
Continual semantic segmentation (CSS) based on incremental learning (IL) is a great endeavour in developing human-like segmentation models. However, current CSS approaches encounter challenges in the trade-off between preserving old knowledge and learning new ones, where they still need large-scale annotated data for incremental training and lack interpretability. In this paper, we present Learning at a Glance (LAG), an efficient, robust, human-like and interpretable approach for CSS. Specifically, LAG is a simple and model-agnostic architecture, yet it achieves competitive CSS efficiency with limited incremental data. Inspired by human-like recognition patterns, we propose a semantic-invariance modelling approach via semantic features decoupling that simultaneously reconciles solid knowledge inheritance and new-term learning. Concretely, the proposed decoupling manner includes two ways, i.e., channel-wise decoupling and spatial-level neuron-relevant semantic consistency. Our approach preserves semantic-invariant knowledge as solid prototypes to alleviate catastrophic forgetting, while also constraining sample-specific contents through an asymmetric contrastive learning method to enhance model robustness during IL steps. Experimental results in multiple datasets validate the effectiveness of the proposed method. Furthermore, we introduce a novel CSS protocol that better reflects realistic data-limited CSS settings, and LAG achieves superior performance under multiple data-limited conditions.
Bo Yuan 0009, Danpei Zhao, Zhenwei Shi 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 PETDet: Proposal Enhancement for Two-Stage Fine-Grained Object Detection
abstract
Fine-grained object detection (FGOD) extends object detection with the capability of fine-grained recognition. In recent two-stage FGOD methods, the region proposal serves as a crucial link between detection and fine-grained recognition. However, current methods overlook that some proposal-related procedures inherited from general detection are not equally suitable for FGOD, limiting the multitask learning from generation, representation, to utilization. In this article, we present a proposal enhancement for two-stage FGOD (PETDet) to better handle the subtasks in two-stage FGOD methods. First, an anchor-free quality-oriented proposal network (QOPN) is proposed with dynamic label assignment and attention-based decomposition to generate high-quality-oriented proposals. In addition, we present a bilinear channel fusion network (BCFN) to extract independent and discriminative features of the proposals. Furthermore, we designed a novel adaptive recognition loss (ARL) that offers guidance for the region-based convolutional neural networks (R-CNNs) head to focus on high-quality proposals. Extensive experiments validate the effectiveness of PETDet. Quantitative analysis reveals that PETDet with ResNet50 reaches state-of-the-art performance on various FGOD datasets, including FAIR1M-v1.0 (42.96 AP), FAIR1M-v2.0 (48.81 AP), MAR20 (85.91 AP), and ShipRSImageNet (74.90 AP). The proposed method also achieves superior compatibility between accuracy and inference speed. Our code and models will be released athttps://github.com/canoe-Z/PETDet.
Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 See, Perceive, and Answer: A Unified Benchmark for High-Resolution Postdisaster Evaluation in Remote Sensing Images
abstract
Visual-language generation for remote sensing image (RSI) is an emerging and challenging research area that requires multi-task learning to achieve a comprehensive understanding. However, most existing models are limited to single-level tasks and do not leverage the advantages of the Visual-Language Pre-training (VLP) model. In this paper, we present a unified benchmark that learns multiple tasks, including interpretation, perception, and question answering. Specifically, a model is designed to perform semantic segmentation, image captioning, and visual question answering for high-resolution RSIs simultaneously. Our model not only attains pixel-level segmentation accuracy and global semantic comprehension, but also responds to user-defined queries of interest. Moreover, to address the challenges of multi-task perception, we construct a novel multi-task data set called FloodNet+, which provides a new solution for the comprehensive post-disaster assessment. The experimental results demonstrate that our approach surpasses existing methods or baseline in all three tasks. This is the first attempt to simultaneously consider multiple remote sensing perception tasks in an integrated framework, which lays a solid foundation for future research in this area. Our data set are publicly available at: https://github.com/LDS614705356/FloodNet-plus.
Danpei Zhao, Jiankai Lu, Bo Yuan 0009
IEEE Trans. Geosci. Remote. Sens.3
2024 Panoptic Perception: A Novel Task and Fine-Grained Dataset for Universal Remote Sensing Image Interpretation
abstract
Current remote-sensing interpretation models often focus on a single task such as detection, segmentation, or caption. However, the task-specific designed models are unattainable to achieve the comprehensive multi-level interpretation of images. The field also lacks support for multi-task joint interpretation datasets. In this paper, we propose Panoptic Perception: a novel task and a new fine-grained dataset (FineGrip) to achieve a more thorough and universal interpretation for RSIs. The new task: 1) integrates pixel-level, instance-level, and image-level information for universal image perception, 2) captures image information from coarse to fine granularity, achieving deeper scene understanding and description, and 3) enables various independent tasks to complement and enhance each other through multi-task learning. By emphasizing multi-task interactions and the consistency of perception results, this task enables the simultaneous processing of fine-grained foreground instance segmentation, background semantic segmentation, and global fine-grained image captioning. Concretely, the FineGrip dataset includes 2,649 remote sensing images, 12054 fine-grained instance segmentation masks belonging to 20 foreground things categories, and 7599 background semantic masks for 5 stuff classes. Furthermore, we propose a joint optimization-based panoptic perception model. Experimental results on FineGrip demonstrate the feasibility of the panoptic perception task and the beneficial effect of multi-task joint optimization on individual tasks. The dataset will be publicly available.
Danpei Zhao, Bo Yuan 0009, Tian Li 0009, Zhuoran Liu 0006, Yue Gao 0008
IEEE Trans. Geosci. Remote. Sens.2
2023 Inherit With Distillation and Evolve With Contrast: Exploring Class Incremental Semantic Segmentation Without Exemplar Memory
abstract
As a front-burner problem in incremental learning, class incremental semantic segmentation (CISS) is plagued by catastrophic forgetting and semantic drift. Although recent methods have utilized knowledge distillation to transfer knowledge from the old model, they are still unable to avoid pixel confusion, which results in severe misclassification after incremental steps due to the lack of annotations for past and future classes. Meanwhile data-replay-based approaches suffer from storage burdens and privacy concerns. In this paper, we propose to address CISS without exemplar memory and resolve catastrophic forgetting as well as semantic drift synchronously. We present Inherit with Distillation and Evolve with Contrast (IDEC), which consists of a Dense Knowledge Distillation on all Aspects (DADA) manner and an Asymmetric Region-wise Contrastive Learning (ARCL) module. Driven by the devised dynamic class-specific pseudo-labelling strategy, DADA distils intermediate-layer features and output-logits collaboratively with more emphasis on semantic-invariant knowledge inheritance. ARCL implements region-wise contrastive learning in the latent space to resolve semantic drift among known classes, current classes, and unknown classes. We demonstrate the effectiveness of our method on multiple CISS tasks by state-of-the-art performance, including Pascal VOC 2012, ADE20K and ISPRS datasets. Our method also shows superior anti-forgetting ability, particularly in multi-step CISS tasks.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 UGCNet: An Unsupervised Semantic Segmentation Network Embedded With Geometry Consistency for Remote-Sensing Images
abstract
In remote-sensing image (RSI) semantic segmentation, the dependence on large-scale and pixel-level annotated data has been a critical factor restricting its development. In this letter, we propose an unsupervised semantic segmentation network embedded with geometry consistency (UGCNet) for RSIs, which imports the adversarial-generative learning strategy into a semantic segmentation network. The proposed UGCNet can be trained on a source-domain dataset and achieve accurate segmentation results on a different target-domain dataset. Furthermore, for refining the remote-sensing target geometric representation such as densely distributed buildings, we propose a geometry-consistency (GC) constraint that can be embedded in both image-domain adaptation process and semantic segmentation network. Therefore, our model could achieve cross-domain semantic segmentation with target geometric property preservation. The experimental results on Massachusetts and Inria buildings datasets prove that the proposed unsupervised UGCNet could achieve a very comparable segmentation accuracy with the fully supervised model, which validates the effectiveness of the proposed method.
Danpei Zhao, Bo Yuan 0009, Yue Gao 0008, Xinhu Qi, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Birds of a Feather Flock Together: Category-Divergence Guidance for Domain Adaptive Segmentation
abstract
Unsupervised domain adaptation (UDA) aims to enhance the generalization capability of a certain model from a source domain to a target domain. Present UDA models focus on alleviating the domain shift by minimizing the feature discrepancy between the source domain and the target domain but usually ignore the class confusion problem. In this work, we propose an Inter-class Separation and Intra-class Aggregation (ISIA) mechanism. It encourages the cross-domain representative consistency between the same categories and differentiation among diverse categories. In this way, the features belonging to the same categories are aligned together and the confusable categories are separated. By measuring the align complexity of each category, we design an Adaptive-weighted Instance Matching (AIM) strategy to further optimize the instance-level adaptation. Based on our proposed methods, we also raise a hierarchical unsupervised domain adaptation framework for cross-domain semantic segmentation task. Through performing the image-level, feature-level, category-level and instance-level alignment, our method achieves a stronger generalization performance of the model from the source domain to the target domain. In two typical cross-domain semantic segmentation tasks, i.e., GTA 5→ Cityscapes and SYNTHIA → Cityscapes, our method achieves the state-of-the-art segmentation accuracy. We also build two cross-domain semantic segmentation datasets based on the publicly available data, i.e., remote sensing building segmentation and road segmentation, for domain adaptive segmentation. Our code, models and datasets are available at https://github.com/HibiscusYB/BAFFT.
Bo Yuan 0009, Danpei Zhao, Shuai Shao 0005, Zehuan Yuan, Changhu Wang
IEEE Trans. Image Process.1
2021 V2RNet: An Unsupervised Semantic Segmentation Algorithm for Remote Sensing Images via Cross-Domain Transfer Learning
abstract
The dependence on large-scale pixel-level annotations brings great challenge to semantic segmentation task for remote sensing images (RSIs). To alleviate this issue, we propose V2RNet, an unsupervised semantic segmentation method which introduces adversarial learning into segmentation network. Our method creatively transfers the segmentation model from the synthetic GTA-V data to the real optical remote sensing data via domain adaptation. Additionally, to unify the source domain semantic structures and target domain image style, we design a semantic segmentation discriminator as auxiliary to optimize the domain adaptation efficiency. Thus the proposed method is effective on typical remote sensing targets such densely arranged, intertwined road. Experimental results on Massachusetts Road data set demonstrate our unsupervised semantic segmentation model achieves comparable segmentation accuracy, which also validates the effectiveness of the proposed method.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001
IGARSS3
2021 Selective focus saliency model driven by object class-awareness
abstract
Abstract Current many salient object detection (SOD) models only focus on highlighting visual conspicuous region but fail to make saliency detection for specific targets. In this paper, a selective focus saliency model driven by object class‐awareness (SF‐OCA) to run saliency detection is proposed. The framework consists of a visual saliency detection flow, a segmentation‐classification flow, and a class‐awareness selection module. It combines bottom‐up visual perception with a top‐down task‐driven manner, which is capable of detecting specific category salient targets and eliminating the interference from other saliency areas, providing a new idea for saliency detection. Experimental results show that the method achieves comparable performance with state‐of‐the‐art models on four public saliency datasets. In addition, a new dataset was also built to test the proposed framework for the selective focus saliency detection. Compared with other SOD methods, the method not only highlights visual saliency regions but can choose more important or more noteworthy targets in a class‐awareness manner. The method also shows better robustness under a variety of conditions including multi‐targets, small targets and complex background.
Danpei Zhao, Bo Yuan 0009, Zhenwei Shi 0001, Zhiguo Jiang 0001
IET Image Process.2