EDBT 2026 Demo / reviewers in the wild / expert
Dong Zhao 0007
dblp:63/550-7
· DBLP profile ↗
24ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0001-9880-8822ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Stable Source-Free Domain Adaptive Semantic Segmentation
Dong Zhao 0007, Qi Zang, Nan Pu, Jinlong Li 0003, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
Int. J. Comput. Vis. | 1 |
| 2025 | ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion ModelabstractData-driven deep learning models have enabled tremendous progress in change detection (CD) with the support of pixel-level annotations. However, collecting diverse data and manually annotating them is costly, laborious, and knowledge-intensive. Existing generative methods for CD data synthesis show competitive potential in addressing this issue but still face the following limitations: 1) difficulty in flexibly controlling change events, 2) dependence on additional data to train the data generators, 3) focus on specific change detection tasks. To this end, this paper focuses on the semantic CD (SCD) task and develops a multi-temporal SCD data generator ChangeDiff by exploring powerful diffusion models. ChangeDiff innovatively generates change data in two steps: first, it uses text prompts and a text-to-layout (T2L) model to create continuous layouts, and then it employs layout-to-image (L2I) to convert these layouts into images. Specifically, we propose multi-class distribution-guided text prompts (MCDG-TP), allowing for layouts to be generated flexibly through controllable classes and their corresponding ratios. Subsequently, to generalize the T2L model to the proposed MCDG-TP, a class distribution refinement loss is further designed as training supervision. Our generated data shows significant progress in temporal continuity, spatial diversity, and quality realism, empowering change detectors with accuracy and transferability. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wenjun Yi, Zhun Zhong |
AAAI | 4 |
| 2025 | Feature Spectrum Learning for Remote Sensing Change DetectionabstractChange detection (CD) holds significant implications for Earth observation, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing methods mainly regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). Nevertheless, their efforts are limited by the complexity of optimizing GANs and the absence of guidance from physical properties. This paper finds that the spectrum transformation (ST) has the potential to mitigate pseudo-changes by aligning in the frequency domain carrying the style. However, the benefit of ST is largely constrained by two drawbacks: 1) limited transformation space and 2) inefficient parameter search. To address these limitations, we propose a Feature Spectrum learning (FeaSpect) that adaptively eliminate pseudo-changes in the latent space. For the drawback 1), FeaSpect directs the transformation towards stylealigned discriminative features via feature spectrum transformation (FST). For the drawback 2), FeaSpect allows FST to be trainable, efficiently discovering optimal parameters via extraction box with adaptive attention and extraction box with learnable strides. Extensive experiments on challenging datasets demonstrate that our method remarkably outperforms existing methods and achieves a commendable trade-off between accuracy and efficiency. Importantly, our method can be easily injected into other frameworks, achieving consistent improvements. Qi Zang, Dong Zhao 0007, Shuang Wang 0001, Dou Quan, Zhun Zhong |
CVPR | 2 |
| 2025 | FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized SegmentationabstractVision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains a challenge. Existing approaches either selectively fine-tune parameters or freeze the VFMs and update only the adapters, both of which may underutilize the VFMs’ full potential in DGSS tasks. We observe that domain-sensitive parameters in VFMs, arising from task and distribution differences, can hinder generalization. To address this, we propose FisherTune, a robust fine-tuning method guided by the Domain-Related Fisher Information Matrix (DR-FIM). DR-FIM measures parameter sensitivity across tasks and domains, enabling selective updates that preserve generalization and enhance DGSS adaptability. To stabilize DR-FIM estimation, FisherTune incorporates variational inference, treating parameters as Gaussian-Distributed variables and leveraging pre-trained priors. Extensive experiments show that Fisher-Tune achieves superior cross-domain segmentation while maintaining generalization, outperforming both selective-parameter and adapter-based methods. Dong Zhao 0007, Jinlong Li 0003, Shuang Wang 0001, Qi Zang, Nicu Sebe, Zhun Zhong |
CVPR | 1 |
| 2025 | Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic Segmentation
Dong Zhao 0007, Qi Zang, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
ICCV | 1 |
| 2025 | Predicting Spectral Information for Self-Supervised Signal ClassificationabstractDeep learning methods have demonstrated remarkable performance across various communication signal processing tasks. However, most signal classification methods require a substantial amount of labeled samples for training, posing significant challenges in the field of communication signals, as labeling necessitates expert knowledge. This paper proposes a novel self-supervised signal classification method called Spectral-Guided Self-Supervised Signal Classification (SGSSC). Specifically, to leverage frequency-domain information with modulation semantics as prior knowledge for the model, we design a previously unexplored pretext task tailored to the format of signal data. This task involves predicting spectral information from masked time-domain signals, enabling the model to learn implicit signal features through cross-domain pattern transformation. Furthermore, the pretext task in the SGSSC method is relevant to the downstream classification task, and using traditional fine-tuning strategies on the downstream task may lead to the loss of certain features associated with the pretext task. Therefore, we propose an attention mechanism-based fine-tuning strategy that adaptively integrates pre-trained features from different levels. Extensive experimental results validate the superiority of the SGSSC method. For instance, when the proportion of labeled samples is only 0.5%, our method achieves an average improvement of 2.3% in downstream classification tasks compared to the best-performing self-supervised training strategies. Shuang Wang 0001, Hantong Xing, Chenxu Wang 0001, Dou Quan, Rui Yang 0038, Dong Zhao 0007, Luyang Mei |
IJCAI | 7 |
| 2025 | SeCoV2: Semantic Connectivity-Driven Pseudo-Labeling for Robust Cross-Domain Semantic SegmentationabstractPseudo-labeling is a dominant strategy for cross-domain semantic segmentation (CDSS), yet its effectiveness is limited by fragmented and noisy pixel-level predictions under severe domain shifts. To address this, we propose a semantic connectivity-driven pseudo-labeling framework, SeCo, which constructs and refines pseudo-labels at the connectivity level by aggregating high-confidence pixels into coherent semantic regions. The framework includes two key components: Pixel Semantic Aggregation (PSA), which leverages a dual prompting strategy to preserve category-specific granularity, and Semantic Connectivity Correction with Loss Distribution (SCC-LD), which filters noisy regions based on early-loss statistics. Building upon this foundation, we further present SeCoV2, which introduces SCC-Unc, a novel uncertainty-aware correction module that constructs a connectivity graph and enforces relational consistency for robust refinement in ambiguous regions. SeCoV2 also broadens the applicability of SeCo by extending evaluation to more challenging scenarios, including open-set and multimodal adaptation, semi-supervised domain generalization, and by validating compatibility with different interactive foundation segmentation models such as SAM Kirillov et al. 2023, SEEM Zou et al. 2023, and Fast-SAM Zhao et al. 2023. Extensive experiments across six CDSS tasks demonstrate that SeCoV2 achieves consistent improvements over previous methods, with an average performance gain of up to +4.6%, establishing new state-of-the-art results. These findings highlight the effectiveness and generalization ability for robust adaptation in diverse real-world environments. Dong Zhao 0007, Qi Zang, Nan Pu, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Joint Style and Layout Synthesizing: Toward Generalizable Remote Sensing Semantic SegmentationabstractThis paper studies the domain generalized remote sensing semantic segmentation (RSSS), aiming to generalize a model trained only on the source domain to unseen domains. Existing methods in computer vision treat style information as domain characteristics to achieve domain-agnostic learning. Nevertheless, their generalizability to RSSS remains constrained, due to the incomplete consideration of domain characteristics. We argue that remote sensing scenes have layout differences beyond just style. Considering this, we devise a joint style and layout synthesizing framework, enabling the model to jointly learn out-of-domain samples synthesized from these two perspectives. For style, we estimate the variant intensities of per-class representations affected by domain shift and randomly sample within this modeled scope to reasonably expand the boundaries of style-carrying feature statistics. For layout, we explore potential scenes with diverse layouts in the source domain and propose granularity-fixed and granularity-learnable masks to perturb layouts, forcing the model to learn characteristics of objects rather than variable positions. The mask is designed to learn more context-robust representations by discovering difficult-to-recognize perturbation directions. Subsequently, we impose gradient angle constraints between the samples synthesized using the two ways to correct conflicting optimization directions. Extensive experiments demonstrate the superior generalization ability of our method over existing methods. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Zhun Zhong, Biao Hou, Licheng Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Generalization-Aware Remote Sensing Change Detection via Domain-Agnostic LearningabstractChange detection has essential significance for the region's development, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing transformation-based methods regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). However, their efforts are limited by two drawbacks: 1) Transformed images suffer from distortion that reduces feature discrimination. 2) Alignment hampers the model from learning domain-agnostic representations that degrades performance on scenes with domain shifts from the training data. Therefore, oriented from pseudo-changes caused by style differences, we present a generalizable domain-agnostic difference learning network (DonaNet). For the drawback 1), we argue for local-level statistics as style proxies to assist against domain shifts. For the drawback 2), DonaNet learns domain-agnostic representations by removing domain-specific style of encoded features and highlighting the class characteristics of objects. In the removal, we propose a domain difference removal module to reduce feature variance while preserving discriminative properties and propose its enhanced version to provide possibilities for eliminating more style by decorrelating the correlation between features. In the highlighting, we propose a cross-temporal generalization learning strategy to imitate latent domain shifts, thus enabling the model to extract feature representations more robust to shifts actively. Extensive experiments conducted on three public datasets demonstrate that DonaNet outperforms existing state-of-the-art methods with a smaller model size and is more robust to domain shift. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Licheng Jiao |
IEEE Trans. Multim. | 3 |
| 2025 | Boosting Generalization of Semantic Segmentation With Unseen Style Seeking-Based Meta-LearningabstractThis article considers a worst and most challenging scene in domain generalization (DG), where a model aims to generalize well on unseen domains while only one single domain is available for training. Existing randomization-based methods achieve this goal by enriching the style of the training data. However, they fail to guarantee the diversity of newly generated data required for generalization and thus lead to insufficient expansion of the training distribution. Thus, we propose a novel single DG (SDG) framework, unseen style seeking-based meta-learning (USSML). In USSML, multiple plausible domains with various styles are first constructed from a single source domain and the combination is performed across generated domains to emulate unseen images, extending the distribution boundaries of the source domain. The domain combination is performed at two levels, i.e., global and instance, to meet the generalization challenge in semantic segmentation. Then, the generated diverse domains are further exploited to force the model to optimize in an unbiased manner across all domains by relearning regions lacking domain-invariant representation capability, driving the model toward domain invariance. A point worth mentioning is that the proposed method is easily integrated into existing segmentation methods with little computational cost to improve their generalization. Extensive experiments are conducted on five popular segmentation datasets and the results have verified the effectiveness of USSML in improving the model's generalization and the superiority of USSML over existing works. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wanqing Li 0001, Dou Quan, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Stable Neighbor Denoising for Source-free Domain Adaptive SegmentationabstractWe study source-free unsupervised domain adaptation (SFUDA) for semantic segmentation, which aims to adapt a source-trained model to the target domain without accessing the source data. Many works have been proposed to address this challenging problem, among which uncertainty-based self-training is a predominant approach. However, without comprehensive denoising mechanisms, they still largely fall into biased estimates when dealing with different domains and confirmation bias. In this paper, we observe that pseudo-label noise is mainly contained in unstable samples in which the predictions of most pixels undergo significant variations during self-training. Inspired by this, we propose a novel mechanism to denoise unstable samples with stable ones. Specifically, we introduce the Stable Neighbor Denoising (SND) approach, which effectively discovers highly correlated stable and unstable samples by nearest neighbor retrieval and guides the reliable optimization of unstable samples by bi-level learning. Moreover, we compensate for the stable set by object-level object paste, which can further eliminate the bias caused by less learned classes. Our SND enjoys two advantages. First, SND does not require a specific segmentor structure, endowing its universality. Second, SND simultaneously addresses the issues of class, domain, and confirmation biases during adaptation, ensuring its effectiveness. Extensive experiments show that SND consistently outperforms state-of-the-art methods in various SFUDA semantic segmentation settings. In addition, SND can be easily integrated with other approaches, obtaining further improvements. The source code is available at https://github.com/DZhaoXd/SND. Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Licheng Jiao, Nicu Sebe, Zhun Zhong |
CVPR | 1 |
| 2024 | Active Domain Adaptive Semantic Segmentation with Regional Relative Entropy for Remote Sensing ImagesabstractThis paper presents a novel approach using active learning to tackle domain adaptation challenges in remote sensing semantic segmentation. Unsupervised Domain Adaptation for Semantic Segmentation (UDASS) aims to transfer a model trained on labeled source domain data to an unlabeled target domain. Existing UDASS methods struggle with the complexity of domain shift factors in remote sensing scenes, such as resolution, imaging mechanisms, geography, and species distribution, falling short of fully supervised performance. To address this, we propose integrating active learning, selecting a valuable (e.g. 2.2%) subset of pixel annotations from the target domain, and combining it with UDASS methods. Our method devises region-relative entropy metric to identify informative yet challenging pixels, facilitating better adaptation. Experimental results on two challenging domain adaptation tasks validate the efficacy of our technique, achieving performance comparable to fully supervised pixel labeling with only 2.2% annotated data. Zhengyao Wang, Dong Zhao 0007, Shuang Wang 0001 |
IGARSS | 5 |
| 2024 | ConDA: Continual Adaptation in Remote Sensing Via Visual Style PlaybackabstractThis study focuses on continual adaptation in remote sensing semantic segmentation, addressing challenges posed by frequent data updates and model forgetting. Remote sensing images exhibit variations in visual styles due to factors like location, time, and weather conditions, creating distinct domains. To counter performance degradation in new domains, we introduce a new challenge task in remote sensing, termed Continual Domain Adaptation (ConDA). Our innovative Visual Style Replay method employs Variational Auto-Encoder (VAE) and knowledge distillation, enabling the model to continuously learn from historical domains without forgetting. The proposed approach, tested on the INRIA dataset under ConDA settings, outperforms existing methods in combating catastrophic forgetting in remote sensing segmentation tasks. This contributes to adapting deep learning models to real-world scenarios with continually evolving remote sensing data. Ketao Zhong, Dong Zhao 0007, Shuang Wang 0001, Yanhe Guo |
IGARSS | 3 |
| 2024 | Mitigating Style Differences in Bitemporal Remote Sensing Images for Change DetectionabstractChange detection has seen significant advancements with the development of deep learning. However, due to variations in sensors or atmospheric conditions, bitemporal images often exhibit visually significant style differences, posing challenges for the detection of changed regions. This paper presents a change detection network designed to effectively address the challenges posed by style differences in bitemporal images. The proposed network comprises a color difference unification module and a generalized feature extraction module, which focuses the network on really changed areas. The color difference unification module harmonizes the color space of bitemporal remote sensing images, thereby mitigating the impact of style differences attributed to objective conditions. The generalized feature extraction module, ensuring robust feature representation for image pairs and further reducing style differences between bitemporal images. Experimental results demonstrate the superiority of our proposed method compared to existing change detection algorithms, confirming its suitability for fulfilling the requirements of change detection tasks. Qi Zang, Dong Zhao 0007, Shuang Wang 0001 |
IGARSS | 5 |
| 2024 | Object-Level Change Detection via Siamese Detection NetworkabstractTraditional change detection methods often lack instance-specific analysis, resulting in inefficient resource allocation and response strategies. This paper introduces a novel network for instance-level change detection. Our approach utilizes a dual-stream encoder with shared-weight backbone to extract robust features, followed by a differential process to highlight changes while suppressing unchanged background. Integration of low-level and high-level semantic information using a Feature Pyramid Network (FPN) enhances the model’s ability to discern subtle changes. Our instance-level bounding box detection module isolates individual change instances, with a subsequent segmentation module delineating precise boundaries. Evaluation on diverse remote sensing datasets demonstrates superior accuracy and computational efficiency compared to existing techniques. This framework not only advances change detection but also offers insights into land cover and land use dynamics. The code is available at https://github.com/DZhaoXd/object-levelchange-detection. Dong Zhao 0007, Hantong Xing, Shuang Wang 0001 |
IGARSS | 3 |
| 2024 | Generalized Source-Free Domain-adaptive Segmentation via Reliable Knowledge Propagation
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Jinlong Li 0003, Nicu Sebe, Zhun Zhong |
ACM Multimedia | 3 |
| 2024 | Connectivity-Driven Pseudo-Labeling Makes Stronger Cross-Domain SegmentersabstractPresently, pseudo-labeling stands as a prevailing approach in cross-domain semantic segmentation, enhancing model efficacy by training with pixels assigned with reliable pseudo-labels. However, we identify two key limitations within this paradigm: (1) under relatively severe domain shifts, most selected reliable pixels appear speckled and remain noisy. (2) when dealing with wild data, some pixels belonging to the open-set class may exhibit high confidence and also appear speckled. These two points make it difficult for the pixel-level selection mechanism to identify and correct these speckled close- and open-set noises. As a result, error accumulation is continuously introduced into subsequent self-training, leading to inefficiencies in pseudo-labeling. To address these limitations, we propose a novel method called Semantic Connectivity-driven Pseudo-labeling (SeCo). SeCo formulates pseudo-labels at the connectivity level, which makes it easier to locate and correct closed and open set noise. Specifically, SeCo comprises two key components: Pixel Semantic Aggregation (PSA) and Semantic Connectivity Correction (SCC). Initially, PSA categorizes semantics into ``stuff'' and ``things'' categories and aggregates speckled pseudo-labels into semantic connectivity through efficient interaction with the Segment Anything Model (SAM). This enables us not only to obtain accurate boundaries but also simplifies noise localization. Subsequently, SCC introduces a simple connectivity classification task, which enables us to locate and correct connectivity noise with the guidance of loss distribution. Extensive experiments demonstrate that SeCo can be flexibly applied to various cross-domain semantic segmentation tasks, \textit{i.e.} domain generalization and domain adaptation, even including source-free, and black-box domain adaptation, significantly improving the performance of existing state-of-the-art methods. The code is provided in the appendix and will be open-source. Dong Zhao 0007, Qi Zang, Shuang Wang 0001, Nicu Sebe, Zhun Zhong |
NeurIPS | 1 |
| 2024 | Transcending Fusion: A Multiscale Alignment Method for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multiscale representations in image content and text vocabulary can enable the models to learn richer representations and enhance retrieval. Current multiscale RSITR approaches typically align multiscale fused image features with text features but overlook aligning image-text pairs at distinct scales separately. This oversight restricts their ability to learn joint representations suitable for effective retrieval. We introduce a novel multiscale alignment (MSA) method to overcome this limitation. Our method comprises three key innovations: 1) a multiscale cross-modal alignment transformer (MSCMAT), which computes cross-attention between single-scale image features and localized text features, integrating global textual context to derive a matching score matrix within a mini-batch; 2) a multiscale cross-modal semantic alignment loss (MSCMA loss) that enforces semantic alignment across scales; and 3) a cross-scale multimodal semantic consistency loss (CSMMC loss) that uses the matching matrix from the largest scale to guide alignment at smaller scales. We evaluated our method across multiple datasets, demonstrating its efficacy with various visual backbones and establishing its superiority over existing state-of-the-art methods. The GitHub URL for our project ishttps://github.com/yr666666/MSA. Rui Yang 0038, Shuang Wang 0001, Yingping Han, Yuanheng Li, Dong Zhao 0007, Dou Quan, Yanhe Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Select, Purify, and Exchange: A Multisource Unsupervised Domain Adaptation Method for Building ExtractionabstractAccurately extracting buildings from aerial images has essential research significance for timely understanding human intervention on the land. The distribution discrepancies between diversified unlabeled remote sensing images (changes in imaging sensor, location, and environment) and labeled historical images significantly degrade the generalization performance of deep learning algorithms. Unsupervised domain adaptation (UDA) algorithms have recently been proposed to eliminate the distribution discrepancies without re-annotating training data for new domains. Nevertheless, due to the limited information provided by a single-source domain, single-source UDA (SSUDA) is not an optimal choice when multitemporal and multiregion remote sensing images are available. We propose a multisource UDA (MSUDA) framework SPENet for building extraction, aiming at selecting, purifying, and exchanging information from multisource domains to better adapt the model to the target domain. Specifically, the framework effectively utilizes richer knowledge by extracting target-relevant information from multiple-source domains, purifying target domain information with low-level features of buildings, and exchanging target domain information in an interactive learning manner. Extensive experiments and ablation studies constructed on 12 city datasets prove the effectiveness of our method against existing state-of-the-art methods, e.g., our method achieves 59.1% intersection over union (IoU) on Austin and Kitsap → Potsdam, which surpasses the target domain supervised method by 2.2%. The code is available at https://github.com/QZangXDU/SPENet. Shuang Wang 0001, Qi Zang, Dong Zhao 0007, Chaowei Fang, Dou Quan, Yutong Wan, Yanhe Guo, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic SegmentationabstractUnsupervised domain adaptation (UDA) in semantic segmentation transfers the knowledge of the source domain to the target one to improve the adaptability of the segmentation model in the target domain. The need to access labeled source data makes UDA unable to handle adaptation scenarios involving privacy, property rights protection, and confidentiality. In this paper, we focus on unsupervised model adaptation (UMA), also called source-free domain adaptation, which adapts a source-trained model to the target domain without accessing source data. We find that the online self-training method has the potential to be deployed in UMA, but the lack of source domain loss will greatly weaken the stability and adaptability of the method. We analyze two reasons for the degradation of online self-training, i.e. inopportune updates of the teacher model and biased knowledge from the source-trained model. Based on this, we propose a dynamic teacher update mechanism and a training-consistency based resampling strategy to improve the stability and adaptability of online self-training. On multiple model adaptation benchmarks, our method obtains new state-of-the-art performance, which is comparable or even better than state-of-the-art UDA methods. The code is available at https://github.com/DZhaoXd/DT-ST. Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Licheng Jiao |
CVPR | 1 |
| 2023 | Learning Pseudo-Relations for Cross-domain Semantic SegmentationabstractDomain adaptive semantic segmentation aims to adapt a model trained on labeled source domain to unlabeled target domain. Self-training shows competitive potential in this field. Existing methods along this stream mainly focus on selecting reliable predictions on target data as pseudo-labels for category learning, while ignoring the useful relations between pixels for relation learning. In this paper, we propose a pseudo-relation learning framework, Relation Teacher (RTea), which can exploitable pixel relations to efficiently use unreliable pixels and learn generalized representations. In this framework, we build reasonable pseudo-relations on local grids and fuse them with low-level relations in the image space, which are motivated by the reliable local relations prior and available low-level relations prior. Then, we design a pseudo-relation learning strategy and optimize the class probability to meet the relation consistency by finding the optimal sub-graph division. In this way, the model’s certainty and consistency of prediction are enhanced on the target domain, and the cross-domain inadaptation is further eliminated. Extensive experiments on three datasets demonstrate the effectiveness of the proposed method. The code will be available at https://github.com/DZhaoXd/RTea. Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Rui Yang 0038, Licheng Jiao |
ICCV | 1 |
| 2022 | MANet: Multi-Scale Aware-Relation Network for Semantic Segmentation in Aerial ScenesabstractSemantic segmentation is an important yet unsolved problem in aerial scenes understanding. One of the major challenges is the intense variations of scenes and object scales. In this paper, we propose a novel multi-scale aware-relation network (MANet) to tackle this problem in remote sensing. Inspired by the process of human perception of multi-scale information, we explore discriminative and diverse multi-scale representations. For discriminative multi-scale representations, we propose an inter-class and intra-class region refinement method (IIRR) to reduce feature redundancy caused by fusion. IIRR utilizes the refinement maps with intra- and inter-class scale variation to guide multi-scale fine-grained features. Then, we propose multi-scale collaborative learning (MCL) to enhance the diversity of multi-scale feature representations. The MCL constrains the diversity of multi-scale feature network parameters to obtain diverse information. And the segmentation results are rectified according to the dispersion of the multi-level network predictions. In this way, MANet can learn multi-scale features by collaboratively exploiting the correlation among different scales. Extensive experiments on image and video datasets which have large scale variations have demonstrated the effectiveness of our proposed MANet. Pei He, Licheng Jiao, Ronghua Shang, Shuang Wang 0001, Xu Liu 0006, Dou Quan, Dong Zhao 0007 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | Cluster Alignment With Target Knowledge Mining for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised domain adaptation (UDA) carries out knowledge transfer from the labeled source domain to the unlabeled target domain. Existing feature alignment methods in UDA semantic segmentation achieve this goal by aligning the feature distribution between domains. However, these feature alignment methods ignore the domain-specific knowledge of the target domain. In consequence, 1) the correlation among pixels of the target domain is not explored; and 2) the classifier is not explicitly designed for the target domain distribution. To conquer these obstacles, we propose a novel cluster alignment framework, which mines the domain-specific knowledge when performing the alignment. Specifically, we design a multi-prototype clustering strategy to make the pixel features within the same class tightly distributed for the target domain. Subsequently, a contrastive strategy is developed to align the distributions between domains, with the clustered structure maintained. After that, a novel affinity-based normalized cut loss is devised to learn task-specific decision boundaries. Our method enhances the model's adaptability in the target domain, and can be used as a pre-adaptation for self-training to boost its performance. Sufficient experiments prove the effectiveness of our method against existing state-of-the-art methods on representative UDA benchmarks. Shuang Wang 0001, Dong Zhao 0007, Yuwei Guo 0001, Qi Zang, Yu Gu 0015, Yi Li 0054, Licheng Jiao |
IEEE Trans. Image Process. | 2 |
| 2021 | Graph Regular Loss for Semi-Supervised Polsar Terrain ClassificationabstractAmongst the utilizations of Polarimetric Synthetic Aperture Radar (PoISAR) data, semi-supervised terrain classification is much in demand. Samples of the same category are distributed in multiple regions of a PolSAR image, resulting in differences in the feature distributions of samples of the same category located in different regions. In addition, some samples of confusable categories, and samples of different categories located near the edges, have small feature differences in a PolSAR image. To address this problem, we introduce a graph regular loss constructed from pseudo-labels to constrain the intra-class similarity and inter-class similarity of features, and improve the discriminative property of features. In addition, since the quality of pseudo-labels affects the optimization of the model by the graph regular loss, we introduce the idea of clustering into the teacher-student model to improve the quality of pseudo-labels. Experiments on real PolSAR data show that our proposed method achieves an excellent performance. Chunlei Han, Yuwei Guo 0001, Qi Zang, Baorui Duan, Dong Zhao 0007, Shuang Wang 0001 |
IGARSS | 7 |