Di Wang 0023

dblp:18/5410-23 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0001-6360-4360ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
abstract
Semi-supervised semantic segmentation (S4) has advanced remote sensing (RS) analysis by leveraging unlabeled data through pseudo-labeling and consistency learning. However, existing S4 studies often rely on small-scale datasets and models, limiting their practical applicability. To address this, we propose S5, the first scalable framework for semi-supervised semantic segmentation in RS, which unlocks the potential of vast unlabeled Earth observation data typically underutilized due to costly pixel-level annotations. Built upon existing large-scale RS datasets, S5 introduces a data selection strategy that integrates entropy-based filtering and diversity expansion, resulting in the RS4P-1M dataset. Using this dataset, we systematically scale up S4 into a new pretraining paradigm, S4 pre-training (S4P), to pretrain RS foundation models (RSFMs) of varying sizes on this extensive corpus, significantly boosting their performance on land cover segmentation and object detection tasks. Furthermore, during fine-tuning, we incorporate a Mixture-of-Experts (MoE)-based multi-dataset fine-tuning approach, which enables efficient adaptation to multiple RS benchmarks with fewer parameters. This approach improves the generalization and versatility of RSFMs across diverse RS benchmarks. The resulting RSFMs achieve state-of-the-art performance across all benchmarks, underscoring the viability of scaling semi-supervised learning for RS applications.
Di Wang 0023, Jing Zhang 0037, Lefei Zhang
AAAI2
2026 CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
abstract
Due to the substantial domain gaps in Remote Sensing (RS) images that are characterized by variabilities such as location, wavelength, and sensor type, Remote Sensing Domain Generalization (RSDG) has emerged as a critical and valuable research frontier, focusing on developing models that generalize effectively across diverse scenarios. However, research in this area remains underexplored: (1) Current cross-domain methods primarily focus on Domain Adaptation (DA), which adapts models to predefined domains rather than to unseen ones; (2) Few studies target the RSDG issue, especially for semantic segmentation tasks. Existing related models are developed for specific unknown domains, struggling with issues of underfitting on other unseen scenarios; (3) Existing RS foundation models tend to prioritize in-domain performance over cross-domain generalization. To this end, we introduce the first vision foundation model for RSDG semantic segmentation, CrossEarth. CrossEarth demonstrates strong cross-domain generalization through a specially designed data-level Earth-Style Injection pipeline and a model-level Multi-Task Training pipeline. In addition, for the semantic segmentation task, we have curated an RSDG benchmark comprising 32 semantic segmentation scenarios across various regions, spectral bands, platforms, and climates, providing comprehensive evaluations of the generalizability of future RSDG models. Extensive experiments on this collection demonstrate the superiority of CrossEarth over existing state-of-the-art methods.
Ziyang Gong, Zhixiang Wei, Di Wang 0023, Xiaoxing Hu, Xianzheng Ma, Hongruixuan Chen, Yuru Jia, Yupeng Deng 0002, Zhenming Ji, Xiangwei Zhu, Xue Yang 0005, Naoto Yokoya, Jing Zhang 0037, Bo Du 0001, Junchi Yan, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Holistic Invariant Retracing for Distortion-Resilient Multi-Modal Learning in Spatial Transcriptomics
abstract
Spatial transcriptomics provides a multi-modal perspective by simultaneously capturing gene expression profiles, spatial coordinates, and histological images. While existing methods focus on maintaining view consistency to handle distribution shifts, they frequently neglect semantic conflicts introduced by distorted views-a common limitation arising from technical data acquisition and processing constraints. These conflicts lead to distorted consensus representations. To address this challenge, we propose Holistic Invariant RetrAcing for mitigating representation distortion (HiraST). Our framework explicitly corrects distorted multi-view representations through two complementary mechanisms: 1) Cross-view invariant retracing, which jointly aligns instance-level features and pseudo-label distributions to retrace invariant information. This dual alignment ensures that semantically similar cells or tissue regions remain consistent across heterogeneous modalities, even in the presence of acquisition-induced distortions; and 2) holistic prototype learning, which leverages low-frequency structural components to recalibrate corrupted views and enhance robustness against noise. Extensive experiments on spatial transcriptomics datasets and incomplete multi-view clustering benchmarks demonstrate our framework's state-of-the-art performance. Meanwhile, HiraST demonstrates strong capability across various downstream tasks. The demo code of this work is publicly available at https://github.com/hexiao0275/HiraST.
Xiao He 0010, Huangxuan Zhao, Di Wang 0023, Dacheng Tao, Bo Du 0001
IEEE Trans. Image Process.3
2025 XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
abstract
The astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the imagery features ultra-high resolution that incorporates extremely complex semantic relationships. Existing benchmarks usually adopt notably smaller image sizes than real-world RS scenarios, suffer from limited annotation quality, and consider insufficient dimensions of evaluation. To address these issues, we present XLRS-Bench: a comprehensive benchmark for evaluating the perception and reasoning capabilities of MLLMs in ultra-high-resolution RS scenarios. XLRS-Bench boasts the largest average image size (8500×8500) observed thus far, with all evaluation samples meticulously annotated manually, assisted by a novel semi-automatic captioner on ultra-high-resolution RS images. On top of the XLRS-Bench, 16 sub-tasks are defined to evaluate MLLMs’ 10 kinds of perceptual capabilities and 6 kinds of reasoning capabilities, with a primary emphasis on advanced cognitive processes that facilitate real-world decision-making and the capture of spatiotemporal changes. The results of both general and RS-focused MLLMs on XLRS-Bench indicate that further efforts are needed for real-world RS applications. We have open-sourced XLRS-Bench to support further research in developing more powerful MLLMs for remote sensing.
Fengxiang Wang 0004, Hongzhen Wang, Zonghao Guo, Di Wang 0023, Yulin Wang 0002, Mingshuo Chen, Long Lan, Wenjing Yang 0002, Jing Zhang 0037, Zhiyuan Liu 0001, Maosong Sun 0001
CVPR4
2025 Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
Fengxiang Wang 0004, Hongzhen Wang, Di Wang 0023, Zonghao Guo, Zhenyu Zhong, Long Lan, Wenjing Yang 0002, Jing Zhang 0037
ICCV3
2025 GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
abstract
Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce **SuperRS-VQA** (avg. 8,376$\times$8,376) and **HighRS-VQA** (avg. 2,000$\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: *Background Token Pruning* and *Anchored Token Selection*, to reduce the memory footprint while preserving key semantics. Integrating these techniques, we introduce **GeoLLaVA-8K**, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench. Datasets and code were released at https://github.com/MiliLab/GeoLLaVA-8K.
Fengxiang Wang 0004, Mingshuo Chen, Di Wang 0023, Haotian Wang 0001, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang 0002, Hongzhen Wang, Wenjing Yang 0002, Bo Du 0001, Jing Zhang 0037
NeurIPS4
2025 RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
abstract
Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. While the linear-complexity Mamba architecture offers a promising alternative, existing RS applications of Mamba remain limited to supervised tasks on small, domain-specific datasets. To address these challenges, we propose RoMA, a framework that enables scalable self-supervised pretraining of Mamba-based RS foundation models using large-scale, diverse, unlabeled data. RoMA enhances scalability for high-resolution images through a tailored auto-regressive learning strategy, incorporating two key innovations: 1) a rotation-aware pretraining mechanism combining adaptive cropping with angular embeddings to handle sparsely distributed objects with arbitrary orientations, and 2) multi-scale token prediction objectives that address the extreme variations in object scales inherent to RS imagery. Systematic empirical studies validate that Mamba adheres to RS data and parameter scaling laws, with performance scaling reliably as model and data size increase. Furthermore, experiments across scene classification, object detection, and semantic segmentation tasks demonstrate that RoMA-pretrained Mamba models consistently outperform ViT-based counterparts in both accuracy and computational efficiency. The source code and pretrained models have be released at https://github.com/MiliLab/RoMA.
Fengxiang Wang 0004, Yulin Wang 0002, Mingshuo Chen, Haotian Wang 0001, Hongzhen Wang, Haiyan Zhao 0001, Yangang Sun, Di Wang 0023, Long Lan, Wenjing Yang 0002, Jing Zhang 0037
NeurIPS9
2025 DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration
abstract
Diffusion models have achieved remarkable progress in universal image restoration. However, existing methods perform naive inference in the reverse process, which leads to cumulative errors under limited sampling steps and large step intervals. Moreover, they struggle to balance the commonality of degradation representations with restoration quality, often depending on complex compensation mechanisms that enhance fidelity at the expense of efficiency. To address these challenges, we introduce \textbf{DGSolver}, a diffusion generalist solver with universal posterior sampling. We first derive the exact ordinary differential equations for generalist diffusion models to unify degradation representations and design tailored high-order solvers with a queue-based accelerated sampling strategy to improve both accuracy and efficiency. We then integrate universal posterior sampling to better approximate manifold-constrained gradients, yielding a more accurate noise estimation and correcting errors in inverse inference. Extensive experiments demonstrate that DGSolver outperforms state-of-the-art methods in restoration accuracy, stability, and scalability, both qualitatively and quantitatively. Code and models are publicly available at https://github.com/MiliLab/DGSolver.
Hebaixu Wang, Jing Zhang 0037, Di Wang 0023, Jiayi Ma 0001, Bo Du 0001
NeurIPS4
2025 PHDMamba: Progressive Hybrid Mamba for Hyperspectral Image Classification
abstract
Although Mamba-based models have demonstrated great potential in hyperspectral image (HSI) classification, existing approaches often rely on patch-wise inputs, causing redundant computation and limiting global spectral–spatial modeling, while the absence of hierarchical representation and explicit interaction further constrains fine-grained fusion. To address these limitations, we propose a Progressive Hybrid Mamba Model (PHDMamba), which is designed to progressively capture long-range spectral-spatial contextual dependencies and gradually fuse spectral and spatial information through adaptive feature interaction, ultimately yielding a unified representation. Specifically, we develop a Progressive Hybrid Mamba Module, which performs stage-wise modeling of long-range dependencies along both spectral and spatial dimensions. In addition, a dedicated Spectral-Spatial Interaction Module is introduced to adaptively integrate contextual spectral and spatial features. Experimental results on three widely used HSI benchmark datasets demonstrate that the proposed method achieves superior classification performance compared to existing approaches. The implementation code will be released publicly.
Yichu Xu, Chengxi Han, Shi Chen 0010, Yuchun Miao, Di Wang 0023
IEEE Geosci. Remote. Sens. Lett.7
2025 Dual selective fusion transformer network for hyperspectral image classification
Yichu Xu, Di Wang 0023, Lefei Zhang, Liangpei Zhang 0001
Neural Networks2
2025 HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model
abstract
Accurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency.
Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 TransWCD: Scene-Adaptive Joint Constrained Framework for Weakly Supervised Change Detection
abstract
Change detection (CD) based on deep learning typically requires costly pixel-level change labels. Recently, weakly supervised CD (WSCD) has emerged as a more label-efficient approach, using scene-level (i.e., image-level) labels to identify pixel-level changes in bitemporal images. With only scene-level labels, existing WSCD methods are typically trained as scene-level change classification models. However, these methods often suffer from label-prediction inconsistency, with false changes frequently predicted in unchanged scenes. To address this issue, we propose TransWCD-SA, an end-to-end classifier-predictor framework. TransWCD-SA consists of a hierarchical transformer-based TransWCD classifier and a scene-adaptive (SA) predictor. This classifier-predictor framework is trained with two-stage joint constraints in an end-to-end learning manner. Specifically, the TransWCD classifier integrates hierarchical transformer blocks and multiscale class activation maps (CAMs), capturing pixel-level changes across various scales under weak supervision. The SA predictor dynamically introduces different pixel-level information for scenes labeled as changed and unchanged. Furthermore, a scene gated constraint is proposed as a penalty for label-prediction inconsistency, which is activated by the Dirac delta function and rectify features of mispredicted pixels in the embedding space. We validate the effectiveness of TransWCD-SA on three datasets: Wuhan University building CD (WHU-CD), learning, vision, and remote sensing CD (LEVIR-CD), and DSIFN-CD, demonstrating significant improvement. The code is available athttps://github.com/zhenghuizhao/TransWCD.
Zhenghui Zhao, Lixiang Ru, Chen Wu 0003, Di Wang 0023
IEEE Trans. Geosci. Remote. Sens.4
2025 Advancing Weakly-Supervised Change Detection in Satellite Images via Adversarial Class Prompting
abstract
Weakly-Supervised Change Detection (WSCD) aims to distinguish specific object changes (e.g., objects appearing or disappearing) from background variations (e.g., environmental changes due to light, weather, or seasonal shifts) in paired satellite images, relying only on paired image (i.e., image-level) classification labels. This technique significantly reduces the need for dense annotations required in fully-supervised change detection. However, as image-level supervision only indicates whether objects have changed in a scene, WSCD methods often misclassify background variations as object changes, especially in complex remote-sensing scenarios. In this work, we propose an Adversarial Class Prompting (AdvCP) method to address this co-occurring noise problem, including two phases: a) Adversarial Prompt Mining: After each training iteration, we introduce adversarial prompting perturbations, using incorrect one-hot image-level labels to activate erroneous feature mappings. This process reveals co-occurring adversarial samples under weak supervision, namely background variation features that are likely to be misclassified as object changes. b) Adversarial Sample Rectification: We integrate these adversarially prompt-activated pixel samples into training by constructing an online global prototype. This prototype is built from an exponentially weighted moving average of the current batch and all historical training data. Serving as an unbiased anchor, the global prototype guides the rectification of adversarial pixel samples. Our AdvCP can be seamlessly integrated into current WSCD methods without adding additional inference cost. Experiments on ConvNet, Transformer, and Segment Anything Model (SAM)-based baselines demonstrate significant performance enhancements, achieving up to 7.37%, 7.46%, and 6.56% IoU improvements on the WHU-CD, LEVIR-CD, and DSIFN-CD datasets. Furthermore, we demonstrate the generalizability of AdvCP to other multi-class weakly-supervised dense prediction scenarios. Code is available at https://github.com/zhenghuizhao/AdvCP.
Zhenghui Zhao, Chen Wu 0003, Di Wang 0023, Hongruixuan Chen, Cuiqun Chen, Zhuo Zheng, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Image Process.3
2024 LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
Jing Zhang 0037, Di Wang 0023, Qiming Zhang 0001, Zengmao Wang, Bo Du 0001
IJCAI3
2024 ITER: Image-to-Pixel Representation for Weakly Supervised HSI Classification
abstract
Recent years have witnessed the superiority of deep learning-based algorithms in the field of HSI classification. However, a prerequisite for the favorable performance of these methods is a large number of refined pixel-level annotations. Due to atmospheric changes, sensor differences, and complex land cover distribution, pixel-level labeling of high-dimensional hyperspectral image (HSI) is extremely difficult, time-consuming, and laborious. To overcome the above hurdle, an Image-To-pixEl Representation (ITER) approach is proposed in this paper. To the best of our knowledge, this is the first time that image-level annotation is introduced to predict pixel-level classification maps for HSI. The proposed model is along the lines of subject modeling to boundary refinement, corresponding to pseudo-label generation and pixel-level prediction. Concretely, in the pseudo-label generation part, the spectral/spatial activation, spectral-spatial alignment loss, and geographic element enhancement are sequentially designed to locate discriminate regions of each category, optimize multi-domain class activation map (CAM) collaborative training, and refine labels, respectively. For the pixel-level prediction portion, a high frequency-aware self-attention in a high-enhanced transformer is put forward to achieve detailed feature representation. With the two-stage pipeline, ITER explores weakly supervised HSI classification with image-level tags, bridging the gap between image-level annotation and dense prediction. Extensive experiments in three benchmark datasets with state-of-the-art (SOTA) works show the performance of the proposed approach.
Jiaqi Yang 0005, Bo Du 0001, Di Wang 0023, Liangpei Zhang 0001
IEEE Trans. Image Process.3
2024 Spectral-Spatial Global Graph Reasoning for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have been widely applied to hyperspectral image classification (HSIC). However, traditional convolutions can not effectively extract features for objects with irregular distributions. Recent methods attempt to address this issue by performing graph convolutions on spatial topologies, but fixed graph structures and local perceptions limit their performances. To tackle these problems, in this article, different from previous approaches, we perform the superpixel generation on intermediate features during network training to adaptively produce homogeneous regions, obtain graph structures, and further generate spatial descriptors, which are served as graph nodes. Besides spatial objects, we also explore the graph relationships between channels by reasonably aggregating channels to generate spectral descriptors. The adjacent matrices in these graph convolutions are obtained by considering the relationships among all descriptors to realize global perceptions. By combining the extracted spatial and spectral graph features, we finally obtain a spectral-spatial graph reasoning network (SSGRN). The spatial and spectral parts of SSGRN are separately called spatial and spectral graph reasoning subnetworks. Comprehensive experiments on four public datasets demonstrate the competitiveness of the proposed methods compared with other state-of-the-art graph convolution-based approaches.
Di Wang 0023, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 HKNAS: Classification of Hyperspectral Imagery Based on Hyper Kernel Neural Architecture Search
abstract
Recent neural architecture search (NAS)-based approaches have made great progress in the hyperspectral image (HSI) classification tasks. However, the architectures are usually optimized independently of the network weights, increasing searching time, and restricting model performances. To tackle these issues, in this article, different from previous methods that extra define structural parameters, we propose to directly generate structural parameters by utilizing the specifically designed hyper kernels, ingeniously converting the original complex dual optimization problem into easily implemented one-tier optimizations, and greatly shrinking searching costs. Then, we develop a hierarchical multimodule search space whose candidate operations only contain convolutions, and these operations can be integrated into unified kernels. Using the above searching strategy and searching space, we obtain three kinds of networks to separately conduct pixel-level or image-level classifications with 1-D or 3-D convolutions. In addition, by combining the proposed hyper kernel searching scheme with the 3-D convolution decomposition mechanism, we obtain diverse architectures to simulate 3-D convolutions, greatly improving network flexibilities. A series of quantitative and qualitative experiments on six public datasets demonstrate that the proposed methods achieve state-of-the-art results compared with other advanced NAS-based HSI classification approaches.
Di Wang 0023, Bo Du 0001, Liangpei Zhang 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2023 SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model
abstract
The success of the Segment Anything Model (SAM) demonstrates the significance of data-centric machine learning. However, due to the difficulties and high costs associated with annotating Remote Sensing (RS) images, a large amount of valuable RS data remains unlabeled, particularly at the pixel level. In this study, we leverage SAM and existing RS object detection datasets to develop an efficient pipeline for generating a large-scale RS segmentation dataset, dubbed SAMRS. SAMRS totally possesses 105,090 images and 1,668,241 instances, surpassing existing high-resolution RS segmentation datasets in size by several orders of magnitude. It provides object category, location, and instance information that can be used for semantic segmentation, instance segmentation, and object detection, either individually or in combination. We also provide a comprehensive analysis of SAMRS from various aspects. Moreover, preliminary experiments highlight the importance of conducting segmentation pre-training with SAMRS to address task discrepancies and alleviate the limitations posed by limited training data during fine-tuning. The code and dataset will be available at https://github.com/ViTAE-Transformer/SAMRS
Di Wang 0023, Jing Zhang 0037, Bo Du 0001, Minqiang Xu, Dacheng Tao, Liangpei Zhang 0001
NeurIPS1
2023 An Empirical Study of Remote Sensing Pretraining
abstract
Deep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights since natural images inevitably present a large domain gap relative to aerial images, probably limiting the fine-tuning performance on downstream aerial scene tasks. This issue motivates us to conduct an empirical study of RS pretraining (RSP) on aerial images. To this end, we train different networks from scratch with the help of the largest RS scene recognition dataset up to now—MillionAID—to obtain a series of RS pretrained backbones, including both convolutional neural networks (CNNs) and vision transformers, such as Swin and ViTAE, which have shown promising performance on computer vision tasks. Then, we investigate the impact of RSP on representative downstream tasks, including scene recognition, semantic segmentation, object detection, and change detection using these CNN and vision transformer backbones. Empirical study shows that RSP can help deliver distinctive performances in scene recognition tasks and in perceiving RS-related semantics, such as “Bridge” and “Airplane.” We also find that, although RSP mitigates the data discrepancies of traditional ImageNet pretraining on RS images, it may still suffer from task discrepancies, where downstream tasks require different representations from scene recognition tasks. These findings call for further research efforts on both large-scale pretraining datasets and effective pretraining methods. The codes and pretrained models will be released athttps://github.com/ViTAE-Transformer/ViTAE-Transformer-Remote-Sensing.
Di Wang 0023, Jing Zhang 0037, Bo Du 0001, Gui-Song Xia, Dacheng Tao
IEEE Trans. Geosci. Remote. Sens.1
2023 Advancing Plain Vision Transformer Toward Remote Sensing Foundation Model
abstract
Large-scale vision foundation models have made significant progress in visual tasks on natural images, with vision transformers (ViTs) being the primary choice due to their good scalability and representation ability. However, large-scale models in remote sensing (RS) have not yet been sufficiently explored. In this article, we resort to plain ViTs with about 100 million parameters and make the first attempt to propose large vision models tailored to RS tasks and investigate how such large models perform. To handle the large sizes and objects of arbitrary orientations in RS images, we propose a new rotated varied-size window attention to replace the original full attention in transformers, which can significantly reduce the computational cost and memory footprint while learning better object representation by extracting rich context from the generated diverse windows. Experiments on detection tasks show the superiority of our model over all state-of-the-art models, achieving 81.24% mean average precision (mAP) on the DOTA-V1.0 dataset. The results of our models on downstream classification and segmentation tasks also show competitive performance compared to existing advanced methods. Further experiments show the advantages of our models in terms of computational complexity and data efficiency in transferring. The code and models will be released athttps://github.com/ViTAE-Transformer/Remote-Sensing-RVSA.
Di Wang 0023, Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 DCN-T: Dual Context Network With Transformer for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is challenging due to spatial variability caused by complex imaging conditions. Prior methods suffer from limited representation ability, as they train specially designed networks from scratch on limited annotated data. We propose a tri-spectral image generation pipeline that transforms HSI into high-quality tri-spectral images, enabling the use of off-the-shelf ImageNet pretrained backbone networks for feature extraction. Motivated by the observation that there are many homogeneous areas with distinguished semantic and geometric properties in HSIs, which can be used to extract useful contexts, we propose an end-to-end segmentation network named DCN-T. It adopts transformers to effectively encode regional adaptation and global aggregation spatial contexts within and between the homogeneous areas discovered by similarity-based clustering. To fully exploit the rich spectrums of the HSI, we adopt an ensemble approach where all segmentation results of the tri-spectral images are integrated into the final prediction through a voting scheme. Extensive experiments on three public benchmarks show that our proposed method outperforms state-of-the-art methods for HSI classification. The code will be released at https://github.com/DotWang/DCN-T.
Di Wang 0023, Jing Zhang 0037, Bo Du 0001, Liangpei Zhang 0001, Dacheng Tao
IEEE Trans. Image Process.1
2022 Fully Contextual Network for Hyperspectral Scene Parsing
abstract
In this article, we propose fully contextual networks (FullyContNets) for hyperspectral scene parsing. Different from the previous approaches that leveraging the local information, the proposed methods can effectively capture the more generic nonlocal contexts. To this end, we first propose the scale attention module (SAM) that can adaptively aggregate the multiple features through obtaining the interfeature dependencies of multiscale with self-attention mechanism, where the weights are determined by measuring the similarity between features. What is more, two fully contextual modules (FCMs) called pyramid fully contextual module (Pyramid-FCM) and atrous spatial pyramid fully contextual module (ASP-FCM) are separately developed to obtain the contextual information that simultaneously lying across positions, channels, and features when combining the intrafeature information aggregation algorithms with SAM on the foundation of existing multiscale modules, such as pyramid pooling (PP) in PSPNet and atrous spatial pyramid pooling (ASPP) in DeeplabV3. We design four schemes for FCMs to obtain more effective contexts. The corresponding FullyContNet-Pyramid and FullyContNet-ASP are separately constructed based on the Pyramid-FCM or ASP-FCM. There are extensive quantitative and qualitative experiments are conducted, depicting the capability of SAM and FCMs and demonstrating the competitiveness of proposed networks on four public hyperspectral scenes when comparing with the current state-of-the-art approaches.
Di Wang 0023, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Adaptive Spectral-Spatial Multiscale Contextual Feature Extraction for Hyperspectral Image Classification
abstract
In this article, we propose an end-to-end adaptive spectral-spatial multiscale network to extract multiscale contextual information for hyperspectral image (HSI) classification, which contains spectral feature extraction (FE) and spatial FE subnetworks. In spectral FE aspect, different from previous methods where features are obtained in a single scale, which limits the accuracy improvement, we propose two schemes based on band grouping strategy, and the long short-time memory (LSTM) model is used for perceiving spectral multiscale information. In spatial subnetwork, on the foundation of existing multiscale architecture, the spatial contextual features which are usually ignored by previous literature are successfully obtained under the aid of convolutional LSTM (ConvLSTM) model. Besides, a new spatial grouping strategy is proposed for convenience of ConvLSTM to extract the more discriminative features. Then, a novel adaptive feature combining way is proposed considering the different importance of spectral and spatial parts. Experiments on three public data sets in HSI community demonstrate that our methods achieve competitive results compared with other state-of-the-art methods.
Di Wang 0023, Bo Du 0001, Liangpei Zhang 0001, Yonghao Xu
IEEE Trans. Geosci. Remote. Sens.1